Hybrid is usually an outcome, not a decision
Something moved to cloud, something could not, and the join between them was never designed — only discovered.
A VPN or Direct Connect circuit nobody has failed over on purpose, and no clear answer to what breaks when it drops.
Untested pathAn on-prem directory and a cloud one, synced by something configured years ago, with the failure mode understood by nobody currently employed.
Fragile syncOne monitoring stack per side, no shared timeline, and incidents that take an hour to place because nothing correlates.
No single viewAn application in cloud that calls a database on-prem on every request, and a performance profile nobody planned.
Undocumented couplingWe run both sides. Deciding what moves is a separate job.
Where this sits relative to the work either side of it, so you can tell what you are and are not buying.
On-Prem
When the workload is staying in your own racks and that is the whole picture.
Go to the pageHybrid, operated
Both sides run as one estate: the connection, identity, observability and the failure modes between them — with one party accountable for the seam rather than two suppliers owning an end each.
Cloud Migration
If the goal is for the hybrid state to be temporary rather than permanent.
Go to the pagePlenty of clients buy one layer and run the rest themselves. That is a normal outcome, not a half-sale.
What running hybrid actually involves
The connection
Direct Connect, ExpressRoute, Interconnect or IPsec, sized against real traffic and with a rehearsed failover — plus routing and DNS that resolve the same way from both sides, which is where most hybrid incidents actually start.
Identity across the seam
One directory as the source of truth, sync you can explain and monitor, and a documented answer to what authentication does when the link is down.
Workload placement
Written down per workload: what stays, what moves, what cannot move and why — usually latency, licensing, data residency or a device on a factory floor. Placement decisions with reasons attached, not preferences.
Data and replication
Which side owns each dataset, how it replicates, what the recovery objective is, and a restore rehearsed on both sides rather than assumed.
One observability plane
Metrics and logs from both sides in one place with a shared timeline, so an incident can be placed in minutes. Prometheus, Grafana and Loki where they fit; provider-native tooling where it is genuinely better.
Change and on-call
One pipeline that can deploy to either side, Terraform describing both, and one owner for an incident that crosses the seam rather than a handoff between two suppliers.
Investigate, implement, hand over
What is genuinely running: accounts, workloads, identity, network, backups, cost, and how a change reaches production today. You get a written picture of the current state and a scope in priority order — what is dangerous, what is expensive, and what can wait.
The priority items, in sequence, as code in your repositories. Rehearsed in the lower environments first, with the rollback run rather than described. Nothing goes in that your own engineers cannot read.
Runbooks, architecture decisions and the reasoning behind them, written down. A working session per area with your team. We stay on-call afterwards only if you want us to.
Your provider covers one side. Count the seam.
This is the honest version of a comparison most pages skip. The provider takes real work off you — and it takes one layer, not the stack.
Your cloud provider
Their side of the circuit, the platform underneath it, and the SLAs on both. Genuinely worth paying for — and it stops at their edge.
What is left
The routing and DNS that has to agree across both sides. Identity sync and its failure mode. Placement decisions. Replication and recovery. A correlated view of incidents. And the on-call rota that covers all of it.
What you are actually choosing between
| Option | Left as code | Runs day to day | Actively cuts cost | Security you can evidence | No key-person risk | Best fit when… |
|---|---|---|---|---|---|---|
| Hire in-house | the work is genuinely continuous — then hiring is the right call. | |||||
| Generalist MSP | you only need the lights kept on. | |||||
| Leave it alone | nothing has broken, been audited, or been queried yet. | |||||
| OpsHero | the work comes in bursts and has to outlast us. |
Hybrid is not right for everybody
If nothing genuinely has to stay on-prem, hybrid is usually the most expensive resting place — two operating models, two failure domains, one team. In that case the honest recommendation is to finish the migration, and Cloud Migration is the page you want. We would rather say that on a page than halfway through an engagement.
None of it is true yet. Worth a conversation.
Your answers stay in your browser.
The ones people actually ask
Is hybrid a permanent answer or a stage?
Both exist. Latency, licensing, data residency and physical hardware are real constraints and make hybrid permanent. Inertia is not a constraint, and makes it a stage. We will tell you which we think you have.
Do you work in our data centre?
We work on the systems in it. Physical hardware, hands and remote access are arranged with you or your facility provider — we will tell you plainly which parts we do not cover.
Can you run only the cloud side?
Yes, though the seam is usually where the interesting failures live. If you split the rota across two providers, agree in writing who owns the connection — otherwise nobody does.
Which side should a workload be on?
The one with a written reason. We produce a placement decision per workload with the constraint that drove it, so the answer survives the next person to ask.
Will we be locked in?
Your repositories, your accounts, your hardware — documented and reproducible on both sides.
Talk to a hybrid engineer
Not a salesperson. You will get a read on your estate, the specific risks in it, and what the work would involve.
- A written picture of what is actually running, and what state it is in
- The risks in priority order — dangerous, expensive, and can wait
- What the work would involve, and in what sequence
- Everything delivered into your own repositories and cloud accounts
