Most teams pick between cloud and on-prem the wrong way: they treat it as a single, estate-wide verdict driven by whichever bill looks worse this quarter, then lift everything one direction. That is how you end up surprised by the steady-state cloud bill after a blanket adoption, or surprised by the capacity limits and ops load after a blanket repatriation. The better approach is to decide on a short list of factors that genuinely change the answer, score each workload against them, and only then choose a destination. This guide gives you that rubric. The framing to hold onto: cloud’s real value is elasticity, agility, and opex, while on-prem’s real value is steady-state run cost and control, so this is a per-workload question, not a religion.
Start with the decision, not the destination
Before you decide anything, characterize each workload by its shape rather than its name. Write down its demand profile over a day and a week, how much its capacity actually varies, and where its data lives and how much of it moves. Steady, predictable, high-utilization workloads and bursty, variable ones want opposite homes, and the estate-wide answer is almost always “some of each.” Get that per-workload shape down first; the destination falls out of it.
What decides where a workload belongs
- Workload shape. Steady-state and predictable demand favors on-prem economics, because amortized owned hardware running near capacity has a low cost per useful hour. Bursty or variable demand that needs to scale on request favors the cloud, because you pay for use instead of provisioning for a peak you rarely hit.
- Elasticity and agility needs. If you must spin capacity up and down quickly, or ship and iterate fast on managed building blocks, the cloud earns its premium. If the workload is flat and rarely changes, that premium buys little.
- In-house capacity to run infrastructure. On-prem takes back capacity planning, hardware refresh, and the on-call and operational load that managed services quietly absorbed. Be honest about whether you have the team for it.
- CapEx versus opex preference. Owned hardware is up-front capital that amortizes; cloud is operating spend that flexes with use. Your finance model and cash position legitimately shape the answer, not just the technical fit.
- Data gravity. Large or heavily accessed datasets pull their compute toward them. Data already on-prem, or subject to residency rules, resists a cloud move; data born in the cloud resists coming back, and egress puts a real one-time cost on either move.
- Managed-service dependence. The further a workload leans on proprietary managed databases, queues, or serverless, the harder repatriation is, because those have no drop-in on-prem equivalent and you take the work back on yourself.
Where each workload lands
Score each workload against the criteria above, then use these as starting points rather than conclusions. Flat, high-utilization systems that run near capacity around the clock, and workloads with heavy data gravity or residency constraints, are the classic candidates to keep or repatriate to on-premise or a private cloud built on VMware or OpenStack. Workloads with swinging or unpredictable demand, or ones that benefit from fast iteration on managed services, are where adopting AWS, Azure, Google Cloud, or OCI pays for its premium. Many estates land on a deliberate split: elastic and fast-moving workloads in the cloud, steady and predictable ones on owned hardware. The right home falls out of each workload’s shape and your team’s capacity, not out of a single estate-wide preference.
How the estate-wide verdict backfires
Three errors show up repeatedly. The first is the blanket decision, lifting everything into the cloud or pulling everything back, which guarantees a mismatch for whichever workloads had the opposite shape. The second is comparing list rates instead of your real utilization, which flatters whichever side you already favor. The third, on the repatriation path, is underestimating the operational load and capacity planning you take back on, and the proprietary managed services with no on-prem equivalent. A quieter fourth: folding one-time egress into the steady-state comparison, which distorts a recurring decision with a one-off cost.
Weigh it per workload before you commit
Turn your leading direction into a proof with written acceptance criteria: baseline the workload’s steady-state cost against the source with your own utilization, replicate a representative pilot in the intended direction, and verify the data-rehydration or replication path and the capacity before scaling. Prove the cost case per workload before committing hardware or reserved-capacity spend, since both are commitments that are awkward to unwind. Keep egress on its own line as a one-time cost, and keep the source as a fallback through hypercare until the target is validated on function, performance, and cost. Choose each workload’s home on its shape and its numbers, not on where the rest of the estate happens to sit.
Use the TCO calculator to model a per-vCPU comparison with your real utilization.