Skip to content
Semi-Functional Tech
Color theme

Pay for What You Use? Azure Would Now Like You to Pay for What You Might Need

By Brad Hager4 min read

The enshittification of the cloud is real.

Remember the pitch? Move your on-premises workloads to the cloud and you'd only pay for what you used. Turn servers off when you don't need them, spin them back up when you do, and watch the savings roll in. Elasticity was the whole point.

That promise is quietly eroding. With the AI boom driving extreme demand for GPUs, memory, and storage, compute in popular Azure regions is getting scarce. Microsoft's answer is On-Demand Capacity Reservations, which let you pay to guarantee that a specific VM size will be there when you need it.

The problem: deallocate at your own risk

Say a VM is misbehaving and you do what every admin has done a hundred times: deallocate it and start it back up to clear a transient issue. When you deallocate, Azure releases the underlying hardware. When you try to start it again, that capacity may already belong to someone else, and your VM fails to start with an allocation error.

In constrained regions like East US, this isn't hypothetical. The old fix-it move is now a gamble.

Quota isn't capacity

A lot of people assume that if their subscription has quota for a VM size, they can deploy it. Not so. Quota is permission to deploy. Capacity is whether Azure physically has the hardware in that region or zone. You can have plenty of quota and still get turned away.

Reserved Instances won't save you either

If you bought 1- or 3-year Reserved Instances, you might think you're covered. You're not. Reserved Instances are a billing discount, not a capacity guarantee. They make your VMs cheaper; they don't make them available.

The fix costs full price, whether you use it or not

On-Demand Capacity Reservations do guarantee capacity, backed by an SLA, with no long-term commitment. You pick a VM size, region (and optionally a zone), and quantity, and Azure sets that hardware aside until you delete the reservation.

Here's the catch: you pay the full pay-as-you-go rate for every reserved slot, whether a VM is running in it or not. Reserve ten VMs, run six, and you're billed for ten. Your Reserved Instance and Savings Plan discounts can apply to the unused slots, which helps, but the core math is simple. To safely turn a VM off, you have to keep paying for it as if it were on.

That's the "pay only for what you use" model turned inside out. The risk of scarcity has been shifted onto the customer, and the price of avoiding that risk is paying for idle capacity, exactly the waste the cloud was supposed to eliminate.

And the scarcest hardware isn't covered

The kicker: some of the most in-demand GPU families, including the ND-series and H100-based NCads v5-series, aren't supported for capacity reservations at all. Neither are Spot VMs, availability sets, proximity placement groups, VMs using Ultra Disk, or VMs resuming from hibernation. If your workload depends on any of these, you're on your own.

The SLA is also worth reading closely. It's backed by service credits, not an absolute guarantee. If reserved capacity isn't available when you need it, you get a percentage back on your bill, not your VM.

To be fair

Microsoft can't conjure data center hardware out of thin air, and GPU and memory shortages are industry-wide. Capacity reservations are a reasonable tool for disaster recovery and mission-critical scale-out. The problem isn't that the option exists. It's that the original promise of elastic, pay-per-use computing increasingly comes with an asterisk, and the cost of that asterisk lands on customers.

What you can do

Know which of your workloads can't tolerate a failed restart, and reserve capacity only for those. VMs have to explicitly opt in to a reservation, so you can protect critical systems without paying to reserve everything.

Think twice before deallocating in constrained regions. A restart from the portal or inside the OS keeps the VM on its host. Only deallocate when you have to.

Build flexibility into your architecture: spread across availability zones, keep alternate VM sizes approved for your workloads, and consider less congested regions for new deployments.

Finally, check whether your VM series is even supported before you plan around reservations. Microsoft's capacity reservation documentation lists supported sizes.

The cloud still offers a lot. But "turn it off and stop paying" now comes with fine print, and it's worth reading before your next routine restart becomes an outage.

Share this post