The problem: deallocate at your own risk
Say a VM is misbehaving and you do what every admin has done a hundred times: deallocate it and start it back up to clear a transient issue. When you deallocate, Azure releases the underlying hardware. When you try to start it again, that capacity may already belong to someone else, and your VM fails to start with an allocation error.
In constrained regions like East US, this isn't hypothetical. The old fix-it move is now a gamble.
Quota isn't capacity
A lot of people assume that if their subscription has quota for a VM size, they can deploy it. Not so. Quota is permission to deploy. Capacity is whether Azure physically has the hardware in that region or zone. You can have plenty of quota and still get turned away.
Reserved Instances won't save you either
If you bought 1- or 3-year Reserved Instances, you might think you're covered. You're not. Reserved Instances are a billing discount, not a capacity guarantee. They make your VMs cheaper; they don't make them available.
The fix costs full price, whether you use it or not
On-Demand Capacity Reservations do guarantee capacity, backed by an SLA, with no long-term commitment. You pick a VM size, region (and optionally a zone), and quantity, and Azure sets that hardware aside until you delete the reservation.
Here's the catch: you pay the full pay-as-you-go rate for every reserved slot, whether a VM is running in it or not. Reserve ten VMs, run six, and you're billed for ten. Your Reserved Instance and Savings Plan discounts can apply to the unused slots, which helps, but the core math is simple. To safely turn a VM off, you have to keep paying for it as if it were on.
That's the "pay only for what you use" model turned inside out. The risk of scarcity has been shifted onto the customer, and the price of avoiding that risk is paying for idle capacity, exactly the waste the cloud was supposed to eliminate.
And the scarcest hardware isn't covered
The kicker: some of the most in-demand GPU families, including the ND-series and H100-based NCads v5-series, aren't supported for capacity reservations at all. Neither are Spot VMs, availability sets, proximity placement groups, VMs using Ultra Disk, or VMs resuming from hibernation. If your workload depends on any of these, you're on your own.
The SLA is also worth reading closely. It's backed by service credits, not an absolute guarantee. If reserved capacity isn't available when you need it, you get a percentage back on your bill, not your VM.
To be fair
Microsoft can't conjure data center hardware out of thin air, and GPU and memory shortages are industry-wide. Capacity reservations are a reasonable tool for disaster recovery and mission-critical scale-out. The problem isn't that the option exists. It's that the original promise of elastic, pay-per-use computing increasingly comes with an asterisk, and the cost of that asterisk lands on customers.
What you can do
Know which of your workloads can't tolerate a failed restart, and reserve capacity only for those. VMs have to explicitly opt in to a reservation, so you can protect critical systems without paying to reserve everything.
Think twice before deallocating in constrained regions. A restart from the portal or inside the OS keeps the VM on its host. Only deallocate when you have to.
Build flexibility into your architecture: spread across availability zones, keep alternate VM sizes approved for your workloads, and consider less congested regions for new deployments.
Finally, check whether your VM series is even supported before you plan around reservations. Microsoft's capacity reservation documentation lists supported sizes.
The cloud still offers a lot. But "turn it off and stop paying" now comes with fine print, and it's worth reading before your next routine restart becomes an outage.