EC2 Auto Scaling Reservations-Then-Balanced: How to Stop Paying for AWS Capacity You Never Launch Into

Most businesses that buy reserved AWS capacity assume the money is safe. You have committed to the spend, you have locked in a discount, and the finance team has a number they can plan around. What almost nobody checks is whether the instances your platform actually launches ever land inside that reservation. Auto Scaling has historically spread instances evenly across Availability Zones to protect resilience, and it did that without much regard for where your pre-purchased capacity was sitting. The result is a quiet and expensive failure mode: you pay for a reservation in one zone while your Auto Scaling group launches on-demand instances in another.
On 30 June 2026, AWS shipped a fix. Amazon EC2 Auto Scaling now supports a distribution strategy called reservations-then-balanced. It tells an Auto Scaling group to launch into your capacity reservations first, and only then spread whatever remains evenly across Availability Zones. You configure it in the AvailabilityZoneDistribution setting of the group and target reservations either by Capacity Reservation Group ARN or by individual reservation IDs. There is no extra charge for the feature, it is available in all commercial Regions, and you continue to pay standard EC2 pricing for the reservations and for any on-demand or Spot instances the group launches. In short, it closes the gap between what you bought and what you run.
Why Reserved Capacity Sits Idle While the Bill Keeps Growing
The mechanics are unglamorous. An On-Demand Capacity Reservation is zonal. When you buy one, you are reserving capacity in a specific Availability Zone, not in the Region as a whole. Auto Scaling groups, meanwhile, have traditionally used a balanced distribution strategy that aims to keep instance counts roughly even across the zones you have enabled. Those two behaviours do not coordinate. If you hold twenty reserved instances in eu-west-2a and your group scales out by six, Auto Scaling may put two in each of three zones. Four of those instances launch outside your reservation at on-demand rates, and four slots inside the reservation continue to bill you for nothing.
This is the sort of thing that never shows up as an incident. There is no alarm, no failed deployment and no unhappy customer. It shows up months later as a line in a cost review that nobody can explain, usually described as an unexplained variance. Teams often respond by buying more commitments, which makes the underlying problem worse rather than better. We have seen this pattern repeatedly in cost assessments, and the root cause is almost always the same: nobody owns the question of whether reserved capacity and launched capacity are actually the same capacity.
What the Data Says About Commitment Waste
The wider picture is not encouraging. Flexera’s 2026 State of the Cloud Report puts estimated wasted cloud spend at 29 percent, the first increase in five years, and attributes the reversal to the surge in cloud-based AI workloads and the cost complexity that comes with them. That is happening at a time when organisations are buying more commitments than ever, which means more money is at risk of landing in the wrong place. ProsperOps analysed anonymised AWS usage and found that 64 percent of organisations used Reserved Instances or Savings Plans in 2024, up sharply from 45 percent the year before, and that median commitment coverage more than doubled from 28 percent to 55 percent over the same period.
The interesting part is where the savings actually land. Median AWS Compute Effective Savings Rate rose from 0 percent in 2023 to 15 percent in 2024, but the top performers barely moved, going from 46 percent to 47 percent at the 98th percentile. Adoption is rising fast, and effectiveness is not keeping pace. Organisations covered only 47 percent of their usage with commitments on average, precisely because commitments carry placement and utilisation risk. Anything that reduces that risk, including a scheduler that reliably fills reservations, makes higher coverage safer to buy.
| Metric | 2023 | 2024 or latest | Source |
|---|---|---|---|
| AWS organisations using RIs or Savings Plans | 45% | 64% | ProsperOps, 2025 Rate Optimization Insights |
| Median AWS compute commitment coverage | 28% | 55% | ProsperOps, 2025 Rate Optimization Insights |
| Average commitment coverage of usage | 37% | 47% | ProsperOps, 2025 Rate Optimization Insights |
| Median AWS Compute Effective Savings Rate | 0% | 15% | ProsperOps, 2025 Rate Optimization Insights |
| Effective Savings Rate at 98th percentile | 46% | 47% | ProsperOps, 2025 Rate Optimization Insights |
| Estimated wasted cloud spend | trending down | 29%, first rise in five years | Flexera, 2026 State of the Cloud Report |
| Organisations citing cloud spend management as a top challenge | not stated | 85% | Flexera, 2026 State of the Cloud Report |
| Practitioners naming optimisation a top priority | not stated | 50% | FinOps Foundation, State of FinOps 2025 |
Methodology note: figures were retrieved on 20 July 2026 from the Flexera 2026 State of the Cloud Report and its accompanying press release, the ProsperOps 2025 AWS Compute Rate Optimization Insights report based on anonymised customer usage data, and the FinOps Foundation State of FinOps 2025 survey covering organisations responsible for more than 69 billion dollars of cloud spend. Sample sizes, definitions and measurement years differ between these sources, so the rows above should be read as directional indicators from separate studies rather than a single like-for-like dataset. Coverage and Effective Savings Rate figures are AWS compute specific. Waste figures cover public cloud spend generally.
How Reservations-Then-Balanced Actually Works
The behaviour is easy to reason about, which is the best thing about it. You set the capacity distribution strategy on the Auto Scaling group’s AvailabilityZoneDistribution configuration to reservations-then-balanced, and you tell it which reservations to target. Targeting by Capacity Reservation Group ARN is the option most teams should reach for, because it lets you add and remove individual reservations from the group without touching the Auto Scaling configuration or redeploying anything. Targeting individual Capacity Reservation IDs is more precise and more brittle, and it suits a fixed, long-lived reservation you rarely change.
When the group scales out, it fills the targeted reservations first, wherever they happen to live. Once those are exhausted, it reverts to standard balanced behaviour and spreads remaining instances evenly across your enabled Availability Zones. That means the feature is additive rather than disruptive. In a steady state where your reservations comfortably cover your baseline, you get near-complete reservation utilisation. In a spike where demand exceeds what you have reserved, you get the same resilience characteristics you had before. The feature supports On-Demand Capacity Reservations, Capacity Blocks and Interruptible Capacity Reservations, which is where it gets more interesting.
Where This Matters Most: GPU Workloads and Capacity Blocks
If your business runs anything GPU-heavy, this is not a marginal optimisation. Capacity Blocks exist because accelerated instances are scarce and you cannot rely on being able to launch a P-series or Trainium instance on demand when you need one. Teams buy Capacity Blocks precisely so a training run or an inference fleet has somewhere to land. Until now, an Auto Scaling group managing that fleet could plausibly fail to use the block you paid handsomely for and then fail to launch elsewhere because the capacity simply was not available. That is the worst of both outcomes: a large bill and an outage.
The timing is deliberate. Flexera’s data shows generative AI rose to the third most widely used public cloud service in 2026, at 58 percent adoption, up from 50 percent, and that AI workloads are the main reason wasted spend has started climbing again. AWS also cut EKS Auto Mode and ECS Managed Instances GPU management fees from 1 July 2026, with P-series and Trainium fees down 60 percent and G-series down 35 percent. The pattern across all of these changes is the same. AWS is making accelerated compute cheaper to manage and easier to place, because a lot of customers are currently doing both badly.
The Resilience Trade-Off Worth Thinking About
There is an honest caveat here. Balanced distribution exists for a reason. If your reservations are concentrated in one Availability Zone and reservations-then-balanced diligently fills them first, your fleet will be more zonally concentrated than it used to be, at least until the reservations are exhausted. For a stateless web tier behind a load balancer with healthy scaling headroom, that is usually an acceptable trade. For a workload where losing a zone means losing quorum, it is not, and you should either spread your reservations across zones deliberately or leave the group on balanced distribution.
The right way to make this decision is to look at what your reservations actually look like today rather than what you intended when you bought them. Pull your Capacity Reservation utilisation from the console or the CLI, compare it against your Auto Scaling group’s actual instance distribution, and see whether the two line up. In most environments we assess, they do not, and the answer is not simply to turn on a new flag. It is to rebalance the reservations first so that filling them preferentially does not create a single-zone dependency you never signed up for.
What to Do About It
Start by measuring, because you cannot fix a utilisation problem you have not quantified. Look at every On-Demand Capacity Reservation and Capacity Block in your account, record its zone and its actual utilisation over the last thirty days, and identify any reservation running below roughly 90 percent. Then map each underused reservation to the Auto Scaling groups that should have been consuming it. That exercise alone usually recovers real money, whether or not you enable anything new.
Once you know where the gaps are, enable reservations-then-balanced on the groups where the resilience profile allows it, prefer Capacity Reservation Group ARNs so your configuration stays stable as reservations come and go, and put a recurring check in place so utilisation is reviewed monthly rather than discovered at the end of a quarter. If your reservations are badly distributed across zones, fix that before you change the scaling behaviour. Finally, revisit your coverage target. The ProsperOps data suggests most organisations under-commit because they do not trust their own placement, and if you can now trust it, buying more coverage becomes a sensible next move rather than a gamble.
If you are not sure whether your AWS commitments are actually being used, HAZERCLOUD can help you find out. We run cost and architecture assessments for growing UK, US and European businesses, covering commitment utilisation, Auto Scaling configuration, zonal resilience and the GPU capacity decisions that increasingly drive the bill. Get in touch at https://hazercloud.com/contact/ for a free consultation or migration assessment, and we will tell you plainly what is working and what is quietly costing you money.