Skip to main content

Cloud

What We Check First in a Cloud Cost Review

The first questions in a useful cloud cost review are about attribution, architecture and risk — not a blanket rightsizing percentage.

By Protogenies Cloud Practice · 8 min read · Updated 2026-08-30

A cloud cost review should not begin with a list of instance types to downsize. Before changing capacity, we need to know whether the bill can be attributed, which workloads are stable, where the architecture itself creates recurring cost, and which controls would stop the same waste returning. The order matters because a cheap system that is harder to recover or operate is not necessarily an improvement.

1. Can the bill be explained?

The first pass is allocation. We map the largest costs to accounts, subscriptions, projects, environments, products and teams. Untagged or shared spend is not treated as a rounding error; it is a signal that optimization will be hard to verify later. A useful allocation model is simple enough that finance and engineering can both reproduce it.

  • Identify the largest services and the resources underneath them.
  • Separate production, non-production and shared-platform spend.
  • Find unallocated costs and decide how shared services should be assigned.
  • Check whether tags or labels are enforced in infrastructure code or only requested by policy.

2. Which costs exist because nobody owns the lifecycle?

Idle resources are the obvious category, but the more useful question is why they survive. Old snapshots, detached volumes, test clusters and non-production systems running continuously usually point to a missing lifecycle rule rather than a one-time cleanup task. We record the control that should prevent recurrence alongside the resource to remove.

3. Which resources are genuinely over-provisioned?

Rightsizing is useful only when utilization data represents normal and peak demand. We compare capacity with observed usage, scaling limits and resilience requirements. A database that looks quiet on average may still need headroom for a weekly batch or failover event. The goal is to remove capacity that has no operational purpose, not to chase a low utilization number in isolation.

4. Where is the architecture generating the bill?

Some costs cannot be fixed from the billing console because they are consequences of system design. Cross-zone traffic, data replication, chatty service boundaries, hot-tier retention and a managed service chosen for convenience can all be rational decisions. They still need to be visible as architecture trade-offs so the team can decide whether the operating benefit is worth the recurring cost.

The useful output is not 'this service is expensive.' It is 'this cost exists because of this design choice, and these are the consequences of changing it.'

5. Are commitments being used as optimization or camouflage?

Reservations, Savings Plans and committed-use discounts are good tools for stable baseline demand. They are a poor substitute for understanding usage. We separate the baseline that is likely to remain from temporary growth, idle capacity and architecture waste before recommending a commitment. Otherwise the organization can lock in the cost it was trying to reduce.

6. Can the team verify the change next month?

Every recommendation should have an owner, a measurement and a way to confirm that the change persisted. That usually means team-level showback, anomaly alerts, budget ownership and a recurring review. A cost review that produces only a spreadsheet creates another one-time project; a useful review leaves the estate easier to explain after the consultants leave.

If you are preparing for a review, start with a recent billing export, the architecture around the largest services and a person who can explain unusual workload patterns. That is enough to distinguish low-risk cleanup from decisions that need deeper engineering work.

About the author

Protogenies Cloud Practice

Cloud architects and platform engineers who design landing zones, migrations and cost models on AWS, Azure and GCP.

See what this team does

Practical resource

Talk to the engineers who would do the work

Bring your current architecture, constraints and the problem you are trying to solve. We will tell you what we would change first, what it depends on, and where we would start.