Cloud cost reviews have a habit of starting the same way. Someone opens the billing console, looks at the total and says "that can't be right". Then they sort by service and find a database somebody spun up for a demo a year and a half ago, still running.

Cloud spend rarely goes wrong because of one big mistake. It goes wrong through lots of small ones that each seemed reasonable at the time. A staging environment mirrors production "just to be safe". An instance size gets picked before anyone knows the real load. A logging setting stays on its default. None of it looks alarming on its own, but together it can be a big share of the bill.

The first round of savings is usually the easiest. Here's where we look, roughly in order of how much it tends to recover for how little effort.

1. Things running for no reason

Start with the dull audit. List every resource and ask who owns it and what it's for.

You'll almost always find orphaned environments: proofs of concept, old feature branches, a "temporary" migration server. You'll find disks left behind when their virtual machines were deleted, and years of automated snapshots with no expiry. There'll be idle load balancers and public IPs pointing at nothing. And there'll be dev and test environments that nobody touches outside business hours but that you pay for all 168 hours of the week.

Scheduling non-production environments to shut down overnight and at weekends is one of the best-value changes you can make. If your team works roughly business hours, those environments sit idle for most of the week.

The lasting fix is tagging. Every resource gets an owner, an environment and a cost centre, applied automatically by your infrastructure-as-code, and anything untagged gets flagged. Once every line on the bill has a name next to it, orphaned resources become much rarer.

2. Things bigger than they need to be

Instance sizes tend to be chosen once, early, by someone guessing at the load, and then left alone. Look at real CPU and memory use over a month. A server that averages low single-digit CPU and rarely climbs higher is paying for headroom it never uses.

Databases are the same. Managed databases are often sized for a peak that comes once a quarter. Where the platform supports it, serverless or auto-scaling tiers can match cost to actual demand, which makes a real difference for spiky work like month-end reporting or seasonal traffic.

Resize carefully. Change one thing, watch it for a week, then move on. You're trying to remove waste, not add risk.

3. Data going where it shouldn't

Data transfer surprises people more than any other line item, because it never appears on an architecture diagram. Getting data into the cloud is generally free. Getting it out, or moving it between regions and availability zones, generally isn't.

Common culprits include chatty services in different zones or regions throwing large payloads at each other all day, and NAT gateways charging per gigabyte for traffic that could have used a private endpoint to the provider's own services. Large files served straight from storage rather than through a CDN are another, since the CDN is usually cheaper per gigabyte and faster for users. So are backups and replicas copied to another region more often than your recovery targets require.

Most providers have a cost explorer that groups spend by usage type. Filter it for data transfer, and if the number surprises you, dig in.

4. Logs nobody reads

Log ingestion and retention is a quiet cost that grows quickly. Debug logging left on in production, verbose request logs on health-check endpoints, and months of retention on logs nobody has looked at since the week they were written all add up. Decide on purpose what you log, at what level and for how long. If compliance says you have to keep older logs, move them to cheap cold storage. You'll almost never search them, and they'll cost a fraction as much to keep.

5. Paying full price for steady workloads

Once the waste is gone, look at what's left. If a workload runs around the clock and will keep doing so, paying on-demand rates leaves money on the table. Every major provider offers commitment discounts, such as reserved capacity or savings plans, that give you a significant reduction in return for a one- or three-year commitment.

Order matters here. Commit after you've right-sized, or you'll lock in a discount on capacity you didn't need.

For work that can survive being interrupted, like batch jobs, CI runners and rendering, spot or pre-emptible capacity is dramatically cheaper and worth designing around.

Make it a habit

A one-off clean-up feels great and then slowly unravels. Teams that keep costs down treat cost as an engineering metric, the same way they treat performance or uptime. They put budgets and anomaly alerts on every account, so a runaway resource gets noticed in a day rather than a quarter. They show cost next to each team's services, so the people who can change it can see it. And they spend a few minutes each month on the biggest movers: what went up, and why.

You don't need a dedicated FinOps team for any of this. You need someone who looks regularly, and tooling that makes looking easy.

If your cloud bill is growing faster than your business, we can run a structured review and hand you a prioritised list of savings, along with the guardrails to stop the waste coming back.