Finance pings you on a Monday: the cloud bill for the weekend is double the usual run rate, roughly 40,000 dollars ahead of forecast. Nobody knows why. What do you do?
What they are really testing: They are testing whether you treat cost as an operational incident with a stop the bleeding phase, and whether you can attribute spend without accusing anyone.
A real interview question
Finance pings you on a Monday: the cloud bill for the weekend is double the usual run rate, roughly 40,000 dollars ahead of forecast. Nobody knows why. What do you do?
What most people say
drag me
“I would tell everyone to turn off anything they are not using and start a cost optimisation review across the whole estate.”
A broad optimisation review is a quarter of work, not an incident response, and it will not find the one thing that bent the curve on Saturday.
The follow-ups they ask next
The spike is in an account nobody claims. How do you find the owner?
Go to CloudTrail for the creating principal, then the tag policy, then the payer account structure, and if there is genuinely no owner, say who you would escalate to.
What alarm would you have set to catch this on Saturday at 03:00?
Give a real number: an anomaly detector or hourly budget alarm on unblended cost per service, tuned so the on call sees it after roughly one hour of abnormal spend, not after the month closes.
What the interviewer is listening for
- Slices the bill by service, region and account before theorising
- Stops the burn before hunting root cause
- Correlates the exact hour the curve bent with the change log
- Ends with an alarm and a guardrail, not a lecture
What sinks the answer
- Starts a broad optimisation project instead of an incident response
- Hunts for someone to blame before the spend is stopped
- Guesses at the cause without ever opening the billing data
- Cannot name a single signal or alarm they would add
If you genuinely do not know
Say this instead of freezing. Reasoning out loud from what you do know beats silence every single time, and a good interviewer is listening for exactly that.
“I have not owned a bill spike of that size, but here is my reasoning: find the exact hour and the exact service that changed, stop the burn safely, then line that hour up against what we shipped.”
Keep going with troubleshooting
Junior
You get paged at 2am. Checkout latency has gone from 200ms to 4 seconds and it started right after the 11pm deploy. Walk me through what you do.
Junior
It is 3am and you get paged: the primary application server is at 100% disk usage and the app is throwing write errors. You are the only one online. Walk me through exactly what you do.
Junior
You push a new version of a service and within two minutes the pods are all showing CrashLoopBackOff. The team is watching you share your screen. Walk me through exactly what you do.
Junior
A developer pings you: the checkout service cannot reach the payments service, the calls just hang and time out. You have never worked on either service. Walk me through what you do.
Junior
A developer messages you: 'My app suddenly gets Access Denied writing to the S3 bucket. Nothing changed.' It worked yesterday. Walk me through what you do.
Mid
Your load balancer shows 5xx errors jumping from 0.1 percent to 12 percent of requests over ten minutes. There was no deploy. What do you do?
Knowing the answer is not the same as recalling it under pressure
Sign in to send the questions you fumble to spaced recall, so they come back right before you would forget them, and learn the concepts behind them with hands-on labs.
Start free