It is 2pm, peak traffic. Half the targets behind your ALB just flipped to unhealthy, and the surviving half are now taking double the load and starting to shed requests. Walk me through the next fifteen minutes.
What they are really testing: They want to know if you protect the healthy half from a cascading collapse before you go looking for why the unhealthy half died.
A real interview question
It is 2pm, peak traffic. Half the targets behind your ALB just flipped to unhealthy, and the surviving half are now taking double the load and starting to shed requests. Walk me through the next fifteen minutes.
What most people say
drag me
“I would look at the CloudWatch dashboards and try to work out why the instances went unhealthy, then fix the root cause.”
It goes hunting for a root cause while the remaining half of the fleet is being overloaded, which is exactly how a partial outage becomes a total one.
The follow-ups they ask next
The health check reason is Timeout, but you can curl the target from your own box. Explain that.
The ALB uses its own security group and a specific health check port. Curl from your box does not prove the ALB can reach it.
Every unhealthy target is in one availability zone. What do you do?
Treat it as an AZ level fault. Shift traffic out of that AZ and confirm the subnet, route table and NAT are healthy.
How would you make the health check less trigger happy without hiding real failures?
Separate liveness from readiness, tune the unhealthy threshold and interval, and never make the check depend on a downstream service.
What the interviewer is listening for
- Scales out or rolls back before diagnosing anything
- Uses the health check failure reason to pick the layer
- Checks whether failures cluster in one availability zone
- Distinguishes a symptom from the actual cause
What sinks the answer
- Starts SSHing into instances while the healthy half collapses
- Assumes an unhealthy target is always a dead target
- Never mentions rolling back a recent deploy
- Has no plan for capturing a timeline afterwards
If you genuinely do not know
Say this instead of freezing. Reasoning out loud from what you do know beats silence every single time, and a good interviewer is listening for exactly that.
“I have not run this exact incident, but my instinct is to protect the healthy half first by scaling out or rolling back, and only then ask why the other half died.”
Keep going with troubleshooting
Junior
You get paged at 2am. Checkout latency has gone from 200ms to 4 seconds and it started right after the 11pm deploy. Walk me through what you do.
Junior
It is 3am and you get paged: the primary application server is at 100% disk usage and the app is throwing write errors. You are the only one online. Walk me through exactly what you do.
Junior
You push a new version of a service and within two minutes the pods are all showing CrashLoopBackOff. The team is watching you share your screen. Walk me through exactly what you do.
Junior
A developer pings you: the checkout service cannot reach the payments service, the calls just hang and time out. You have never worked on either service. Walk me through what you do.
Junior
A developer messages you: 'My app suddenly gets Access Denied writing to the S3 bucket. Nothing changed.' It worked yesterday. Walk me through what you do.
Mid
Your load balancer shows 5xx errors jumping from 0.1 percent to 12 percent of requests over ten minutes. There was no deploy. What do you do?
Knowing the answer is not the same as recalling it under pressure
Sign in to send the questions you fumble to spaced recall, so they come back right before you would forget them, and learn the concepts behind them with hands-on labs.
Start free