Tell me about a time your eval data said one thing and an important stakeholder insisted the opposite. How did you resolve it?
What they are really testing: Whether measurement survives contact with power, and whether you treat anecdotes as noise or as signal about eval blind spots. The senior move is taking the anecdote seriously as data about your data.
A real interview question
Tell me about a time your eval data said one thing and an important stakeholder insisted the opposite. How did you resolve it?
What most people say
drag me
“A director thought the model was worse but our metrics showed it was better, I walked them through the eval results and they came around.”
The resolution is "I showed them my dashboard until they stopped arguing", which assumes the eval was right. Not investigating their examples means a real blind spot would have survived, and the director would have been correct and dismissed.
The follow-ups they ask next
What if his examples had turned out unrepresentative?
Show the trace-through respectfully, here is what the data says across your team 50 cases, keep the two cases in the suite anyway, and credit the check: the process of taking it seriously is what maintains trust for the time they are right.
How do you build eval sets that avoid this class of blind spot?
Sample from real traffic stratified by segment, workflow and input length, weight by stakes rather than pure volume, review coverage with the teams who live in the product, and treat every escaped complaint as a missing slice.
What the interviewer is listening for
- First move was investigating the anecdote, not defending the dashboard
- Found and named the blind spot: slice coverage versus traffic reality
- Turned it into standing process: slice reporting and a trace-it-in-a-day rule
What sinks the answer
- Used the aggregate metric to win the argument rather than test it
- Story ends with the stakeholder being wrong, conveniently
- No change to how evals get built afterwards
If you genuinely do not know
Say this instead of freezing. Reasoning out loud from what you do know beats silence every single time, and a good interviewer is listening for exactly that.
“Tell it as [eval said X, stakeholder said not-X, both credible], [I traced their actual cases and checked eval coverage], [found the blind spot and fixed the slice, or showed the data respectfully], and [the process change: slice-aware evals and a standing trace-anything offer].”
Keep going with behavioural
Junior
This field changes monthly. Tell me about a time something you had built became outdated fast, and how you handled it.
Mid
Tell me about a time an AI feature you shipped behaved badly in production. What happened and what did you change?
Mid
Describe a time you pushed back on using AI for something. How did you make the case, and what happened?
Mid
Tell me about a technical decision you got wrong on an AI project. How did you find out, and what did you do?
Mid
Describe a time you were pressured to ship an AI feature before you thought it was ready. What did you do?
Senior
Tell me about an AI feature that worked technically but users did not adopt or trust. What did you learn?
Knowing the answer is not the same as recalling it under pressure
Sign in to send the questions you fumble to spaced recall, so they come back right before you would forget them, and learn the concepts behind them with hands-on labs.
Start free