MidBehavioural

Tell me about a time an AI feature you shipped behaved badly in production. What happened and what did you change?

What they are really testing: Whether you have really operated one of these systems, and whether failure turned into process. Candidates who have never had an LLM feature surprise them in production have not shipped one that mattered.

A real interview question

Tell me about a time an AI feature you shipped behaved badly in production. What happened and what did you change?

What most people say

drag me

Our chatbot once gave some wrong answers, so we improved the prompt and it got better, and we learned to test more.

No stakes, no mechanism, no numbers, and "test more" is not a change, it is a wish. The interviewer learns nothing except that either the failure was trivial or the reflection was.

The follow-ups they ask next

  • How did you communicate this to non-engineering stakeholders?

    Impact first in their terms, customers affected, cost, then the fix and the prevention, without model mysticism: "it read the wrong document, here is how we made that impossible" lands better than embeddings talk.

  • What would you have needed to catch it before support did?

    A groundedness or citation check on money-topic answers, and feedback-rate monitoring sliced by topic, both of which existed after. The honest answer names the specific missing alarm, not "more testing".

What the interviewer is listening for

  • Quantifies user and business impact without being asked
  • Diagnosis reached the mechanism, retrieval, not just the symptom
  • Failure became eval cases and a release rule, not a resolution to be careful

What sinks the answer

  • Cannot produce a single concrete failure from their AI work
  • Blames the model or the provider with no mechanism
  • Fix was a prompt tweak and hope, no structural change

If you genuinely do not know

Say this instead of freezing. Reasoning out loud from what you do know beats silence every single time, and a good interviewer is listening for exactly that.

Structure it as [feature, failure, impact in numbers], then [how traces revealed the mechanism], then [immediate fix] versus [structural prevention: eval cases, guardrails, release rules], ending with [the rule you now apply everywhere].

Keep going with behavioural

All 57 ai engineer questions

Knowing the answer is not the same as recalling it under pressure

Sign in to send the questions you fumble to spaced recall, so they come back right before you would forget them, and learn the concepts behind them with hands-on labs.

Start free