The CTO asks whether you should build your AI capability on frontier APIs or self-host open-weight models. How do you frame the decision?
What they are really testing: A strategy question with a numerate answer. The tells: knowing the true cost of self-hosting beyond the GPU sticker price, naming the few triggers that genuinely force it, and recommending an architecture that keeps the decision reversible.
A real interview question
The CTO asks whether you should build your AI capability on frontier APIs or self-host open-weight models. How do you frame the decision?
What most people say
drag me
“Self-hosting open models gives us control and avoids per-token fees, so at scale it should be cheaper, I would run the numbers on GPU costs versus our API bill.”
GPU rental versus token fees is the amateur comparison. It omits the serving stack, the utilisation problem, the on-call burden and the quality gap on hard tasks, and it frames as one decision what should be a per-workload portfolio.
The follow-ups they ask next
What actually breaks even, roughly?
Sustained, predictable volume where a small self-hosted model replaces a per-token bill exceeding the loaded cost of GPUs plus an engineer or two, at high utilisation. Bursty or growing-unpredictably traffic favours APIs almost regardless of volume.
The board worries about provider dependence. Answers besides self-hosting?
Multi-provider readiness through the gateway, evals that make switching measurable in a day, contracts with capacity commitments, and open-weight fallbacks kept warm for critical paths, dependence managed, not eliminated at 10x cost.
What the interviewer is listening for
- Frames it per workload with quality, volume and constraints first
- Prices the invisible self-host costs: serving stack, utilisation, on-call, frontier lag
- Recommends reversibility via the gateway and names evals as the real asset
What sinks the answer
- GPU rent versus token price as the whole analysis
- Ideological answer in either direction, control or convenience
- One irreversible company-wide decision instead of a portfolio
If you genuinely do not know
Say this instead of freezing. Reasoning out loud from what you do know beats silence every single time, and a good interviewer is listening for exactly that.
“Per workload, not one bet: [frontier APIs default, since model velocity dominates], [self-host where data locality, air-gap, latency floors or massive stable volume force it], priced honestly including [serving stack, utilisation risk, on-call, frontier lag], all behind [a gateway so placement stays a quarterly routing decision]. The durable asset is [evals and integration, not GPU ops].”
Keep going with strategy
Principal
Every quarter brings a new AI paradigm the company is urged to adopt. As the senior AI voice, how do you decide what the organisation adopts, watches, or ignores?
Principal
You have 8 product teams all building AI features independently. Do you build a central AI platform team, and if so, what does it own?
Principal
The board asks: we spent 2 million dollars on AI this year, what did we get? How do you make AI investment measurable, before and after the spend?
Senior
Design the model-serving layer for a company with 15 LLM features. One gateway or per-team integrations, and what lives in it?
Senior
Support reports the AI assistant got noticeably worse this week. Nothing was deployed. Walk me through the investigation.
Senior
When do multi-agent architectures actually earn their complexity over one well-tooled agent, and what fails in them?
Knowing the answer is not the same as recalling it under pressure
Sign in to send the questions you fumble to spaced recall, so they come back right before you would forget them, and learn the concepts behind them with hands-on labs.
Start free