One frontier model or a router — where a small team places its bet
The hypothesis put to the council
A small AI product team should commit to a single frontier model and build deeply against it, rather than abstract behind a multi-model router.
The situation
A composite. Every number below is a working fact for the debate, and describes no real person or company.
Three engineers. B2B AI product doing structured document analysis, live 9 months. $14,000 MRR across 60 accounts, growing 11% month-over-month. Inference is 26% of revenue.
The product calls one frontier model directly and uses that provider's tool-calling, structured output and prompt caching. Caching currently saves an estimated 30% of inference cost.
The proposed router would take 3-5 weeks plus ongoing maintenance, and would require dropping provider-specific features the prompts depend on. Measured cost on their own eval set: 4-7 percentage points lower accuracy. Reliability history is one 90-minute outage in nine months, which produced three support tickets and no churn. Two customers asked about vendor lock-in in procurement questionnaires; neither made it a condition.
The verdict
Validated
Validated. The concrete penalties of the abstraction as described — 4-7 points of accuracy and the loss of a 30% caching saving — exceed the concentration risk the case file actually documents, which is one 90-minute outage, three tickets, zero churn and zero lost deals.
The vendor-lock-in questions in procurement are worth answering, but they were questions rather than conditions. Building a router to answer a question nobody made binding is expensive theatre.
Where the verdict flips
Accuracy delta on the team's own eval set between the provider-native path and the portable path, at 2 percentage points
At or above 2 points, commit to one provider. Below 2 points, the abstraction becomes worth testing.
This is the move that makes the verdict useful rather than ideological. Single-vendor versus portable is normally argued from principle, and it is settleable with a number the team already has the means to produce. If going portable costs you almost nothing in accuracy, the optionality is cheap and worth having. If it costs you five points, you are trading your product's quality for a hedge against an outage that has cost you three support tickets.
Where the council disagreed
Kept as it came out of the session. The chairman states which way the evidence points, but the split is the useful part.
Whether migration is optional or merely deferred
The strongest counter is that model retirement is scheduled, not hypothetical — so the migration cost gets paid eventually regardless. Paying it once, deliberately, while calm beats paying it reactively during a deprecation window that happens to coincide with a large deal.
Whether procurement questions become procurement conditions
Two customers asked about lock-in without making it binding. The dissenting read is that this is what the beginning of a pattern looks like, and that the answer gets harder as deal sizes grow.
Whether the accuracy penalty is a property of routers or of this router
The measured 4-7 point drop came from a lowest-common-denominator prompt. A router that keeps provider-native prompts per backend might not pay it — at the cost of maintaining several prompt sets instead of one.
The strongest case against this verdict
Stated so its advocate would accept it.
Model retirement is a scheduled event with advance notice, so the team will pay a migration cost eventually. Paying it once deliberately, via a router built while nothing is on fire, beats paying it reactively inside a deprecation window that coincides with a large deal.
That case wins if the measured accuracy delta on the team's own eval set stays under 2 percentage points, and if a signed contract is ever lost specifically because single-vendor capability was absent.
First three moves
This week
Extract every model call into one internal module that preserves native structured output, tool calling and prompt caching. Add an eval runner that takes a model identifier and records accuracy, cost per document, schema validity and latency against the existing eval set.
Next week
Run that eval against one alternative provider using that provider's own best prompt — not a translated or lowest-common-denominator version. Record the real delta and write a one-page substitution plan naming primary model, fallback, measured deltas and switchover steps.
Ongoing
Decline the router project and any contractual promise of multi-vendor failover or customer-selectable models. Put the substitution plan into procurement answers as the documented exit strategy — which is what the lock-in question was actually asking for.
The substitution plan is the real product of this decision. It costs a week, answers the procurement question in writing, and preserves the option to migrate without paying the accuracy tax up front.
How this goes wrong
Failure mode
Silent behavioural drift. An auto-upgraded model or a deprecated parameter degrades accuracy gradually, and nobody notices until customers complain — which is the one failure mode a single-provider strategy is genuinely exposed to.
Early warning signal
The internal eval set has not been run against the currently deployed model in the last 14 days. That is a calendar check, not a judgement call.
What we checked
The single factual claim in this session was checked against primary sources afterwards, and 1 survived as sourced fact. The most consequential are below; the outcome of every one is in the fact-check ledger. The reasoning above stands on its own. Most of the numbers the council reached for do not.
Verified and citable
- Model retirement is a scheduled process with published notice. OpenAI documents at least six months for generally available models, at least three months for specialized variants, and as little as two weeks for preview models — so the session's assumption of roughly 60 days was conservative for a GA model, which if anything strengthens its argument. Source
Cut — no locatable source
The council's assumptions, not established facts
- The 2-percentage-point threshold, which is the council's own derivation rather than an external benchmark. The mechanism for testing it is what matters — the number is a starting line, not a finding.
This is the shortest session in the corpus and it cites almost nothing external, because nearly every input it needed was already in the case file. That is worth noticing: the verdicts that reason from numbers the reader already has are the ones with the least to fact-check and the least to get wrong. The session also produced no evidence-audit section of its own, unlike the other six.
How this session ran
- Council
- Claude Opus 5 · GPT-5.5 · Gemini 3.1 Pro · Grok 4.3 · Kimi K2.6
- Chairman
- Grok 4.3
- Written for
- Small teams shipping AI products
- Shape
- Five independent proposals, anonymized peer review, chairman synthesis. One round.
Also in the ledger
A Council Audit on your own decision
Every external claim is checked against primary sources before you read it. What survives is cited. What doesn't is printed as cut, the same way it is above.
You get one recommendation and the number at which it flips, the strongest argument against it, and the first three moves in order.
$299. Written brief in 48 hours. No call.