Skip to content

The fact-check ledger

Every verdict is checked against primary sources after the council runs, and independently of it. This is the outcome for all seven, including the one that failed.

14 of 43 external claims survived as sourced fact. Another 15 were real but could not be traced to a primary document, 10 were cut, and 4 were false. 6 of 7 verdicts shipped.

VerdictCheckedVerifiedAttributedCutFalseOutcome
Capacity ceiling81430shipped, numbers downgraded
Inference COGS93420shipped
Free tier71132did not ship
Client concentration53110shipped
Enterprise pull62400shipped
Niche down73112shipped
Model lock-in11000shipped
Total4314151046 of 7

The four outcome columns are disjoint and sum to the claims checked. “Attributed” means the claim is real, but appears as the council’s own assumption rather than as sourced fact.

A false claim does not kill a verdict — a load-bearing one does

Free tier died because its false claims were the argument: the whole question was which channel yields more paid accounts per visitor, and the benchmarks answering it had been misread. It is not published.

Niche down also carries a false claim — an invented competitor product — but it sat in a supporting list while the verdict rested on runway arithmetic and a small sample the council flagged itself. Cut the claim, and the argument stands. So it ships, with the claim removed and recorded here.

Where the false claims came from

Two of the four were not invented by a council member. They came from a search tool’s own AI summary, which attributed statistics and client names to a real page containing neither, and were then repeated as sourced. A model with browsing will do this silently and confidently, and nothing in the output marks it.

That is the specific failure this check exists to catch, and it is why the check is done by fetching the cited page rather than by asking a model whether the claim looks right.

This table has been wrong too

Three rows were corrected on 2026-08-09 after auditing this summary against the per-verdict records. Two had double-counted claims. The third, niche down, had understated its own false-claim count as one when the record showed two.

A ledger whose purpose is to be uncomfortable is worth little if it is not audited against the primary record, including when the error flatters us.

Read the verdicts themselves — each publishes its own “What we checked” section, listing what was verified with sources and what was cut for lacking one.