Inference is a real cost of goods — stay flat or go usage-based
The hypothesis put to the council
A B2B AI product whose inference cost runs a third of revenue should move to usage-based pricing rather than defend a flat subscription.
The situation
A composite. Every number below is a working fact for the debate, and describes no real person or company.
Two people, B2B AI product, live eleven months. $18,000 MRR across 140 accounts on a single flat $129/mo plan with unlimited use inside a fair-use policy that has never once been enforced. Gross margin 68%, and inference is essentially the entire cost of goods.
The cost is badly skewed. The top 12% of accounts consume 61% of inference spend. Nine accounts are gross-margin negative. The median account costs about $18/mo to serve.
Growth is 9% month-over-month, logo churn 4.1%/mo, and the cost to serve an active account is up 40% in six months because the accounts that stay use the product more. The team has exactly one pricing change in them this year. Their stated fear: metering makes people ration, and rationing kills the habit that drives retention.
The verdict
Invalidated
The 32% COGS ratio is not the load-bearing trigger. The nine gross-margin-negative accounts — costing 2.1–2.6× what they pay — are addressable by enforcing the fair-use policy that already exists and has never been used. Doing that lifts aggregate gross margin to roughly 81% while touching the other 131 accounts not at all.
A pricing-model change is a large, one-shot, trust-expensive move. The case file describes a problem that a policy already on the books can solve.
Where the verdict flips
Same-cohort six-month cost-to-serve growth, at 25%
Above 25%, go usage-based. At or below, stay flat and enforce the ceiling first.
This is the sharpest thing the council produced, and it is worth stating plainly: the observed 40% rise may not be real. At 4.1% monthly logo churn the population changes underneath the average. If cheap accounts churn out faster than expensive ones, blended cost-per-account rises while no individual customer changes behaviour at all. That is survivorship, not escalation, and it argues for doing nothing.
The case file even says the accounts that stay use the product more — the signature of a compositional effect rather than a behavioural one. Measuring the same accounts in month 5 and month 11 separates the two. It is one SQL query, and it decides a pricing strategy.
Where the council disagreed
Kept as it came out of the session. The chairman states which way the evidence points, but the split is the useful part.
Whether the cost rise is behavioural or compositional
Two proposals treated the blended 40% as decisive evidence for metering. One, backed by peer review, argued it may be survivorship at 4.1% monthly churn. The chairman sided with the latter, and the whole verdict rests on that reasoning.
Whether hybrid pricing carries a measurable uplift
Two proposals cited a specific percentage-point advantage for hybrid pricing. Peer review found the delta small, not AI-specific, and sourced only to secondary summaries. It did not survive the session, let alone the fact-check afterwards.
Whether new tiers or circuit breakers are the middle path
Two proposals wanted them. The objection that carried: both still require metering infrastructure and carry the same trust risk as overages, while enforcing the existing policy requires neither.
The strongest case against this verdict
Stated so its advocate would accept it.
The 40% six-month rise continues, dragging blended gross margin from 68% to negative inside 24 months even with the median account unchanged. And the top cohort's revealed willingness to consume is the only visible path to expansion revenue in a business shedding roughly 40% of its logos a year.
That case wins if the nine negative accounts are multi-seat organizations whose usage tracks their own growth — expandable — rather than single power users running loops, who are simply fixable by enforcement. And it wins if same-cohort growth is at or above 25%.
First three moves
This week
Run the same-cohort query on accounts active in both month 5 and month 11 and measure their individual cost change. One query, and it decides the strategy.
Weeks 2–3
Build shadow metering — an events table plus daily aggregation — and enforce the existing fair-use ceiling on the nine negative accounts with founder calls and 30 days' notice.
By end of month
Publish one quantified ceiling on the existing $129 plan for new accounts only, with graceful model downgrade at the limit. Tell the 140 existing accounts their usage is now visible and the policy is now enforced.
Explicitly declined: overage billing, multiple tiers, and grandfathering beyond 90 days.
How this goes wrong
Failure mode
The unenforced ceiling — you publish a limit and then never hold it, which is exactly how the company arrived here with a fair-use policy nobody has ever applied.
Early warning signal
The first account granted an exception above the published limit at the base price. Not the tenth. The first.
What we checked
All 9 factual claims in this session were checked against primary sources afterwards, and 3 survived as sourced fact. The most consequential are below; the outcome of every one is in the fact-check ledger. The reasoning above stands on its own. Most of the numbers the council reached for do not.
Verified and citable
- Cursor's June–July 2025 pricing backlash, refunds and public apology. Their own post: “We recognize that we didn't handle this pricing rollout well, and we're sorry.” Source
- Heavy-tailed resource usage is documented in peer-reviewed work, not folk knowledge — 7,170 sample distributions across three storage clusters over 478 days found tails heavier than log-normal. Source
- Stripe Billing meters bill at 0.7% of billing volume. Source
Cut — no locatable source
- A specific hybrid-pricing revenue uplift figure attributed to Freemius. Not present anywhere in their published material. The council's own peer review had already flagged it as unverifiable.
- “Metering causes a 30–40% drop in prompt volume.” No controlled data exists for this in any domain. It is the founders' fear stated as a statistic.
The council's assumptions, not established facts
- Inference prices falling ~50× per year, attributed to Epoch AI. Epoch's published finding is a range of 9×–900× per year depending on the capability milestone, with 40×/year as their worked example. The direction is right; the single number is not theirs.
- That Stripe's docs require guardrails to live in the product layer. Stripe never says this outright, but their meter architecture ingests events asynchronously with no real-time authorize hook, so the conclusion follows from the mechanics.
- Net revenue retention benchmarks from Bessemer's 2021 report. Real and correctly quoted, but published over a year before ChatGPT — a pre-AI SaaS cohort used as an AI-product baseline.
Three of the claims that failed this check had already been flagged by the council's own evidence audit during the session, which recorded that no verified source supported the 32% COGS threshold as a decision rule at all — the premise it was handed.
How this session ran
- Council
- Claude Opus 5 · GPT-5.5 · Gemini 3.1 Pro · Grok 4.3 · Kimi K2.6
- Chairman
- Grok 4.3
- Written for
- Solo founders and small teams shipping AI products
- Shape
- Five independent proposals, anonymized peer review, chairman synthesis. One round.
Also in the ledger
A Council Audit on your own decision
Every external claim is checked against primary sources before you read it. What survives is cited. What doesn't is printed as cut, the same way it is above.
You get one recommendation and the number at which it flips, the strongest argument against it, and the first three moves in order.
$299. Written brief in 48 hours. No call.