If your product calls a language model every time a user clicks, flat-rate pricing is a bet that your heaviest users never find you. In 2025, two well-funded companies lost that bet publicly within weeks of each other, and their write-ups are the best free pricing education a SaaS founder can get right now. The conclusion up front: when a feature has real marginal cost, your price needs a usage component or an explicit, published cap, decided while you are calm, not retrofitted mid-controversy. This week's business press is full of the same story playing out again at consumer AI startups; the pattern is structural, not a management failure at any one company.
Two companies, one diagnosis
Cursor sold its Pro plan as 500 fast requests a month, then in June 2025 rebuilt it as "$20 of frontier model usage per month at API pricing," with extra usage purchasable at cost. Its own diagnosis: newer models create "an order of magnitude" of cost variance between a simple request and a hard one, so counting requests had become fiction, one user's 500 requests could cost fifty times another's. The messy part wasn't the economics but the communication: "unlimited" turned out to apply only to the auto-routed mode, what the company called rate limits was really "a usage credit pool for the month," and Cursor ended up refunding surprised users' overages from the transition weeks.
Replit ran the same movie with different scenery. Its agent charged a flat $0.25 per checkpoint until the company replaced it with effort-based pricing: simple requests now cost less than the old quarter, complex agent runs more. The stated reason mirrors Cursor's, flat pricing "assumed consistent task complexity," and agent tasks had diverged into quick fixes and long autonomous runs that cost wildly different amounts to serve.
Different products, identical structure: both priced like software with zero marginal cost, both watched task-level costs spread by orders of magnitude, both repriced toward usage, and the company that suffered most reputationally suffered for the rollout, not the economics.
The arithmetic that catches up with you
The underlying anchor is public. Frontier API rates run around $5 per million input and $25 per million output tokens at the Opus tier (Anthropic's published pricing, checked July 24, 2026), which means a single long agentic session, chewing through millions of tokens of context and output, can cost real dollars. Cursor's post gives the scale from the other side: its $20 allowance corresponds to roughly 225 Sonnet-class requests.
Now run our illustrative arithmetic on a $20/month flat plan with 100 users. Suppose the median user consumes $3 of inference, a middle band $15, and the top five users $60 each, a mild version of the skew both companies describe. Revenue: $2,000. Inference cost: roughly $700. On average, comfortable. But averages are exactly the wrong lens: each power user loses you $40 a month, and power users are not a fixed 5%, they are the people for whom your product works best, they concentrate over time, and a flat price is an open invitation for the $200-of-usage customer to arrive. Meanwhile every pricing page that says "unlimited" is marketing to them specifically. That is the quiet asymmetry: with zero-marginal-cost software, your best users were your best customers; with per-click COGS, your most enthusiastic users can be your worst customers, and growth makes it worse.
Four structures that survive contact with power users
- Flat with a published fair-use cap. Keeps the simple pitch; the cap converts tail risk into a known ceiling. Works when cost variance is moderate. The cap must be on the pricing page from day one, a cap discovered by users mid-month is how you end up writing an apology post.
- Included credits plus overage at cost, where Cursor landed. The subscription buys a bundle priced off your actual COGS; heavy use pays its way. Best default for prosumer and developer tools whose buyers already understand metering.
- Pure usage pricing, Replit's effort-based direction. Cleanest margin protection, hardest sell for buyers who need predictable bills; suits products where usage maps visibly to value delivered.
- Hybrid seat + meter. A per-seat price covering the workflow value, plus a meter on the expensive verbs only (agent runs, generations, exports). Fits team SaaS where finance wants a predictable base and you need protection on the tail.
The choice reduces to two questions. How wide is your per-user cost variance, if p95 is under about 3x the median, flat with a cap is fine; if it is 20x, you need a meter somewhere. And how much billing unpredictability will your buyer tolerate, developers tolerate meters, consumers mostly don't, which pushes consumer AI products toward caps and credits rather than raw usage.
Three operating rules before you pick
- Instrument per-user COGS from the first week. Both cautionary tales are stories of companies discovering their distribution late. You cannot design a cap, a credit bundle, or a meter without knowing your median, your p95, and how fast the tail is growing.
- Set the flat component so the 95th-percentile user is profitable, not the median. Pricing to the median is how the top decile silently becomes a subsidy program. This is the usage-cost sibling of a lesson we've argued before: founders anchor low and never revisit.
- When you reprice, over-communicate. Cursor's refunds were a communication bill, not a pricing bill: the new structure was defensible, the words "unlimited" and "rate limits" were not. Publish the arithmetic, define every term, grandfather generously, and announce before the invoice does it for you.
One more note for the revenue side of the same spreadsheet: metered and overage billing raises the number of charges you attempt, which raises your exposure to failed payments, build recovery in before the meter ships. And if you are on the buying side of these repricings rather than the selling side, the defense is the same as ever: know your own per-use cost before your vendor's new meter teaches it to you.
Flat pricing didn't break because founders got greedy or careless. It broke because a generation of pricing intuition was trained on software that cost nothing to run, and that assumption quietly expired the day your product started buying tokens on every click. Price like you sell compute wrapped in workflow, because now you do.
If pricing AI features has you rethinking the whole price list, this pairs well with setting a price you can defend and auditing what AI tooling actually costs you.
Discussion
Sign in with Google or just a name. No email link, no password to remember.