---
title: LLM Gateway at $1.2M: when the flat-rate math broke
date: 2026-09-18T09:00:00.000Z
edited: 2026-09-18T09:00:00.000Z
author: Ismail Ghallou (Smakosh)
canonical: https://smakosh.com/blog/llm-gateway-flat-rate-economics
reading_time: 8 min (1819 words)
Tags: ai, growth
---

# LLM Gateway at $1.2M: when the flat-rate math broke

![The LLM Gateway models directory](/assets/blog/llmgateway-models.png)

Seven weeks ago I published [a post about DevPass](/devpass-growth-branding) that contained this sentence:

> Those three prices and that 3× multiplier have been constant since day one — every "pricing change" in the product's history was packaging and framing, never the number.

On October 15 the multiplier drops from 3× to 2×. The prices stay the same.

This is the post about why. It's less fun than the last two, and it's the one I'd have wanted to read a year ago: what actually happens to a flat-rate subscription sitting on top of metered AI infrastructure, and what we built in the meantime so the business stops depending on getting that one number right.

## Where things stand

[LLM Gateway](https://llmgateway.io) crossed **$1.2M in gross all-time revenue** this week — roughly 16 months from the first commit. Across product lines, that breaks down as:

| Line                           | Gross all-time |
| ------------------------------ | -------------- |
| Credits (pay-as-you-go)        | $918K          |
| DevPass plans & top-ups        | $231K          |
| Enterprise licensing           | $66K           |
| Reset Passes, chat plans, misc | $5K            |

Behind it: **22,019 sign-ups**, 19,022 verified, **3,460 paying customers**. In my [first LLM Gateway post](/building-llm-gateway) in June, we'd crossed 8,800 signups and 714 paying.

The traffic side grew faster still. That post reported 34.4M requests and 214B tokens. Today the gateway has served **77.7M requests** and **1.53 trillion tokens** across **1,558 distinct models** — and 1.1T of those input tokens were cache reads, which turns out to be the whole story.

## Why flat-rate AI pricing is hard

DevPass sells a multiplier: pay $29, $79 or $179 a month, get that much times three in model usage at provider catalogue prices. It's a genuinely good offer, and it's the reason the last two quarters happened.

It also contains a piece of arithmetic I never took seriously enough. **A 3× multiplier breaks even at 33.3% utilization.** If a Lite subscriber pays $29 and burns more than $29 of the $87 we granted them, that subscriber costs us money. Not at 100% of the allowance — at one third of it. Every flat-rate plan on metered infrastructure has a number like this, and it is the only number that decides whether the plan works.

Two things moved ours, in roughly the same eight weeks.

**Agentic coding consumption stopped looking like chat consumption.** A developer using Claude Code or [DevPass Code](https://devpass.llmgateway.io) all day isn't sending a few thousand tokens per turn. They're sending a large, mostly-cached context on every step of a loop, for hours, unattended. That's what 1.1T cache-read tokens out of 1.53T looks like from the inside: individually cheap, enormous in aggregate, and growing with tool quality rather than with headcount.

**We grew into the part of the curve where the heavy users live.** At two hundred subscribers, an average hides everything. At several thousand it stops hiding anything. The subscribers who sit closest to their cap aren't outliers to be smoothed over — they're the customers the product is _for_, and they behaved exactly as the offer invited them to.

That second point is the one I'd hand to anyone shipping a plan like this. The risk was never the average subscriber. It was that I kept looking at the average.

## What we did about it

**Sold the overflow instead of the cap.** [Pay-as-you-go overflow](https://llmgateway.io/changelog) shipped on August 5: exhaust your weekly premium allowance and you can opt into billing the difference against your credits balance rather than stopping dead. Reset Passes do the same job one tier up. Both turn a hard wall into a transaction, and both are opt-in — we still refuse to sell a Reset Pass late in a cycle when it can't pay for itself.

**Stopped subsidising abuse.** A flat-rate plan with a 3× multiplier is an arbitrage invitation, and people accepted it. Over the last six weeks we shipped duplicate-card fingerprinting, abusive-IP risk flagging at signup, country-level signup blocks, a new-account credit-purchase kill switch, per-organization concurrency and spend caps, and a tiered content filter. None of it appears in a changelog anyone celebrates. It closed a real part of the gap.

**Changed the multiplier, with 30 days' notice.** From October 15, DevPass grants 2× the plan price instead of 3×. Prices don't move. Existing subscribers transition at their next renewal rather than immediately; new subscriptions and upgrades get the new allowance right away. The notice went up on the signup page, the plan chooser, the pricing page and a dated section of the terms — [shipped a month before it takes effect](https://github.com/theopenco/llmgateway/pull/4080), specifically so nobody buys a plan this week that quietly becomes a different plan next month.

Two-times moves break-even utilization from 33.3% to 50%. I want to be careful about how much credit I take for that, because it assumes usage behaviour stays put, and it won't. Shrinking an allowance doesn't shrink demand — it moves where the cap bites, and it bites first on exactly the heavy users most likely to leave over it. The economics get better. Some of the revenue in the numerator goes with them. I'd rather say that now than discover it in November.

## The part of the business that just works

While I was busy with the subscription, the boring line compounded. Credits — plain pay-as-you-go, buy a balance, spend it at catalogue prices plus a platform fee — is now the large majority of gross revenue, and it has no utilization assumption to get wrong. What the customer spends and what we pay move in the same direction, by construction.

That's the real lesson of the quarter. DevPass sells certainty to the customer and buys volatility from the provider. That's a legitimate trade and people pay well for it — but we'd priced it on eight weeks of data, in a market that reprices itself monthly.

So we spent the quarter building revenue that doesn't depend on winning that trade every month.

**Enterprise.** $66K in licensing so far, on the back of the deeply unglamorous work enterprises actually buy: teams with SCIM and Entra directory sync, zero-data-retention enforcement that fails closed across provider routing, project-level admin roles, license gating and seat warnings, hash-only API key storage (keys are HMAC-SHA-256 fingerprints now, and the plaintext columns are dropped), and expanded provider disclosures in the legal pages.

**[Airside](https://airside.llmgateway.io).** This is the structural one. Provider onboarding used to be sales calls and a spreadsheet, which quietly favoured whoever had a BD team. Airside is a self-serve carrier console: a provider verifies their own domain, claims their listing, runs a preflight verification against their live endpoint, files prices for review, and tunes the discount and landing fee that win them routed traffic. $2,500 one-time to list, waived by invite for providers we already work with, plus a landing fee on the traffic they win.

It makes the gateway a two-sided marketplace — margin can come from carriers competing for routed requests, not only from marking tokens up to developers. It launched three weeks ago with its first carriers. Ask me again in six months.

## On publishing verified revenue

Both previous posts pointed at [TrustMRR](https://trustmrr.com), which read our Stripe account directly and published the number for anyone to check. We've since expired that key and stopped sharing revenue there.

I want to be straight about that, because I used it as a credibility prop twice and it would be easy to just quietly drop the link. Early on it was the best asset we had: two people with no funding and no audience, claiming numbers nobody had a reason to believe. A live feed from Stripe solved that in a way no screenshot could.

It stopped paying for itself. The audience that matters now asks for SOC 2 Type II, zero-data-retention guarantees and a signed DPA — none of which an MRR badge speaks to — while a continuously updating revenue feed is genuinely useful to competitors pricing against us. The figures in this post come from our own Stripe reporting, and they're no longer third-party verified. That's a real downgrade in how much you should take my word for, and it's the right call anyway.

## Distribution kept compounding

The one thing that behaved exactly as promised. In June, organic search was 33.4K clicks from 2.33M impressions. Today it's **73.7K clicks from 5.05M impressions** at an average position of 8.8 — both more than doubled in three months, off pages we wrote up to a year ago. The open-source core went from 1,293 to **1,644 stars** and 186 forks over the same stretch.

None of that required a decision this quarter. That is the entire point of it.

## The numbers, 16 months in

- **$1.22M** gross all-time revenue across all products
- **22,019 sign-ups**, 3,460 paying customers
- **77.7M requests**, **1.53T tokens**, 1,558 distinct models routed
- **73.7K clicks / 5.05M impressions** from organic search, average position 8.8
- **1,644 GitHub stars**, 186 forks on the open-source core
- Enterprise licensing live; Airside live with its first carriers

Not everything landed. Lounge, our chat product — same team, same brand system, same distribution surface as DevPass — hasn't found anything like the same traction. Developers wanted a subscription for the tool they already lived in. They did not want another chat window.

## What I took away

**Every flat-rate plan has one number, and it isn't the price.** It's break-even utilization, it's set by your multiplier, and it decides everything. I could have written it on a napkin in January. I worked it out properly in September.

**Averages hide the only thing that matters.** A headline utilization figure tells you about the median subscriber, who was never the risk. The distribution tells you about the business. I should have been looking at the shape of it from the first hundred customers, not the mean.

**Instrument cost at the same time as revenue.** We had beautiful instrumentation on signups, starts, cancellations, cap hits and funnels — every growth number — months before we had the equivalent on the cost side. So for a while the dashboards said one thing and the bank account said another. Build both halves, even when the second one is boring and the number is small.

**Announce the bad change early and in full.** The last post argued that honesty is a growth tactic rather than a constraint. This was the expensive version of that test: 30 days' notice, prices held flat, existing subscribers riding to renewal, and a warning on the signup page that costs us conversions today. I'd rather pay that than spend down the thing the last two years were actually built on.

If you're running a flat-rate plan on metered infrastructure and want to compare notes on where your break-even actually sits — [reach out](https://smakosh.com/contact). I have a spreadsheet now.

---

_Originally published at https://smakosh.com/blog/llm-gateway-flat-rate-economics_
