Economics
How unmetered stays in business.
Every few months another “unlimited” AI plan gets capped, throttled or repriced. This is why that keeps happening, and why a plan priced in lanes doesn’t have to.
Sold per seat, paid for per token.
An inference provider’s real cost is GPUs: bought or rented, powered and cooled, by the hour. Tokens are what comes out of them. A GPU that idles for an hour costs the same as one that runs flat out.
Seat-based unlimited plans charge per person per month. But what a person costs to serve depends on how many tokens they pull through those GPUs, and among developers using agents that number varies by orders of magnitude. Someone who asks a few questions a day and someone running three agents overnight pay the same price and cost wildly different amounts.
That holds while heavy users are rare. Coding agents made them common. The math breaks, and the provider reaches for the only levers left: caps, rolling windows, slower queues, smaller models. The word unlimited stays on the page while the limits multiply underneath it.
Price the thing that costs money.
A GPU serving a model works on many requests at once, in a batch. Each place in that batch is a slot of real capacity: a known share of a known machine for a known length of time.
A lane is one of those slots, reserved for your team. Because it is a fixed share of fixed hardware, the most a lane can ever produce in a month is set by the hardware, not by how hard you push it. We know that ceiling before we sell the lane, and we price the lane above it.
The same logic sets the rest of a plan. A larger model needs more of the machine per request, so it holds more lanes. A long context holds GPU memory for as long as the request runs, so the largest windows come with the largest plan. And a request streams faster when it shares the GPU with fewer others, so higher plans run at higher speeds. Each is a cost we can count in advance; none of them is a count of your tokens.
What that buys you.
- Your heaviest day costs what your lightest day costs.
- A lane run flat out for a month costs us the same as one that sits idle, so it costs you the same too.
- No limits to discover.
- The only limit is printed on the pricing page: the number of lanes. There is no hidden daily budget behind it.
- Nothing to claw back.
- Heavy use of a lane doesn’t raise what the lane costs to run, so there is no reason to slow you down, shorten your context or swap the model under you.
The honest limit.
Lanes are a real constraint. Run more agents at once than you have lanes and some of them wait for a turn. We think that is the right trade: a limit you can see and plan around, instead of one you discover halfway through a Thursday afternoon.
If you are waiting often, add lanes. The price of a lane is the whole price; nothing you do inside it changes the bill.