What Building Atlas Changed About How I Think About Quotas
Lessons from building and locally validating a quota-aware LLM gateway: account for work, make reservations explicit, and explain every rejection.
Table of Contents
I started Atlas because “requests per minute” felt like the wrong abstraction for shared LLM capacity.
Two requests can have completely different costs. One may produce a short classification; another may hold a streaming connection while generating thousands of tokens. If both consume one request from the same quota, the accounting is simple and the resource protection is weak.
Atlas is my local, production-oriented reference implementation of a quota-aware LLM gateway. It has not operated under sustained production traffic, so the useful story is not a latency claim. It is what building and testing the control path taught me.
Quotas should account for work
The gateway needs a unit that is closer to resource consumption than request count. Tokens are an improvement, but even token counts are model- and phase-dependent: prompt processing and token generation do not have identical costs, and different models have different serving profiles.
For a practical gateway, I would start with a weighted usage unit:
usage = input_tokens * input_weight
+ output_tokens * output_weight
+ fixed_request_cost
The weights do not need to be a perfect GPU cost model. They need to be stable enough to enforce a product policy and explainable enough that users understand why capacity was consumed.
Streaming requires a reservation
Non-streaming requests can be charged after the response is complete. Streaming creates a harder problem: the gateway does not know the final output size when it admits the request.
Atlas treats this as a reservation problem:
- estimate or cap the maximum charge;
- reserve that amount atomically;
- stream the response while recording actual usage;
- settle the reservation when the stream completes;
- reclaim it safely if the client disconnects or the worker fails.
This is where idempotency stops being an abstract distributed-systems rule. Without a stable request identity, a retry can reserve twice. Without expiry and settlement, interrupted streams leak quota. Without atomic updates, concurrent requests overspend the same balance.
Admission and routing are different decisions
I originally wanted the gateway to make one “smart routing” decision. In practice, it is clearer to separate two questions:
- Is this tenant allowed to start this work?
- Which healthy backend should serve it?
Admission depends on policy, quota, and the estimated cost of the request. Routing depends on backend health, model availability, and load. Combining them makes it harder to explain why a request was rejected and harder to retry safely when only the selected backend failed.
The separation also leaves room for different policies. A tenant may have quota but no compatible backend. A backend may be healthy but reserved for a different service tier. Those are different failures and should remain different in the API.
A rejection should be actionable
“429 Too Many Requests” is technically valid and operationally thin.
For quota enforcement to feel fair, the response should identify the limiting dimension, current usage, applicable limit, and retry condition. Operators also need to trace that decision back to the policy version and accounting events that produced it.
This changed the metrics I cared about. Request count and latency were not enough. I wanted reservation age, unsettled usage, rejection reason, quota utilization by model and tenant, and disagreement between estimated and actual cost.
That last metric is especially important. If estimates are consistently high, capacity sits unused. If they are consistently low, admitted work exceeds the intended budget.
Local validation has a boundary
The repository currently has 46 automated tests covering quotas, streaming reservations, forecasting, health, and usage behavior. It also includes load-test and streaming harnesses.
That gives me confidence in state transitions and failure cases exercised by the tests. It does not establish production SLOs, multi-region consistency, or behavior under months of real tenant traffic. I prefer to state that boundary directly rather than turn a design target into an outcome claim.
What I would keep in a production version
Building Atlas left me with four rules I would carry into a real deployment:
- charge for an understandable approximation of work, not just request count;
- reserve before admitting open-ended work such as streaming;
- separate admission policy from backend routing;
- make every rejection and accounting adjustment explainable.
Quota systems are often described as protection from abusive users. I now think that framing is too narrow. A good quota system is the contract that lets multiple users share expensive, variable capacity without having to trust one another’s traffic patterns.