RC RANDOM CHAOS

Monitoring won't save you from a runaway AI bill

Why every AI system needs deterministic, granular hard budget caps wired into the control plane to keep LLM and agent costs predictable.

· 11 min read
Monitoring won't save you from a runaway AI bill

A single misconfigured retry loop inside an autonomous agent can turn a $40-a-day workload into a four-figure overnight bill, and the team usually finds out from the invoice, not the logs. The fix is not better monitoring. It is a hard budget cap - a limit that physically stops execution - wired into every AI system by default, at every boundary where tokens get spent: per request, per user, per agent run, per tool call, and per day. Not a soft alert that pages someone after the money is gone. A ceiling the code cannot cross.

Treat the cap as infrastructure, not a feature you add when something goes wrong. The default posture for any LLM-backed system should be fail-closed on spend: if a request would exceed its allotted token or dollar budget, it stops and returns a controlled error, the same way a well-built service rejects a payload that is too large or a connection that times out. You already do this for memory, for request size, for rate limits. Token spend is the one resource most teams leave uncapped, and it is the only one that can scale to infinity while you sleep.

The reason this matters more for AI than for a normal API is that the cost of a single call is not known in advance. A traditional endpoint costs roughly the same every time it runs. An LLM call costs whatever the model decides to generate, multiplied by however many times an agent decides to call it. You are metering a process whose output length and call count are both decided at runtime by a probabilistic system. That is exactly the situation where you need a deterministic ceiling bolted on from the outside, because the thing inside the box will not impose one on itself.

Underneath, the cost of an AI system is a product of three variables, and every one of them is live at runtime: tokens per call, number of calls, and the price of the model you routed to. Token count moves because output length is non-deterministic - the same prompt can return 200 tokens or 2,000 depending on sampling. Call count moves because agents loop, retry, branch, and fan out; a single user action can trigger one model call or forty. Model price moves because routing logic, fallbacks, and ‘use the big model if the small one fails’ patterns quietly send traffic to the most expensive tier precisely when the system is already struggling. Multiply three variable quantities together and the output is not a predictable line on a graph. It is a distribution with a very long tail, and the tail is where your budget dies.

The compounding is the part people underestimate. In a plain pipeline, cost grows linearly with usage. In an agentic system, cost grows with the depth and breadth of the execution tree. An agent that calls a tool, reads the result, reasons about it, and calls another tool is accumulating the full conversation context into every subsequent call. Token spend per step rises as the context grows, so step ten is far more expensive than step one, and a loop that fails to terminate does not cost you a flat rate per iteration - it costs you more on each pass as the history balloons. This is why a stuck agent does not produce a steady drip of charges. It produces an accelerating curve.

The spend also lives in two different places that most architectures never separate. There is the data plane - the actual model calls doing the work - and there is the control plane, the orchestration logic deciding what to call and how often. Budget caps belong in the control plane, because that is the only layer with a complete view of how much a given request, user, or job has already consumed. The individual model call cannot see the budget; it only sees its own prompt. If your enforcement lives inside the prompt or inside the model’s instructions, it is advisory at best. The model can be coaxed, confused, or simply wrong about how much it has spent. A real cap is a counter held by the orchestrator, checked before every call, that refuses to dispatch the next request once the number is hit. Deterministic control wrapped around a non-deterministic core - that is the whole pattern, applied to money.

The most common mistake is treating budget control as a monitoring problem. Teams wire up a dashboard, set a billing alert at 80 percent, and call it governance. Monitoring tells you the house is on fire. It does not close the gas valve. By the time a daily spend alert fires, the runaway agent has already run for an hour, and alerts are reactive by construction - they trigger after the threshold is crossed, which means after the money is spent. A cap is preemptive. It is checked before the call, and it blocks. Those are different categories of control, and conflating them is how organisations end up with beautiful observability and surprise invoices in the same quarter.

The second mistake is capping at the wrong altitude. Provider-level spending limits and billing caps exist, but they sit at the top of the account, aggregated across every workload, and they usually take effect with a delay measured in hours. A single account-wide limit cannot tell the difference between your revenue-critical production traffic and a test script someone left running in a loop. When it finally trips, it takes down everything, including the parts that were behaving. Caps have to be granular and local - scoped to the individual request, the specific user, the particular agent run - so that one bad actor or one broken loop fails in isolation without starving the rest of the system. A budget is not one number at the boundary of your bill. It is many small numbers enforced at the boundaries of each unit of work.

The third mistake is the belief that caps degrade quality, so they get left out of the default and added only after an incident. The opposite is true in practice. A system with no spend ceiling does not produce better answers; it produces unbounded ones, and unbounded behaviour is the enemy of a production system. A cap forces the uncomfortable but necessary design questions up front: what is a single request actually worth, how many tool calls should solving this task reasonably take, and what should happen when the work exceeds that envelope? Answering those questions makes the system more predictable, not less capable. The teams that skip caps are not protecting quality. They are deferring a decision about cost until the moment they have the least control over it - mid-incident, reading a billing page, trying to find the kill switch they never built.

The cap is a counter, and it lives in the orchestration layer, not in the model and not in the prompt. Every model call goes through a thin spend-accounting wrapper that does three things in order: estimate the worst-case cost of the call before it fires, check that cost against the remaining budget for every scope the call belongs to, and dispatch only if all of them have room. After the response returns, the wrapper reconciles the estimate against the actual token usage reported by the API and decrements the counters. No call bypasses the wrapper. If a code path can reach the model without passing through it, your cap has a hole, and the one uncapped path is exactly where the runaway will happen.

Budgets nest, and the enforcement checks them from the inside out. A single call carries a per-call ceiling - set max_tokens on every request so the most one call can cost is bounded before it runs. That call sits inside an agent run, which carries a cumulative ceiling: total tokens and total steps across the whole loop, checked before each iteration. The run sits inside a per-user or per-tenant daily budget, held in a fast shared store like Redis so it survives across requests and processes. Each layer has its own counter keyed by its own identifier, and a request is allowed only if it fits inside the tightest remaining budget of every layer it touches. One broken loop burns its own run budget and stops; it never reaches into the tenant’s daily allowance or anyone else’s.

Worst-case estimation is what makes the pre-flight check honest. Before a call you know the input token count, and you have set a max output, so the ceiling cost is input tokens times input price plus max output times output price. Refuse the call if that number would breach any remaining budget - do not gamble that the output will come back short. For agent loops, carry two hard counters alongside the token budget: a step count and a cumulative-token total. When either trips, terminate the loop deliberately and return a structured partial result - what was done, why it stopped, what remains - rather than throwing an exception that discards work already paid for. A cap that loses the half-finished job on the way out is a second failure stacked on the first.

Failing closed does not have to mean failing ugly. When a cap trips, the controlled response can route down a cheaper path instead of dying: fall back to a smaller model for the remaining steps, return a cached or partial answer, or hand off to a human queue with the context attached. The branch is chosen by your code, deterministically, not by the model’s judgement about its own spend. Keep the soft alert and the hard cap as separate numbers - page someone at eighty percent so a human can look while the system is still running, and block at one hundred so the system stops without waiting for that human. Monitoring and enforcement are both present, and neither is standing in for the other.

Take a customer-support triage agent handling inbound tickets. It reads the ticket, searches a knowledge base, pulls account history through a tool, reasons over the results, and either drafts a reply or escalates. The budget is a small config object attached to every run: fifty cents per conversation, a hard ceiling of fifteen tool calls and sixty thousand tokens per run, and two hundred dollars per day per tenant. Those numbers are not arbitrary. They come from measuring what a normal, healthy resolution actually costs and setting the ceiling a comfortable margin above it, so the cap only ever trips on genuinely abnormal behaviour.

Now the knowledge-base tool starts returning malformed results for a particular query. The agent reads the garbage, decides it is incomplete, and retries - then reasons again, retries again, each pass appending the full previous context to the next call. With no cap, this is the accelerating curve from earlier: by the time anyone notices, the loop has run forty times, the context has ballooned past a hundred and eighty thousand tokens, and the single conversation has cost more than the tenant’s entire normal day. With the per-run ceilings in place, the cumulative-token counter crosses sixty thousand at around step nine. The orchestrator does not dispatch step ten. The loop terminates, the agent returns a structured escalation - unable to resolve automatically, routed to a human, partial context attached - and writes a log line recording the scope that tripped and the exact tokens consumed.

The difference is not a cheaper invoice after the fact. It is that the broken query cost fifty cents and produced a clean handoff instead of costing three figures and producing a silent hole in the budget. The on-call engineer sees a budget-exceeded, escalated event in the logs the next morning, traces it to the malformed tool response, and fixes the actual bug - the retrieval problem - instead of first spending an hour hunting for a kill switch. The cap turned an open-ended cost event into a bounded, observable, debuggable one. That is the whole return: the failure still happened, but it happened inside a box you built, at a price you chose in advance.

An AI system without hard budget caps is not a cheaper system. It is a system carrying an unbounded liability that simply has not been billed yet. The cost variables - tokens per call, number of calls, price of the model - are all decided at runtime by something probabilistic, and probabilistic processes do not bound their own spend. The only reliable bound is a deterministic one wired around the outside: a counter in the control plane, checked before every call, scoped tightly enough that one failure fails alone.

Build the cap as default infrastructure, the same way you already treat request size limits, timeouts, and rate limits. Put the counter in the orchestrator, where it can see the full picture. Scope it at every boundary where tokens get spent - per call, per run, per user, per day. Set max_tokens on everything. Make the trip fail closed into a controlled, structured response, and keep your alert threshold separate from your blocking threshold. None of this is exotic engineering. It is the ordinary discipline you apply to every other finite resource, finally applied to the one resource most teams leave open.

The choice is not whether to decide what a request is allowed to cost. The choice is when. You can decide now, calmly, by reading your own usage data and setting a number - or you can decide mid-incident, reading a billing page, looking for a switch you never built. The teams that cap by default are not sacrificing quality for safety. They are refusing to let a probabilistic system write a cheque their budget has to cover. Design the ceiling before you need it, because the moment you need it is the moment you have the least control over it.

Share

Keep Reading

Latest on the Wire

Full wire →

New signal daily · RSS

Stay in the loop

New writing delivered when it's ready. No schedule, no spam.