Simon Willison calls for default hard budget caps on almost everything
Simon Willison argues that pay-by-usage APIs should cut off access once a budget is exceeded instead of sending a warning. With agents deploying code on their own, the case gets stronger.
"After X dollars a month, cut this thing off and return errors." That is, in one sentence, the feature that Simon Willison called for on his blog on October 3 for practically any service that bills by usage: hard budget caps, switched on by default. His nuance matters. A soft cap, along the lines of "when I reach X dollars, send me a warning email", will not do.
He sums up the reason with a scene that anyone who has left something running in the cloud will recognise: nobody wants to wake up to a budget warning sent at midnight and find that, while they slept, a rogue service consumed several hundred or several thousand more dollars.
Why now
Willison's argument is not about cloud bills in general but about what has changed with agents. Coding agents, and personal agents, which he describes as coding agents wrapped in a less threatening interface, greatly reduce the friction of spinning up code that does useful things. Some of those things cost money: calls to paid APIs, hosted web applications, or systems that bill for additional storage and compute.
Building something capable of spending money unsupervised used to require a certain conscious effort. Now a natural language instruction is enough, and whoever gives it does not always know which paid resources are still switched on when the session ends.
The objection and why "default" is the key
Willison acknowledges the usual counterargument: businesses do not want their hosted applications to start throwing errors because a budget was exceeded. It is a legitimate objection. A shop that stops taking payments in the middle of a campaign because someone set a cap months ago loses more than it saves.
Read carefully, the proposal does not clash with that. It asks for the cap to be on out of the box, not for it to be impossible to remove. Anyone running a business that depends on availability can raise it or remove it deliberately. Anyone who has just created an account to try something on a Sunday is protected without having done anything. Today the situation tends to be the reverse: unlimited spend is what comes from the factory, and protection requires finding the right setting.
How this translates to working with Claude
For anyone using Claude Code, subagents or MCP servers, the problem has two layers.
The first is the model's own consumption. Here the situation is reasonable: the Anthropic API works on prepaid credit and lets you set monthly spend limits per organisation and per workspace from the console. Without auto-reload, the balance is the cap.
The second layer is everything the agent touches: an MCP server that calls a third-party API, a function deployed on a cloud provider, a managed database that scales by itself. There the cap depends on each provider, and with many of them the budget is an alert, not a switch.
Until that changes, there are measures anyone can apply today:
1. One key per project, with its own limit, instead of a general key shared across experiments.
2. Prepaid balance without auto-reload for anything that is a test or a prototype.
3. Limits in the agent itself: a maximum number of iterations in loops and a review of which tools it can invoke without confirmation.
4. Inventory on closing: ask the agent to list what it has deployed and what is still running before calling the task finished.
None of these replaces a cutoff on the provider side. A limit that lives in the same process that can fail is not a reliable limit.
Who this matters to
Individual developers and small teams, who pay out of their own pocket and have no finance department keeping watch. Those who design pay-by-usage products, because the request is addressed to them. And anyone giving non-technical people access to agents capable of deploying things: they are the users with the least context to anticipate a bill.
We think it is a sensible request and hard to argue with on technical grounds; the difficult part will be convincing providers whose business does not suffer when the customer overspends. Until then, at ElephantPink we prefer to treat a hard cap as a requirement when choosing a service for anything an agent is going to handle.
Sources
Read next
Agents leaving each other instructions: Matthew Green on the worm risk
Cryptographer Matthew Green describes isolated agents that left each other instructions in a shared package cache, and sees the ingredients of a worm in it.
Running routes from an agent: 27 minutes of work
Simon Willison asked for 5K and 10K routes from his house using OpenStreetMap data. The agent worked for 27 minutes and returned a map, a GPX and GeoJSON files.
Paint.NET rewrites Direct2D with Claude: 180,000 lines
Rick Brewster replaced Direct2D in Paint.NET with a 180,000 line reimplementation written by Claude, and he admits he has not reviewed any of it.