AI agents need budgets, not blind autonomy
The agent trend is no longer only about better answers. Agents are being placed inside messages, commerce, and operating systems. The missing primitive is a budget for what they may do.
The AI conversation moved another step this week. Agents are showing up inside text messages, commerce flows, and operating systems. Apple is tightening macOS access controls in response to new agent risks, and the software community is asking for default hard budget caps. Those are different headlines with the same engineering lesson: an agent that can act needs a limit before it needs more autonomy.
A chatbot can be wrong and still stop at the edge of its reply. An agent can be wrong and keep calling tools. It can retry a payment, send a message to the wrong person, edit the wrong records, or burn through a model quota while its owner sees only a friendly status label. The prompt may say be careful. The system still needs to say no.
Autonomy without a budget is just an incident waiting for a better interface.
Ayush Adhikari
Step 1: Define the budget in units you can measure
Start with the things that can run away. A useful budget can combine a ten-minute deadline, twenty tool calls, one dollar of model spend, fifty changed records, and zero destructive actions. The exact numbers depend on the job. The important part is that the limit is explicit and observable.
export type AgentBudget = {
expiresAt: number;
toolCallsRemaining: number;
centsRemaining: number;
recordsWritable: number;
destructiveActions: 0;
};
export function canCallTool(
budget: AgentBudget,
estimatedCents: number,
): boolean {
return (
Date.now() < budget.expiresAt &&
budget.toolCallsRemaining > 0 &&
budget.centsRemaining >= estimatedCents
);
}
Step 2: Enforce the budget where the tool runs
Do not rely on the model to remember its own allowance. The API route, job worker, or tool server must check the budget immediately before the side effect. A prompt can explain the rule. Only the component holding the credential can enforce it.
- Pass a request-scoped budget to every tool call.
- Reserve spend before a paid model call or external write, then reconcile the actual cost.
- Make retries consume budget too. A retry is still an action.
- Return a typed limit error so the agent can stop instead of guessing.
Step 3: Put approval in front of irreversible work
Read-only search, draft generation, and a test run can usually stay inside the agent loop. Sending an email, charging a card, deleting a row, changing an IAM policy, or deploying to production should pause it. Approval is not a failure of agent design. It is the boundary between preparing an action and owning its consequences.
The approval should show the action Do not ask a person to approve “continue.” Show the tool, target, arguments, expected cost, and what cannot be undone. A human can review a concrete action. They cannot review an agent's mood.
Step 4: Record the evidence, not just the result
When an agent changes data, keep the request id, model, tool name, actor, budget before and after, approval decision, and result. This is useful for debugging a failure and for answering a less comfortable question later: why did the system think this action was allowed? Logs should make the decision reconstructable.
Step 5: Test the limit like a feature
A budget that is never exhausted is a comment with extra steps. Test the agent at zero calls, one cent over its spend, an expired deadline, a duplicate request, and a tool that returns an error after the side effect. The expected behavior should be boring: stop, report the boundary, and leave the system in a known state.
- Copy: deadlines, spend limits, tool-call caps, scoped credentials, and audit records.
- Copy: approval for destructive or externally visible actions.
- Do not copy: a giant system prompt treated as an authorization layer.
- Do not copy: unlimited retries hidden inside a helpful-looking agent loop.
The market will keep presenting autonomy as the product. The engineering work is quieter: decide what an agent may do, how much it may spend, when it must stop, and who can approve the next step. Give the agent a budget first. Then give it room to be useful.
AI agent budget questions
What does a budget mean for an AI agent?
A budget is a limit on the agent's actions, such as time, model calls, tokens, money, tool invocations, or records changed. It is an operational boundary, not just a value in the prompt.
Why are prompts not enough to control an agent?
A prompt describes intent, but it does not reliably enforce permissions or stop a runaway loop. The service that owns the tool must enforce limits and reject actions beyond them.
Which agent actions should require approval?
Require approval for irreversible or externally visible actions, including sending messages, spending money, deleting data, changing permissions, deploying code, and running production migrations.