Try Budgets Beat Agent Sprawl: A Containment Card for Governed Execution
Give each agent cell a finite try budget, a dated stop, and an evaluator who cannot spend the ledger it grades.
The dashboard looks busy. Packets went out. Drafts piled up. Seats got scanned. Three new agents shipped this week.
Someone asks what the run proved. The room answers with volume.
That is how agent work turns into sprawl. Activity scores as fitness. The cell keeps spending because the scoreboard still lights up.
Governed execution starts with a smaller unit. Each agent cell gets a finite try budget, a dated stop, and an evaluator who can kill the run without spending the same ledger it grades.
Sprawl is what happens when activity scores as fitness
Okta defines agent sprawl as agents proliferating without centralized tracking, inventory, or governance. Teams stand up chatbots and workflow agents on their own. Nobody can say what exists, who owns it, or what data it can reach. Updated 19 March 2026, that page still treats sprawl as an inventory and identity problem first.
The inventory gap is real. There is a second gap that shows up after the registry is filled in.
The company can now name the agents. It still scores them on motion.
Packets sent become proof. Drafts produced become throughput. Seats scanned become coverage. Token spend becomes “the agents are working.” Leadership gets a weekly count and calls it an operating review.
Fitness is a verified change in the job’s outcome. A covering experiment looks for cash-shaped proof, or another commitment the founder has already named as binding. A service workflow looks for a verified change in the job: a cleaner handoff, a send that held, a production change that stayed correct. A good stop counts too, if it saved a bad week. Packets, drafts, and seats scanned are inputs to that judgment. They are not the judgment.
When activity is the scoreboard, every new agent looks healthy. The estate grows because growth is how the estate proves it exists.
Registry and ownership are necessary — and still not containment
The 2026 advice from identity vendors, CIO advisors, and public frameworks is worth taking. It is also incomplete if you stop where they stop.
Okta’s prevention list is the right first layer:
- an enterprise-wide registry;
- agents as first-class identities;
- lifecycle management from creation to retirement;
- least-privilege and time-bound access;
- a cross-functional governance forum.
If leadership cannot answer how many agents are deployed, Okta’s signs-of-sprawl checklist already applies. Fill that gap. Do not skip it.
BCG’s 2026 CIO guide then asks for an Enterprise AI Control Plane above every platform. The plane holds common identity, an agent and tool registry, and runtime policy enforcement. It also offers “golden paths” so the compliant build is the fast one. The registry they describe should capture ownership, configuration, tool schema, and versioning. It should also track token budget and usage, with spending limits and circuit breakers before cost becomes an incident. That is a serious control-plane claim. Token and usage caps still measure compute and API spend. They do not, by themselves, decide how many irreversible actions a cell may take, when the run ends, or who may kill it.
Singapore’s Model AI Governance Framework for Agentic AI v1.5 (published 20 May 2026, updated 5 June 2026) bounds risk in a different language. Organisations should judge an agent by the scope of actions it can take, the reversibility of those actions, and its autonomy. Humans stay accountable. “Human-in-the-loop” has to be adapted for automation bias. Name the checkpoints that need approval, especially high-stakes or irreversible actions. Then audit whether that oversight still works. Permissions plus a human checkpoint are necessary. A cell can still keep researching and acting if research and irreversible action share one ledger. The same leak appears when the person who grades the run can also spend it.
Arcade’s 11 September 2026 lifecycle framework is the closest public map to a full life. Register, scope, deploy and update, operate, account, and decommission. A governed agent is owned, permission-bound, version-pinned, observable, cost-attributable, and provably retired. Cost that cannot be traced to an agent cannot be traced to an owner. Retirement means the agent can no longer act: credentials revoked, triggers disabled, evidence retained. That is the right end of the life. Cost-to-agent is still activity accounting unless fitness is defined on a different scoreboard.
Take the registry. Take the owner. Take identity, runtime policy, irreversible-action checkpoints, cost attribution, and retirement proof. Then add the missing unit: a finite number of Act tries, a date the run dies, and an evaluator who cannot spend the Act ledger.
The Agent Operating Record already connects a registry entry to a bounded job, source-of-truth rules, approval points, and a lifecycle decision. The card below is the containment layer that record is missing.
The missing unit: tries + dated stop
A try is one Act against the world: a send, a spend, a publish, or a production change the cell cannot cheaply undo. Research is not a try. A draft is not a try. A weekly status note is not a try.
Give the cell a number. Give the cell a date. The run ends when either one is hit.
That pair does three jobs a registry cannot do:
- It makes autonomy expensive in the right unit. The cell spends tries on irreversible action, not on existing.
- It makes the stop inspectable. “We will know when it is working” is not a stop. A Friday and a remaining-try count are a stop.
- It gives the founder or owner a kill clock that does not wait for a security incident.
In the public AI Nature Company journey, the covering experiment moved onto this unit. The closed-beta constitution was amended from a cash ledger to unit tries. The kill clock is tries plus a dated stop. One owner is accountable for the outcome. The founder still signs outbound. Packets sent are not fitness. The stance is experiment-until-fitness: the cell has to show cash-shaped proof, or it stays small.
A service business can use the same unit without copying that experiment. Name one job the company already does. Cap how many irreversible actions the agent may take this window. Write the date the window ends. Decide now what evidence would justify another window.
Recurring autonomy without those two fields is a standing invitation to keep the lights on.
Separate Sense from Act on the ledger
If research and irreversible action share one budget, the cell learns the wrong lesson.
Spend the budget on research and the send never happens. The review then says the cell did nothing. Spend the budget on sends and the research was never paid for. The next packet is a guess wearing a blazer. Mix them and nobody can say whether the list was wrong or the try was wrong.
Sense observes. It builds lists, enriches them, and assembles packets. On the Nature Company covering cell, Sense research does not debit try spend. Act outbound does, on complete or fail, and only from Sense-ready packets. The founder still signs the send.
Learning stays split by force:
- Bound adapts who the job is for.
- Sense updates whether a list, an enrich, or a door was true.
- Act updates whether a try was effective.
That split is the operating rule, even if your stack has different names.
Give Sense a research ceiling if the team will otherwise drown in drafts. A seats-examined cap, a packet-ready cap, or an hours-of-research cap is fine. Keep that ceiling off the Act ledger. Draft volume is a research limit, not a fitness score.
Then put Act on a try budget. A try is spent when the irreversible action completes or fails. An unsent draft does not spend a try. A send that bounced still spent one, because the world was touched.
The send gate still sits between the packet and the world. The ledger says how many times that gate may open this window. The gate says who may open it, on which channel, with what evidence.
Independent Defend: evaluators who cannot spend
The person who produced the run should not clear the run.
That is older than agents. It gets ignored anyway, because the builder is nearby, the evaluator is busy, and shared praise feels like closeout.
Defend is independent verification. On the covering cell, an evaluator effectuates Defend. It can reject. It can raise stop-risk or real falsification. It cannot spend Act tries. It cannot become the owner of the thing it just praised. A producer cannot clear its own claims.
IMDA’s accountability chapter is useful here if you keep it operational. Humans remain accountable. Multiple actors in the agent lifecycle can diffuse that accountability. The useful move is to name who may approve an irreversible action, who may grade the window, and who may kill the cell. Then check whether the human at the checkpoint is still actually checking. Automation bias and alert fatigue are how a “human in the loop” becomes a stamp.
If the evaluator can also spend Act tries, the grade and the spend collapse into one incentive. The cell will keep finding reasons to spend. If the owner who sets ICP also impersonates the evaluator, the cell will keep finding reasons the ICP was right.
Write the separation on the card before the first try. After the first good packet, everyone will argue they can wear two hats for a week.
Fill the Try-Budget Containment Card
Copy the card. Fill it for one cell, one job, one window. Leave blanks visible. The blanks are the containment gaps.
| Field | What to write | |---|---| | Bounded job | One job the cell is allowed to do this window, in workflow language. | | Owner (Bound) | The human who sets ICP, targets, and the membrane. One owner per outcome. | | Sense work | Research the cell may do. No Act-try debit, or a separate research ceiling. | | Act try budget | How many irreversible actions this window. Count complete and fail. | | Dated stop | Calendar end, or tries-exhausted, whichever comes first. | | Fitness definition | The verified outcome that would justify another window. Packets, drafts, and seats scanned do not qualify. | | Evaluator (Defend) | Who can reject or kill the run, and confirmation they cannot spend Act tries. | | Founder/owner kill conditions | The patterns that end the cell now, before the date. | | Send/approval gate | Who may turn a packet into a send, spend, publish, or production change. | | Retirement proof | How you will show the cell can no longer Act when the window ends. |
Decision rule:
No recurring autonomy without a try budget, a dated stop, and an evaluator who cannot spend the Act ledger. If the only scoreboard is activity, you are financing sprawl.
Worked example, generalized from the Nature Company covering cell
The public covering experiment is fractional sit-the-seat work on posted full-time roles. It is a closed beta on personal DataSaa infrastructure, not a commercial offer. Use it as a filled card, then rewrite the fields for the job you actually run.
| Field | Covering-cell example | |---|---| | Bounded job | One covering experiment: put a fractional operator on posted full-time roles, through official application doors, inside a dated window. | | Owner (Bound) | Founder/owner sets ICP and targets. One owner per outcome. Owner does not impersonate selection. | | Sense work | Lists, enrich, and packets. Sense research does not debit Act tries. A seats-examined ceiling is allowed. Draft volume is not the goal. | | Act try budget | A finite unit-try budget. Outbound spends a try on complete or fail, and only from a Sense-ready packet. | | Dated stop | A calendar date or tries-exhausted, whichever comes first. | | Fitness definition | Experiment-until-fitness: cash-shaped proof or another founder-accepted commitment. Packets sent are activity. | | Evaluator (Defend) | Independent evaluator. Speaks on stop-risk or falsification. Cannot spend Act tries. Cannot take ownership of a run it just praised. | | Founder/owner kill conditions | A chatbot org chart. A marketplace GTM bot farm. Treating packets-sent as success. An ICP the covering bet cannot prove. | | Send/approval gate | Founder signs every outbound. Official doors only. | | Retirement proof | Tries gone or date hit; remaining Act paths disabled; evidence kept; cell marked experiment-closed rather than quietly still sending. |
The same card works for a service-business job. Bound might be the practice lead who names the client slice. Sense might be a research agent that may not debit sends. Act might be an operator agent with twelve tries through 3 October. Defend might be a partner who can kill the window and cannot hit send.
If a field is “we will decide later,” the cell is already financing sprawl. Later is how standing agents get born.
What to kill early
Kill these before they grow a constituency.
A chatbot org chart. Eight named agents in a chat, each with a job title, looks like a company. It is a swarm with opinions. The Nature Company founder already killed that pattern. Forces can stay as physics. Hosts can be temporary. Standing personas recreate the sprawl the registry was supposed to prevent.
A marketplace bot farm. A catalog of add-on bots that “do GTM” or “do ops” multiplies Act surfaces. The farm usually has no Bound owner, no Sense/Act split, and no dated stop. Volume becomes the product.
Volume goals. Packets per week, seats scanned, drafts completed, agents deployed, tokens burned. Those numbers can sit on a research ceiling or a cost report. They cannot sit on the fitness line. Transform metabolizes Sense and Act into evidence or economics. Activity stays off the fitness line.
Resurrecting consumed work. A packet that already spent a try, failed, and taught the door was wrong should not be replayed to dress the scoreboard. Consumed work belongs in the evidence ledger. Re-sending it to keep the cell “active” is how a dated stop gets cheated.
If you need a weekly number, count remaining tries, days to stop, evaluator rejects, and whether any fitness evidence appeared. That is a containment report. It is a short meeting.
Where this sits next to send gates and weekly reviews
This card does not replace the artifacts already on this site. It sits beside them.
The send gate names the approver, channel, evidence threshold, and outcome log for one artifact. The send-gate playbook walks the stage transition from draft-ready to sent. Use those when a packet is about to leave the building. The containment card says how many times that gate may open this window, and when the window dies.
The weekly operating review asks which workflows ran, who owns them, what drifted, and what to keep, fix, or stop. Bring the card into that review. Remaining tries and days-to-stop should be visible without opening a vendor dashboard.
The Agent Ownership Scorecard asks who owns the outcome, the decision rights, the risk boundary, and the retirement rule. Fill the scorecard so the cell has a human contract. Fill this card so the contract has a budget and a kill clock.
Direction Before Speed still holds: write the workflow, owner, risk boundary, and stop-or-scale metric before adding capability. The try budget is that stop-or-scale metric in a unit the cell can actually spend.
For the wider spine, read From AI Sprawl to an Operating System. Scorecards, owners, and cadence are the management layer. A try-budget card is how one cell stays contained while that layer is being built.
If you want the public experiment these fields were generalized from, the dated dispatches live on the Nature Company journey log.
Optional next step
Pick the agent cell that already runs without a kill clock. Fill the ten fields this week. Do it on one page. Do not add another agent until the blanks are closed.
If the only number you can put on fitness is a packet count, stop the cell. Write a better fitness line, or admit the run is research-only and take it off the Act ledger.
A registry will tell you the cell exists. A send gate will tell you who may touch the world. A weekly review will tell you whether anyone looked. The containment card tells you how much Act remains and when the run dies. It also names who can kill the run without spending the tries they are supposed to grade.
If you already have one expensive job a service business does every week, one job is the quieter door. Ask for one inspectable run with memory and a gate. The card still works if you never send that email.
Sources
- Okta, “What is Agent Sprawl?”, updated 19 March 2026.
- BCG, “How CIOs Can Govern AI Agents at Scale in 2026” (Enterprise AI Control Plane), 14 August 2026.
- IMDA, Model AI Governance Framework for Agentic AI v1.5, published 20 May 2026, updated 5 June 2026. PDF: framework document.
- Arcade, “What is AI agent governance? A lifecycle framework for governing production AI agents”, 11 September 2026.