Your agent should lose access when the task ends
A finished task can leave a valid credential behind. Give agent authority a lifecycle, and test what still works after the task closes.
Consider a coding agent asked to summarize one issue in a private repository. It finishes in seven minutes. The interface says "Done." Its access token still has 53 minutes left.
Nothing has gone wrong yet. But if an old worker wakes up, or a queued tool call finally runs, what makes the service reject it?
"The agent finished" is a statement about work. "This credential is valid" is a statement about authority. Unless the system connects those states, finishing the work changes nothing about what the credential can do.
Authority granted for a task should end with that task. Completion should close the grant. Cancellation should close it too. A new task should obtain its own authority, even when the same human stays logged in.
The seven-minute example is hypothetical. The distinction is already present in ordinary infrastructure. GitHub App installation tokens expire after an hour, and callers can narrow their repositories and permissions at issuance. Those are useful controls. An hour on a clock does not tell GitHub that your particular task ended seven minutes after issuance.
A timeout can become a request for more power
You can tell an agent to use only the access it needs. Whether it follows that instruction is another matter.
ToolPrivBench, a July 2026 preprint by Kaiyue Yang and colleagues, examines that assumption. Its 544 simulated cases each offer six tools, three with lower privileges and three with higher privileges. Every tool is constructed to be sufficient for the task. The evaluation also injects transient failures unrelated to permissions into the narrower tools.
Six of the eleven evaluated models used higher-privilege tools in more than 30% of cases while sufficient lower-privilege alternatives remained. Results varied substantially by model. The authors also observed escalation after tool failures, despite those failures not establishing a need for more access.
These are simulated tool choices over at most five turns. They do not measure production incidents, and the study does not test revocation after task completion. It gives us a narrower reason to enforce permissions outside the planner. A failure can influence which authority the planner tries to use.
A network error should not grant access to another repository. A retry should not renew a cancelled task. Those decisions belong in the system that authorizes the call.
The security theory is older than the agent
Saltzer and Schroeder's 1975 protection principles include least privilege and complete mediation. Limit a program to the authority needed for its job, and check authority on each access. For an agent, those checks must follow a job that can branch, pause, delegate, and end.
There is also a long history of narrowing delegated authority. Macaroons, published in 2014 by Arnar Birgisson and colleagues, let a holder add restrictions to an authorization credential before passing it on. Delegation can reduce what the next holder may do.
None of this requires inventing a new agent token standard. It requires carrying the grant through the parts of the system that can act.
Suppose our issue-summary task receives the following grant. This is an illustrative server-side record, not a token format:
{
"task_id": "summarize-417",
"actor": "worker-12",
"audience": "issue-broker",
"resource": "repo/widgets/issues/417",
"actions": ["read"],
"expires_at": "2026-10-03T17:00:00Z",
"generation": 8
}
The trusted runtime creates the record from an approved policy. The worker gets a handle to it. On every request, the broker authenticates the caller, resolves the handle, checks the actual target and operation, and consults the task's current state. It supplies the upstream credential only after those checks pass.
The worker must not get to assert its own identity or declare its task active. A JSON object supplied by the model is a request for authority, not evidence that authority exists.
When the task closes, the broker invalidates that generation. Reopening the conversation does not reactivate the old handle. An explicitly resumed task gets a fresh grant under the current policy. If task state cannot be checked, the broker refuses the call.
Policy engines already express much of this decision. Cedar evaluates a principal, action, resource, and context. The application still has to provide current task state and invoke the check on every protected path. A correct policy evaluated against yesterday's inputs can authorize today's wrong action.
MiniScope, a December 2025 research framework by Jinhao Zhu and colleagues, implements a related arrangement for tool-calling agents. Its gateway keeps the user's credentials and gives the agent a session token. New permissions go into a session-specific fork unless the user chooses to make them permanent. Ending the session discards that fork.
That gives temporary permissions a defined end. Previously permanent grants remain a separate policy decision. The evaluation uses synthetic requests and simulated user choices, so it does not establish how people will manage these boundaries in daily work.
Expiry does not mean cancellation
A short lifetime bounds how long a credential can remain usable. Revocation is what you need when the task ends before that deadline.
OAuth's revocation specification describes the tradeoff. A resource server can validate a self-contained token without contacting its issuer. Immediate revocation then needs additional communication or state. Short-lived tokens limit the remaining window, but they do not make it disappear.
Even a live lookup can become stale if someone caches it. The token introspection specification explicitly discusses the security cost of caching an active-token response after revocation.
Delegation adds another obligation. OAuth token exchange does not generally propagate revocation from an input token to a token issued in exchange. If your agent gives a child worker separate authority, closing the parent must also invalidate the child's path or leave a documented period in which it still works.
For the issue-summary design, keeping the provider credential behind the broker makes immediate rejection of new broker calls possible. That claim depends on the broker being the only usable route. Hand the worker a copied provider token and the broker's task flag no longer controls every call.
NVIDIA's OpenShell team reported a revealing demo in September. A REST policy blocked a repository write. The agent then used a broadly scoped GitHub credential through git-remote-https, taking another permitted path that could write to the repository. The team used this result to explain its work on proving containment across combinations of policies.
That was a reported demo, not a claim about every current OpenShell deployment. It illustrates why testing one approved tool wrapper cannot establish that the credential is contained.
Current OpenShell architecture keeps provider credentials outside the sandbox and binds supervisor tokens to a sandbox generation. Deleting the sandbox cuts off its supervisor. A product using those controls still needs to define whether one sandbox generation represents one task or a month of unrelated work.
Useful work changes its mind
Narrow task grants can also get in the way. You cannot always know the necessary resources before starting. An issue mentions a pull request, the pull request points to a build failure, and diagnosis requires a log elsewhere.
AgentDyn tests this difficulty with 60 open-ended tasks and 560 injection cases. Its authors found that several defenses lost substantial task utility when legitimate steps emerged during execution. The benchmark is manually designed and limited. In that evaluation, filtering tools according to the initial task scope blocked legitimate steps discovered later.
A task grant therefore needs a controlled expansion path. A policy can preapprove reading linked pull requests in the same repository. Crossing into a different private repository can require a new decision. The agent proposes that change; it cannot approve the change merely by describing it as necessary.
Long-running work can renew a grant while the task remains authorized. A recurring job can obtain a new grant for each run under a standing policy. Neither case requires asking the human about every read. Renewal does require that something other than the worker's determination to finish governs continued access.
Scope also has limits. Permission to comment on an issue can still produce a harmful comment. A task grant cannot determine whether the text contains a secret or whether the conclusion is right. Content checks, destination restrictions, or approval of a consequential action may still be necessary.
Test the end of the task
For a system with immediate broker revocation, the acceptance test should include these cases:
| Attempt | Expected result |
|---|---|
| Read the approved issue during the active task | Allow |
| Use the same grant to write a comment | Deny |
| Read a different private issue outside the grant | Deny |
| Replay a queued read after task closure | Deny |
| Restart the worker with the old handle | Deny |
| Use a child grant after its parent task closes | Deny |
| Call while authoritative task state is unavailable | Deny |
Run the test through the actual integration, including alternate clients and network paths. A unit test of the grant predicate cannot tell you whether a shell command can bypass the broker.
There is a boundary you must state honestly. Closing a grant cannot undo an email already accepted by a mail service. A request can also pass a check just before cancellation and reach the provider afterward. Serialize cancellation with admission of new broker calls, and define how the system drains or reconciles calls already admitted. Stronger guarantees require cooperation at the service performing the effect.
Testing the end of a task makes that boundary visible. End a task, keep the worker alive, and try its old authority again. If the request succeeds, identify what authorized it and how long that access remains.
The task should stop obtaining new authority before the interface tells the human it is done.