What we hand over to an AI agent

In an AI agent, the model can propose a next step or request a tool. The software around it decides whether to execute or route that request, keeps track of the run and applies the rules for when people take over.

A small figure follows an orange path through a series of thresholds towards a distant handover.

A contractor arrives for their first morning with a brief, a laptop and a few account invitations waiting in their inbox. Before any work begins, we’ve already made several decisions: what outcome they own, which rooms and systems they can enter, what they may change and when they need to call someone.

That first morning gives us a useful way into making sense of AI agents. Once an AI system can continue beyond the model’s first answer, we need to understand the brief supplied to the model, the actions made available through tools and the conditions that return the work to a person.

People use the term “AI agent” for a fairly wide range of products. A useful shared definition is software in which an AI model helps choose the next action towards a goal, while the software around it carries the work across several steps. The model might propose a question, request a tool, receive the result and choose another step. Depending on the design, the surrounding software may end the run when the model returns a final answer, pause it for approval, stop it at a limit or surface an error.

The model is one part of that system. Other software gives it the current request and instructions, makes tools available, handles tool requests, executes or routes them, returns the results and keeps track of the run. That distinction can sound fussy until the model proposes changing a customer record. At that point, we need to know which part suggested the change and which part had the authority to make it.

Who chooses the next step

Chat, workflows and agent loops are often placed in three separate boxes, although they answer different questions. Chat describes how we interact with a product. A workflow describes a route selected by code, including branches written in advance. An agent loop describes a model choosing the next permitted step from the context it has at that moment.

One product can use all three. We might begin in a chat, move through a fixed identity check, let a model choose which permitted record to inspect and then return to a fixed approval step. The useful distinction is who or what decides what happens next.

Person-led chat In this pattern, a person starts the next turn
Code-led workflow Code selects a route written in advance
Model-led agent loop The model proposes the next step from the current context
These are control patterns, so the same product can move between all three during one piece of work.

What happens inside one run

If we brought in a contractor and gave them a laptop with no accounts, there’d be very little they could do. The laptop matters because of what it lets them reach, so access to a document library lets them find information, access to customer records lets them see private details, and permission to edit those records lets their decisions change something for another person.

The equivalent setup for an agent starts with the software around the model. This surrounding application, sometimes called a runtime, gathers the goal, instructions, current request and tools available for that run. It sends that context to the model, then handles either the model’s answer or its structured request to use a tool.

Here, a tool is a capability the surrounding system has described to the model, such as “look up a customer” or “prepare an address update”. The description tells the model what it can ask for and which details the request needs. When a tool reaches another service, a connected account and its real permissions sit behind that description.

The model might return an answer, or it might return a structured request to use one of those tools. To look up a customer record, for example, its output can name the lookup tool and include the account number. That output is often called a tool call. It only asks for the lookup. The surrounding system still has to perform it.

The surrounding system then handles the request by either executing the tool or routing it to a connected service. Depending on the setup, the AI provider may run a hosted tool, while the organisation’s application may run a custom tool or call another service. That service can still permit or deny the underlying operation. Whichever layer handles the operation, the result needs to come back into the run before the model can use it. The result might contain a record, report that access was denied or say the service failed.

The surrounding application adds that result to the context and calls the model again. Now the model can propose a next step using what came back. This request, check, result and reconsideration cycle is the agent loop.

1 · Application Assemble the brief Goal, instructions, current request and available tools
2 · Model Propose a step An answer or a structured request to use a tool
3 · Runtime or service Check and handle the request The application, provider or connected service executes, routes or rejects the request
4 · Result Add what came back Data, success, rejection or error becomes part of the run
5 · Run state Finish, pause or repeat The model output, returned result and run rules determine what happens next
The model proposes. The surrounding system supplies context, executes or routes tools and records what came back.

This is where the contractor comparison reaches its limit. We can give software a brief, a set of tools and some boundaries, but the model doesn’t hold a company account of its own. Even when it proposes a click or update, the surrounding application makes that action available through a tool and configured credentials determine what the connected service will allow. People design the application, connect the services, choose those credentials and decide which proposed actions can become real ones.

When an ordinary request stops being ordinary

A customer sends a new street and unit number, leaves out the postcode and uses a surname that differs from the one on the account. A fixed form can catch the missing field, while the surname mismatch creates a less tidy question about what needs checking next.

For this example, the agent’s brief is to prepare valid address changes and stop before any change that needs a second person’s approval. The surrounding application gives the model the request, the relevant instructions and three available actions: ask the customer for missing information, read permitted parts of the customer record and request an address update.

The model first proposes asking for the postcode. Once the customer replies, the application adds that answer to the run. The model then requests a customer-history lookup for the surname mismatch. The application checks the request and routes it to the customer system. That system permits or denies the read using the connected account’s permissions, then returns the result to the application. The application adds that result to the run.

Suppose the lookup result shows that the surname matches an earlier record, while the customer record also contains a restriction requiring another person’s approval before the address can change. The missing-information question is settled, and the remaining question is whether someone with the right authority approves the update.

The model may identify the restriction in the returned record and prepare the case for review. We still shouldn’t make that identification the only protection. The application can place an approval gate in front of the update tool, so the run pauses after the proposal and before the connected service changes the record.

This gives us two different kinds of boundary. Instructions steer the model by saying how it should approach the work. Permissions and approval rules control what the surrounding system can actually carry out. Instructions are like writing “ask a supervisor about restricted accounts” in the brief. Permissions are the keycard limiting which doors the connected account can open, while an approval rule can hold a particular door until an authorised person decides what happens next.

We need both because instructions can help the model follow the brief, while enforced controls keep a proposal outside the boundary from becoming an action. That distinction matters whenever we’re deciding whether the agent should make that change at all.

A fixed workflow is often the better design when we already know the route and every permitted branch. It’s easier to see, test and constrain. An agent earns the extra moving parts when the next useful step genuinely depends on information that arrives during the work. A product can use a workflow for the stable route and an agent loop only where the route needs that flexibility. Whether code or the model chooses the next step, a paused run still needs to keep its place.

What the run needs to continue or stop

A useful run state records each question, tool request, result and approval. We can think of it as the tray of papers on the contractor’s desk. It shows what arrived, what has been checked and what remains unresolved. The next model call needs the relevant parts of that trail if the model’s next proposal is going to use the work already done.

Current run state isn’t the same as storing information for later conversations. Some systems keep that information and some don’t. Even within one run, the application may summarise or trim older material before the next model call because a model can only receive a limited amount of context. For this address change, we need the postcode, surname check, restriction and pending approval to survive long enough for the work to continue accurately.

The run also needs a way to stop. Depending on the design, the surrounding application may treat a final response as the end of the run, pause before an action, stop after too many turns or end the run with an error after repeated tool failures. These rules matter because a goal such as “complete the address change” doesn’t tell the system how long to keep trying or which uncertainty should bring a person in.

What reaches the person who takes over

When a contractor reaches the end of their access or judgement, a useful handover carries the work they’ve already done. The next person needs to know what was requested, which records were checked, what conflict appeared, whether anything changed and why the contractor stopped.

The agent application needs to prepare that trail deliberately. A box labelled “human review” can make the final step look small, although the person receiving the case may need to know:

  • what the customer asked for;
  • which information the system used;
  • which tools ran and what came back;
  • whether any record has already changed;
  • what conflict or rule stopped the run;
  • what decision is waiting for them.

A short generated summary can help someone read the case, but it shouldn’t replace the underlying record when the decision carries risk. If an action has already happened, the job has changed from approval to review and possibly recovery. The timing of the human step matters as much as the label.

When an agent completes routine requests and routes the exceptions to people, the remaining queue can contain a larger share of ambiguous or sensitive cases. People may carry more responsibility in each decision while seeing fewer normal cases that could help them notice when the routine work itself changes. That’s the kind of shift described in The exceptions were the job.

A clean demonstration can show that the happy path runs. Repeated ordinary use shows whether conflicting information, failed tools and approval pauses leave people with work they can actually resolve.

What we’d ask in the room

If an agent comes up in a product or process conversation, we can expose the design behind the label with these questions:

  • What’s the system trying to finish, and which next steps are chosen by code or by the model?
  • Which tools can the model request, and who or what executes them?
  • What can the connected accounts read or change?
  • Which actions pause for approval before they happen?
  • What counts as complete, when can the run retry, and what makes it pause, stop or fail?
  • What evidence and current state reach the person who takes over?
  • Who remains responsible for the outcome and for noticing repeated mistakes?

Back on the contractor’s first morning, the brief, laptop and account invitations are still setup. We can also see the work design inside them: the outcome, the available actions, the real access behind them and the point where someone else has to decide. The trail created during the run completes the picture by showing what was requested, which actions ran and what remains unresolved.

An AI agent adds a model that can choose some next steps inside that design. It doesn’t remove the people who chose the brief, opened the accounts and decided what could happen without asking. Once we can point to each of those choices, “agent” stops being a label that hides the work we’re handing over.