AI agents that do a named job
One job, in your accounts, with a person accountable for the exceptions.
An agent earns its place on work that arrives constantly, follows rules with judgment at the edges, and can be checked afterwards. It reads your systems, decides inside limits you set, and hands anything unusual to a person. Where the steps never vary, a plain workflow is cheaper and steadier.
Work it out in a minute
One task, four questions. Says which shape of AI fits, or that none of them do yet.
What we take on
One job, named before anything is built
Triage the shared inbox, chase the invoice that has passed terms, prepare the weekly report, answer from your own documents. A general assistant with no particular job is where the budget goes and the value does not.
It works in your systems, not a copy of them
Connected to the tools you already run on, through your own accounts, with read and write limited to exactly what the job needs and nothing beside it.
Limits written down before it runs
What it may decide alone, what it has to ask about, and what it never touches. Agreed in advance and visible, rather than discovered from something it did on a Sunday.
Exceptions reach a person, with the context
The value sits in the routine nine tenths. The rest arrives at somebody with what the agent saw already attached, which is faster than that person starting from nothing.
Logged, so it can be checked and fixed
Every action recorded with what it read and why it acted. That is what makes an agent reviewable by your team, and what makes a wrong answer a fixable bug.
Sometimes the answer is a workflow
Where the steps never vary and there is no judgment in them, a plain automation is cheaper to run and easier to trust. We say so when that is the answer.
What an agent is right for, and what it is not
Start from the work rather than the technology. Describe one job the way you would describe it to a new hire: what arrives, what has to happen to it, what finishing looks like, and who to ask when it is unclear. If that description is hard to write, the agent is not the first problem.
Three things make a job suit an agent. It arrives often enough to be worth automating. It follows rules most of the time with judgment at the edges. And somebody can tell afterwards whether it was done correctly. Miss the third and you have built something nobody can trust.
Where all the steps are fixed, an agent is the expensive way to run an automation. A plain workflow does the same job for less, breaks in ways that are obvious, and does not need anybody to review its reasoning.
The part that decides whether this works is the boundary. An agent that can read everything and change anything is a risk nobody signed up for. An agent with read access to two systems, write access to one field, and a limit on what it can send is a tool.
Handover matters as much here as anywhere else. The prompts, the limits, the connections and the log all live in your accounts, documented, so the agent is yours rather than something only its builder can change.
Start with one job and run it beside the person who does that job today. A month of that shows what it handles, what it hands over, and whether the second job is worth doing. It also produces the written process, which was worth having anyway.
Questions we get
What is the difference between an agent and an automation?
An automation follows steps you defined. An agent decides which steps to take, inside limits you set, and can handle input that varies. That flexibility is the point and it is also the cost.
Where the input never varies, an automation is the better buy. The decision guide on agents against automation and hiring works through it with a real example.
How do we know it did the right thing?
Because it logs what it read, what it decided and why, and because the job was defined well enough that correct and incorrect are distinguishable. Both of those are design decisions made before anything is built.
A job where nobody can say afterwards whether the answer was right is a job to leave with a person for now.
What can it get wrong, and what happens then?
It can misread an unusual case and act as though it were a normal one. The limits are what contain that: a narrow write scope, a value ceiling, a rule that anything outside a pattern goes to a person instead.
Everything it did is in the log, so a wrong answer is traceable and the rule that allowed it gets tightened.
Which model does it use?
Whichever suits the job, and it is a decision worth making on evidence rather than on brand. Some jobs want the strongest reasoning available and some run happily on something small and fast.
The AI benchmarking service exists for exactly this question, and the answer changes as the models do.
Do we need our data in one place first?
Not in one place, but reachable. The agent needs access to the systems that hold what the job touches, and it needs those records to be current enough to act on.
Where they are not, that is the first piece of work and it pays for itself whether or not an agent follows.
Who owns it when it is built?
You do. It runs in your accounts, on your keys, with the prompts, the limits and the connections written down and handed over as the work happens rather than at the end.
The page on who owns the code covers the same ground for anything else built for you.
Two places where this question comes up most: San Francisco and Seattle.
More in the guides and every answer in one place.
Shaheer leads the work, with engineers, writers, filers and analysts behind him. C-suite operations for a San Francisco AI company, Six Sigma on the process side, Anthropic certified on the Model Context Protocol, ten years across eight industries. See what we have built