AI agents, automation, or another hire
Three ways to absorb more work, and they are not interchangeable.
How we ship AI
- 1TestYour real tasks, run through Claude, GPT and Gemini. Scored before anything is built.
- 2BuildInside Gmail, Slack, your CRM or Sheets. Agents, MCP connections, retrieval on your documents.
- 3GuardLimits it cannot cross, a person on the exceptions, every action logged.
- 4MeasureHours and error rate before and after. If it does not pay, it comes out.
ClaudeGPTGeminiMCPAgentsRetrieval (RAG)Evalsn8n / ZapierYour CRM
Match the option to the shape of the task. Deterministic automation for high volume rule based work where a wrong answer is expensive. An agent for language and messy inputs where a review step is acceptable. A person where accountability, relationships or a decision has to be owned.
Is it a fit?
Cost of wrong
What one bad answer costs, in money or in trust.
Shape of input
Fixed fields, or language that arrives differently every time.
Accountability
Whether somebody has to own the decision in front of others.
Start here. Six questions, and the score says which of the three fits this task.
Score this taskWork it out in a minute
Six questions on one task, scored on volume, rules, stability and judgment. Returns a straight verdict.
The three options, plainly
Deterministic automation
Fixed rules.
An AI agent
Handles language, unstructured input, and cases the rules never anticipated.
Another person
Carries accountability, relationships, and judgment that has consequences.
The hybrid that usually wins
A rule based pipeline with a model doing the one messy step in the middle and a person on exceptions.
Why the wrong choice takes six months to show
The failure mode is rarely dramatic. An agent gets pointed at a task that had a perfectly good rule, and for a while it works. Then an input shifts, the output is plausible but wrong, and because it is plausible nobody catches it for a quarter.
The reverse happens just as often. A team writes a rigid automation for work that varies, then spends a year bolting exceptions onto it, until the exception handling is larger than the original rule and nobody is willing to touch it.
See alsohire an assistant, or automate the task AI agent development an AI engineer or a consultant hiring and talent
Which one fits
The choice is settled by seven signals in the work itself. Read across the row that matches your situation. Where the rows disagree, the cost of a wrong answer wins the argument.
| Signal in the work | Deterministic automation | AI agent | Another person |
|---|---|---|---|
| Input format | Structured and predictable: forms, fields, records. | Messy: email, documents, chat, voice, mixed formats. | Ambiguous, where resolving the ambiguity is the actual job. |
| Volume | High and repeating. The build pays back on repetition. | Medium to high, where writing a rule for every case is impractical. | Low volume and high value, or varied every time. |
| Cost of a wrong answer | Can be high. It will not improvise. | Should be low, or survivable with a review step before the consequence. | High and consequential, where somebody has to be answerable. |
| How it fails | Loudly. It stops, and you find out. | Quietly. It produces something plausible and wrong. | Visibly, and usually for a reason you can address. |
| Time to working | Days to weeks, once the process is written down. | Days to weeks, plus the evaluation set you have to build first. | Weeks to months including hiring, notice, and ramp. |
| What it asks of you | A process described clearly enough to encode. | Examples of good and bad output, and somewhere a person can review. | Management, context, and work that is worth their time. |
| Where it degrades | When the process changes and nobody updates the rules. | When real inputs drift away from the examples it was set up on. | When the role turns into a queue instead of a job. |
Nothing in this table is about capability. All three options can do most of these tasks. The question is which one keeps working after the person who set it up has moved on to something else.
Five questions that settle it
Run these in order. Most decisions are made by the first or the fourth, and teams that answer all five rarely pick the wrong option.
- Can you write the rule down in one sentence?If you can, and the sentence contains no "it depends", use deterministic automation.
- What happens when it gets one wrong?Follow the worst realistic error all the way to its end.
- How many times a month does it run?Below roughly ten, the build usually costs more than the task ever will.
- Is the work heavy, or is the work waiting?Waiting is a queue problem, and neither a tool nor a hire fixes it.
- Who owns the output once it runs itself?Every automated or agent run process still needs a named person who notices when it stops.
Questions we get
Are AI agents replacing rule based automation?
No. They cover a different part of the map. Rule based automation is not going anywhere for structured, high stakes work, because determinism is a feature there rather than a limitation.
How do we tell whether a task is rule based?
Write it out as if-then statements and see how far you get. If you finish and no "usually", "normally" or "it depends" survived the exercise, it is rule based.
What does an AI agent need to work reliably?
Three things: a narrow scope, a set of examples showing what good and bad output look like, and a place where a person can see what it produced.
Should we hire someone to run our automations?
Somebody has to own them, but that is rarely a full role at the start. Attach ownership to an existing person who has the time to notice failures, and revisit once the surface is large.
More in the guides and every answer in one place.
Your results so far
Kept in this browser, sent nowhere.
By Shaheer Shaikh, technology and operations consultant · Updated October 3, 2026
Read next
Shaheer leads the work, with engineers, writers, filers and analysts behind him. C-suite operations for a San Francisco AI company, Six Sigma on the process side, Anthropic certified on the Model Context Protocol, ten years across eight industries. See what we have built
