Ask what they would refuse to build, who owns the prompts and accounts at the end, and what happens when the model is wrong. If they cannot name a case where AI is the wrong answer, keep looking.
Is it a fit?
Refusal
They can name a case where AI is the wrong answer.
Ownership
The model, the prompts and the accounts end up in your name.
Wrongness
They have an answer for what happens when the output is wrong.
Start here. The ownership questions sit on one page. Take them into the call.
Score the taskWork it out in a minute
Six questions on one task, scored on volume, rules, stability and judgment. Returns a straight verdict.
The plan
Do these six, in order
A time against each one and a way to tell it is finished.
Write down the one task first
Name the process, who does it now, and how many hours a month it takes.
1 hour · Done when the task fits in a sentence with an input and an output
Take the baseline before anybody sees it
Hours, error rate or queue age today, measured rather than estimated.
2 hours · Done when a number exists that somebody can check later
Send the same short brief to three firms
One page: the task, the baseline, and the question of what you should not build.
1 hour · Done when all three received an identical brief on the same day
Ask the six questions on every call
Refusals, ownership, wrong answers, which part is not AI, week-one measures, resale.
3 hours · Done when you have a written answer to all six from each firm
Compare what they refuse, not what they promise
Score on specificity: named thresholds, named failure modes, a real recommendation.
1 hour · Done when one proposal is clearly more specific than the others
Buy one deliverable, not a program
One process, one measure, one date, with ownership of code and accounts in writing.
1 week · Done when the contract names the deliverable, the date and who owns the output
What you get
Ask what they would tell you not to automate
Everybody says AI is not always the answer. Few can name a task in your business where it is not. Somebody who has done this work has refused jobs, and can tell you which ones and why in under a minute.
Ask who owns it at the end
The code, the prompts, the fine-tuning data, the vector store, the cloud accounts and the model keys. If any of those stay with them, you have not bought a system, you have rented one and the rent goes up.
Ask what happens when it is wrong
Every model is wrong sometimes. The answer you want describes a specific route: which cases go to a person, how the person knows, and what the system does while it waits. No answer here means no plan for the only certain event.
Ask which part is not AI
Most working systems are mostly ordinary software with a model doing one step. A proposal where AI does everything is either a demo or a rebuild of something a rules engine already does better and cheaper.
Ask what they measure in week one
Good answers are dull and countable: hours on a task, error rate, queue age, time to first response. If the first measurement is a launch date or a model benchmark, nobody has agreed what success is.
Ask what they resell
Plenty of firms are implementation partners for one vendor and will find that vendor is the answer. That is not disqualifying, it is disclosure. You just need to know it before the recommendation, not after.
What separates an adviser from a reseller
Almost every AI consultancy says the same six things on its home page. The words do not separate them, so use the questions above and watch which ones produce a specific answer and which produce a paragraph.
Be careful with case studies. Impressive numbers with no denominator are common: hours saved with no baseline, accuracy with no error class, adoption with no headcount. Ask what the number was before, and who measured it.
Questions we get
What should the first engagement look like?
One process, one measurement, one date. Small enough that being wrong costs a few weeks rather than a quarter, and specific enough that everybody agrees afterwards whether it worked.
If the first proposal is a transformation program, you are buying the wrong thing first.
Do they need experience in our industry?
Less than people expect. What matters is whether they have shipped something that survived contact with real users, and whether they ask good questions about how your work moves.
Industry knowledge you can supply in a week. Judgment about what breaks in production you cannot.
Is a big firm safer?
Safer on procurement, not usually on outcome. A large firm sends the people it has free, and for a company of under a few hundred staff that is often somebody learning on your budget.
Ask who does the work, by name, and how much of their week you get.
Should we ask for a fixed scope?
Ask for a fixed first deliverable, which is different. Assessments and pilots can be scoped tightly because their output is a document or one working thing.
Fixing the scope of a build before anybody has seen the data is how both sides end up unhappy.
Ask what they would refuse to build, and have your AI policy written before the first call. Worth doing before the first call: preparing for AI, so you arrive with a task and a baseline. If you are still deciding whether to hire instead, an AI engineer or a consultant.
More in the guides and every answer in one place.
Most of this is doable in house.
Read next
Why most of these stall, what to run before hiring anyone, and the cheaper answers people skip past.
Shaheer leads the work, with engineers, writers, filers and analysts behind him. C-suite operations for a San Francisco AI company, Six Sigma on the process side, Anthropic certified on the Model Context Protocol, ten years across eight industries. See what we have built