Six models. Your data. One table you can decide from.
Claude, GPT, Gemini, Grok, Copilot and Perplexity, tested on your questions, scored the same way.
In short
You bring the decision and the data. We put the same questions to Claude, GPT, Gemini, Grok, Copilot and Perplexity, score every answer the same way, and hand you a heatmap, a consensus table and one recommendation. Where they agree, you can act. Where they split, you know what to check. Every answer and score stays with you.
Work it out in a minute
One task, four questions. Says which shape of AI fits, or that none of them do yet.

What the table looks like
Claude vs GPT vs Gemini vs Grok vs Copilot vs Perplexity, on one sheet
Sample output. Scores out of 100 on a client question set. Yours will differ, which is the point.
| Model | Accuracy on your data | Cites its sources | Same answer twice | Reads your files | Cost per 1,000 | Verdict |
|---|---|---|---|---|---|---|
| Claude | 92 | 84 | 95 | 96 | 71 | Shortlist |
| GPT | 90 | 78 | 86 | 91 | 74 | Shortlist |
| Gemini | 87 | 83 | 79 | 88 | 88 | Shortlist |
| Grok | 79 | 64 | 68 | 70 | 82 | Not for this |
| Copilot | 83 | 76 | 77 | 90 | 73 | Not for this |
| Perplexity | 80 | 97 | 66 | 58 | 85 | Sources only |
What we take on
Your questions, not a leaderboard
Public benchmarks test trivia. You are tested on your contracts, your tickets, your spreadsheets, and the questions your customers ask.
Six models, one scoring sheet
Claude, GPT, Gemini, Grok, Copilot and Perplexity answer the same set, blind, and are scored on the same criteria. Accuracy, sources, consistency, cost, speed.
A heatmap you can read in a minute
Every model against every criterion, weaker to stronger, on one page. The pattern is visible before anyone reads a number.
A consensus table
Where all six agree, marked. Where they split, marked, with the answers side by side so a person can settle it.
One recommendation, with the reason
Which model, for which task, at what cost, and what would change the answer. Written for the person who signs, not the person who codes.
Repeatable
The same set can be run again when a model updates or a price changes, and the new heatmap sits beside the old one.
Yours to keep
Every prompt, every answer, every score and the scoring sheet itself. Show it to a board, an auditor or the next vendor.
What you have at the end
AI benchmarking, done for a decision, means putting the same questions from your own work to Claude, GPT, Gemini, Grok, Copilot and Perplexity, scoring every answer on the same sheet, and reading the result as a heatmap and a consensus table. That is what you have at the end.
A heatmap. Six models down the side, your criteria across the top, every cell colored by how that model did on your questions. Weaker to stronger, at a glance, before anyone reads a number.
A consensus table. For each question, whether the six agree or split, and where they split, the answers side by side. Agreement is where you can act. Disagreement is where the risk is, and now it has a name.
A cost line. What each model would cost at your volume, next to what it got right, so the cheapest and the best are on the same page and the trade is plain.
One recommendation. Which model for which task, in writing, with the reason and the thing that would change the answer. Short enough for a board pack.
The evidence. Every prompt, every answer, every score, and the sheet the scores came from. Yours. Run it again next quarter and put the two heatmaps side by side.
And a decision that took a week instead of a quarter, made on your data instead of a slide from a vendor.
The plan
Do these seven, in order
A time against each one and a way to tell it is finished. The sections after this say why each step sits where it does.
Name the decision
What is being chosen, for which task, and what a wrong choice costs. One page.
half a day · Done when the decision fits in one sentence and everyone agrees on it
Build the question set
Fifty to two hundred real questions from your own material, with the answer a good person would give.
2 days · Done when a colleague could score an answer without asking you
Agree the scoring sheet
The criteria, the weights, the cost at your volume. Fixed before the first run.
half a day · Done when you have signed off the sheet
Run all six, blind
The same questions to every model, answers stored, names hidden from the scorers.
1 day · Done when every answer is on file with its model hidden
Score and chart
The heatmap, the consensus table, the cost line.
1 day · Done when the pattern is visible without reading a number
Decide, in writing
Which model for which task, the reason, and what would change the answer.
half a day · Done when one page a board can read
Keep the set
The questions, the sheet and the results stay with you, ready to run again.
1 hour · Done when you could re-run it without us
Questions we get
Which models do you test?
Claude, GPT, Gemini, Grok, Copilot and Perplexity by default, the current version of each. Any model with an interface can be added, including open models you host yourself.
If you already have a shortlist, we test the shortlist.
What gets measured?
Whatever the decision depends on. The usual set is accuracy on your material, whether the answer cites a source you can check, whether the same question gets the same answer twice, how it handles your files, speed, and cost at your volume.
You see the criteria and the weights before anything runs.
Does my data leave my control?
Your questions and documents are used for the test and nothing else. Where a model offers a no-training setting, it is on. Sensitive material can be masked or replaced with matched samples before it goes anywhere.
You get a record of what was sent where.
How is this different from the public benchmarks?
Public benchmarks score models on public tests, which the models have often seen. They say nothing about your contracts, your tickets or your tone.
This scores the models on your work, blind, with the answers kept.
How long does it take?
A focused decision, one task and six models, is about a week: two days to build the question set and the scoring sheet with you, two days to run and score, one day to write it up.
A wider comparison across several tasks runs two to three weeks.
What do I do with the result?
Sign the contract, pick the model under the product, or set the policy for which tool staff use for what. The recommendation is written to be acted on.
When a model updates, the same set runs again and you see whether the answer moved.
More in the guides and every answer in one place.
Shaheer leads the work, with engineers, writers, filers and analysts behind him. C-suite operations for a San Francisco AI company, Six Sigma on the process side, Anthropic certified on the Model Context Protocol, ten years across eight industries. See what we have built