Who benchmarks AI agents, and can you trust it?
1 of the 84 agents here publishes a comparison against a named rival. Then one published scores. What makes a vendor's benchmark checkable, and what does not.
1 of the 84 agents in this directory publishes a comparison of itself against a named rival. In the infrastructure half of the same market, 4 of 27 do.
That gap is the background to what happened on 26 September, when one of these products published scores.
The consumer half does not compare itself
Zeme is the exception, with 8 pages against Poke, Folk, Zillow, StreetEasy, Apartments.com, ChatGPT and Claude. Every other listing here is silent about its rivals on its own site.
The clearest evidence of the norm is a product leaving it. Bloome ran twelve vendor-written comparison pages against LinkedIn Jobs, Indeed, Teal, Jobright and the rest. When it renamed itself Otty in September 2026, all twelve started answering 410 Gone, which is the code a site sends when it has deleted something deliberately. The about page and the blog went the same way. A company that had built the largest set of rival comparisons in this directory decided, at the first opportunity, not to carry them.
The builder half does it constantly
The other side of this site is infrastructure: the APIs an agent runs on. There, comparison pages are the genre. Who writes the best iMessage API comparisons exists because vendor-written pages own that search result, and four of the tools here publish their own.
The difference is who is reading. A developer choosing an API is comparing on price and deliverability and will read a vendor's table with the appropriate scepticism. A person choosing an assistant to text is not doing procurement, and the vendors seem to know it.
Then somebody published scores
On 26 September 2026 Wajo published a benchmark measuring assistants on two axes it defines: finishing an errand that goes wrong, and doing it without overstepping the user. It scored its own agent against base models and against two products listed here, OpenClaw and Hermes. Its own agent came first on both axes and had the lowest error rate.
None of those numbers is recorded as a field on any listing here, including the two it ranks. That is the standing rule about figures from somebody else's scoreboard, and it applies hardest when the scoreboard belongs to a competitor.
What separates a benchmark from an advertisement
This is not a new problem and there is a usable test for it. An audit published in August 2026 took 42 benchmark rows from nine AI vendors and asked one question of each: starting from the vendor's own page, could a buyer work out which test produced the number? It looked for four things: which version of the benchmark, which subset of its tasks, which harness ran it, and at what effort setting.
The gap that test is looking for is large. Independently run public results for one class of models topped out around 59 to 61 percent where vendor self-reports reached 80.
The failure mode is not always dishonesty, either. OpenAI stopped reporting one widely cited coding benchmark in early 2026 after auditing 138 tasks and finding more than 60 percent unsolvable as written, with evidence that models could reproduce the reference solutions from the task identifier alone. The benchmark was broken, and everybody quoting it had been quoting a broken thing.
So the question to ask of any vendor number is not whether the company is honest. It is whether you could run the test yourself and get the same answer.
On that test, Wajo did the right things
It published a paper rather than a chart. It described the method, and the method contains the detail worth keeping regardless of who ran it: the tasks happen in a simulated world of businesses with invented names, built so that no run could reach a real person. For a benchmark of agents that place phone calls, that is a choice other people should copy.
It also says the whole evaluation setup will be released: the simulated world, every task in it, and the graders that score a run.
That last promise is the whole thing. If it lands, the scores stop being a claim and become something anybody can re-run, which would be this market's first shared measure. If it does not land, the numbers stay what they are now, which is a company reporting that it won.
What to watch
Whether that evaluation setup is actually published, and whether anybody outside Wajo runs it.
Then whether a second product publishes a benchmark against named rivals. One company breaking a norm is a decision. Two is the start of a genre, and the genre is the one where everybody benchmarks themselves and everybody wins.