AI
RevenueBench: The Benchmark RevOps Has Been Missing
Every public leaderboard tests what a model can do in general. Nobody had tested what it can do with your actual job until now.
Cliff Simon
September 30, 2026
Public model benchmarks are almost useless for the work I actually do. They'll tell you a model can code, write, or solve a math problem, and none of that tells me whether it can clean up a CRM, catch a renewal date that contradicts itself across two systems, or build a forecast off an export nobody bothered to standardize. That's the work. That's what I need to know a model can do before I put it anywhere near my pipeline.
Amani Phipps, Sr. Revenue Architect at Bonusly, built the benchmark that actually answers that question. It's the closest thing I've seen to hard evidence for an argument that I've made before. AI earns its keep in specific places, on specific work, and the only way to know whether it works is to test it against the job itself.
The report is called RevenueBench. Amani ran 119 models against 40 real revenue-operations tasks, built from the actual jobs a GTM team runs every week like forecasts, hygiene audits, renewal calls, the monthly close. Every answer is graded by code against a computed truth, not by another model's opinion, with traps planted throughout like blank owner columns, renewal dates that disagree with each other, call records tied to deal IDs that no longer exist.
A few findings from the report are worth knowing before you dig in yourself.
- The top open-source score across all 40 tasks came in at 0.985, and the full leaderboard, every prompt, every model's raw answer, cost, and latency, is sitting in the open for anyone to check.
- The top-scoring model on the entire leaderboard cannot read a screenshot, because it's a text-only model, which means the “best” model on paper is disqualified the moment your workflow involves an image.
- The field is sorted into six behavioral archetypes rather than a single ranking. The Complete Analysts score elite across the board with no weak flank, Reliable Operators are cheap and steady on routine work but thinner on complex joins, and Capable but Loose models score well until they start inventing names and account details that were never in the data. Which one matters to you depends entirely on what you're asking it to do.
This report will be rebuilt monthly instead of published once and left to rot, and it's graded against a real answer key instead of another model's judgment.
If you're trying to figure out where open models make sense for your team, this is where I'd start.
More signals
Subscribe