All resources

How to Read an AI Vendor's Benchmark Claim

Every AI vendor quotes benchmarks. Almost none test what you actually care about. Here's how to read them — and the five questions that expose cherry-picking fast.

1:44

AIbenchmarksvendor evaluationLLMs

Transcript

The vendor's deck says 94% accuracy. That number is real. It just wasn't tested on anything close to your data, your use case, or your failure modes. Benchmarks don't lie. They just tell you a very carefully selected version of the truth.

Here's what a typical vendor benchmark table looks like — and where to look first. The model name is usually the latest version. The date is often the evaluation date, not the model release date. The test set is a public academic dataset that every model has been implicitly optimized against. The score column shows the metric that made this model look best — not necessarily the one you'd choose. Accuracy here is macro-averaged across classes you don't care about. And the competitor scores? Pulled from older versions. Look for the asterisk. There's always an asterisk. The footnote says "evaluated on a subset." That subset was not random.

Five questions cut through any benchmark claim in under two minutes. First: what dataset? Academic benchmarks don't reflect messy real-world inputs. Second: what date? A 2025 evaluation of a current model means nothing. Third: whose eval? Vendor-run evals have obvious incentive problems. Fourth: does it match your use case? Summarization benchmarks don't predict extraction performance. Fifth: what are the failure modes? Any vendor that can't show you where the model breaks is hiding the part that matters most.