When a Benchmark Stops Measuring
Eye for AI · 4 September 2026
Every good test of machine ability eventually becomes a bad one.
Volume III · Number 6 · September 2026
Independent and reader funded. No advertising, no sponsored posts, no affiliate links. editor@eyeforai.blog
Contents/Notes/Reading a Capability Announcement
Note
A short field guide to the genre: what to look for, what to discount, and the two sentences that carry all the information.
Abstract
Capability announcements follow a fixed form, and almost none of the information sits in the part most readers reach. This note sets out the genre, names the two sentences that carry the content, and lists what to discount entirely.
Capability announcements follow a stable form. A headline claim with a superlative. Two or three benchmark comparisons against named competitors. A selection of demonstrations, usually video, usually edited. A safety section. A note about availability. The form is stable because it works, and it works because most readers stop after the headline.
Almost none of the information content is in the first four sections. The benchmarks are chosen after the results are known, which is not fraud but is selection. The demonstrations are best cases, and the number of attempts behind each one is never stated.
The first is the availability sentence. General availability today, at a stated price, with a stated rate limit, is a strong claim: the organisation is committing capacity and taking liability. A waitlist, a limited preview or a research demonstration is a much weaker one, and the gap between the two has historically run to many months and occasionally to never.
The second is the pricing sentence, when there is one. Price per unit tells you roughly what the thing costs to serve, subject to whatever the provider is willing to subsidise, and a capability priced at ten times its predecessor is telling you something about compute per request that no benchmark table will.
Comparisons against a competitor's older release. Aggregate scores where the component tasks are not listed. Any percentage improvement stated without the baseline. Demonstrations of tasks with no cheap verifier, since those are the ones where failure is least visible. And safety sections that describe process rather than result — an evaluation was conducted is not a finding.
None of this implies bad faith on the part of the authors. It is what marketing documents are for. The error is reading them as if they were papers.
Longer, on the same themesAll essays
Eye for AI · 4 September 2026
Every good test of machine ability eventually becomes a bad one.
Eye for AI · 21 August 2026
A demonstration proves that a system can succeed.
Eye for AI · 7 August 2026
The clearest historical evidence on automation and employment is not about looms or robots.