When a Benchmark Stops Measuring
Eye for AI · 4 September 2026
Every good test of machine ability eventually becomes a bad one.
Volume III · Number 6 · September 2026
Independent and reader funded. No advertising, no sponsored posts, no affiliate links. editor@eyeforai.blog
Contents/Essays/‘AI Safety’ Means Four Different Things
Policy
A single phrase now covers four research programmes with different evidence bases, different timescales, and occasionally opposed prescriptions. The conflation is costing all four of them.
Abstract
Four distinct programmes share one phrase. This essay separates them, shows that the conflation is useful to several parties and therefore stable, weighs the evidence base behind each, and demonstrates that their prescriptions conflict — so a rule written for one will often obstruct another. It ends with a request rather than a position.
The first is product safety: does this system, deployed in this context, harm the people using it? Does the medical triage tool underperform for a particular group? Does the chatbot give dangerous advice to a teenager? This is ordinary consumer protection, it has a mature methodology, and its evidence base is the strongest of the four because the harms are observed rather than projected.
The second is misuse: can a competent bad actor use this system to do something they could not otherwise do at that cost? Fraud at scale, non-consensual imagery, targeted harassment, assistance with weapons development. The relevant question is always marginal uplift over existing tools, and this is where most public argument goes wrong, because a capability that sounds alarming is only a safety issue if it is meaningfully easier than the alternative.
The third is systemic effect: what happens to institutions when this is everywhere? Information ecosystems, labour markets, concentration of compute and capital, dependence of public services on a handful of private providers. These harms are diffuse, slow, and hard to attribute, which makes them chronically underweighted relative to how much they probably matter.
The fourth is loss of control: could a sufficiently capable system pursue objectives its operators did not intend and resist correction? This is the one that gets the coverage, it has the weakest empirical base of the four by a wide margin, and it is also the only one where being wrong in the optimistic direction is unrecoverable.
Bundling has advantages for everyone involved, which is why it persists. For researchers working on speculative long-horizon risk, association with documented near-term harms lends urgency and evidentiary weight the underlying work does not have on its own. For firms, a regulatory conversation about hypothetical future systems is considerably more comfortable than one about the discrimination in the model they shipped last quarter.
For advocacy organisations working on near-term harm, the bundle is a mixed inheritance. It brings attention and money into a space that had neither, and it also means their carefully documented findings about hiring algorithms arrive in the same news cycle as speculation about extinction, and get filed by policymakers under the same heading. The reasonable irritation this produces has hardened into a factional dispute that helps nobody.
The costs of bundling fall unevenly and they are real. Regulatory attention is finite. When a single agency mandate is written to cover all four, it will in practice be staffed and executed according to whichever framing is loudest, and the other three get a paragraph.
Product safety harms are documented, replicated, and in several cases litigated. We have specific findings on differential error rates in deployed classification systems, on recommendation systems and adolescent mental health, on automated benefits decisions producing wrongful denials at scale. The evidence is not perfect and effect sizes are contested, but this is a normal empirical literature with normal disagreements.
Misuse evidence is mixed and the honest summary is that uplift has been smaller than early warnings suggested in some domains and larger in others. Fraud and social engineering: real, measurable, already happening at scale. Biological and chemical uplift: the published red-team work is genuinely inconclusive, the classified work is unavailable for scrutiny, and the controlled studies that exist have small samples and questionable controls.
Loss-of-control evidence consists of laboratory demonstrations of specification gaming, reward hacking, and deceptive behaviour under contrived conditions. These are real results and they are not nothing: a system optimising a proxy will exploit the proxy, reliably, and this has been shown many times. What they do not establish is the extrapolation from a system exploiting a scoring bug in a game to a system strategically resisting shutdown. That extrapolation is an argument, sometimes a careful one, but it is not a measurement and it is regularly presented as though it were.
This is the part that gets least attention and matters most. The four programmes do not merely differ in emphasis; their policy recommendations point in different directions.
Open weights are the clearest case. For systemic-effect concerns, open release is close to essential: it is the only realistic check on concentration, it enables independent audit, and it prevents a handful of firms from becoming the sole arbiters of what these systems will do. For misuse concerns, open release removes every safeguard permanently and irreversibly, since a released weight cannot be recalled. Both of these are correct. They cannot both be acted on.
Compute thresholds show the same structure. Regulating above a training-compute threshold targets frontier loss-of-control risk while exempting almost every system that has actually harmed anyone, since the documented product-safety failures come overwhelmingly from small, cheap, unremarkable models deployed carelessly in consequential settings. A threshold regime is not a compromise between the four programmes. It is a decision to fund one.
The practical fix is unglamorous: say which one you mean. A claim about safety that does not specify the programme cannot be evaluated, because the standard of evidence, the relevant expertise and the appropriate intervention all differ. I would extend the same demand to critics: 'AI safety is a distraction' is four separate claims, and at least one of them is clearly false.
I would also ask for honesty about confidence levels within each programme. It is entirely coherent to hold that product-safety harms are demonstrated and urgent, that misuse uplift is real but overstated in the domains that get the most coverage, that systemic effects are the most underrated of the four, and that loss-of-control risk is poorly evidenced but severe enough in the tail to justify serious funded work. That is roughly where I sit, and it is not a position either faction is set up to represent.
What I am confident about is the negative claim. Anyone who tells you the four questions have the same answer, or that three of them are settled, is doing something other than reasoning from the evidence. The evidence is uneven, and a serious position has to be uneven in the same places.
Editor’s note
A disclosure of priors: I think systemic effects are the most neglected of the four and I have written more about them than the others. Readers should weigh the section above with that in mind.
References and further reading
Filed under Policy · Evaluation
Elsewhere in this issueAll essays
Eye for AI · 4 September 2026
Every good test of machine ability eventually becomes a bad one.
Eye for AI · 21 August 2026
A demonstration proves that a system can succeed.
Eye for AI · 7 August 2026
The clearest historical evidence on automation and employment is not about looms or robots.