Eyefor AI

Volume III · Number 6 · September 2026


Independent and reader funded. No advertising, no sponsored posts, no affiliate links. editor@eyeforai.blog

Contents/Sections/Evaluation

Section

Evaluation

Abstract

How we measure machine ability, why the measurements decay, and what a serious evaluation would have to look like.

The problem in one paragraph

Nearly every public claim about what these systems can do rests on a benchmark, and nearly every benchmark is a proxy that was accurate when it was built and has been under attack ever since. Evaluation is not a side discipline within machine learning research. It is the part that determines whether any of the rest of it means anything, and it receives a small fraction of the attention and none of the glamour.

What we look for in this section

Whether the measurement predates the training data. Whether the scorer has an interest in the outcome. Whether the metric is continuous or discontinuous, and whether that choice was made for the researcher's convenience or the user's reality. Whether anybody has run the perturbed variant, and whether the number of prior attempts against the same test set is disclosed.

We are not looking for reasons to dismiss results. Dismissal is cheap and the field has plenty of it. We are looking for the conditions under which a number is worth acting on, which is a harder and more useful question.

Recurring themes

Goodhart's law as an empirical prediction rather than an aphorism. The difference between capability and reliability, which is the difference between a demonstration and a product. Task-grounded evaluation and why almost nobody runs it. And the persistent gap between predictable aggregate loss and unpredictable downstream competence.

Annotated bibliographyAll essays

  • 01

    When a Benchmark Stops Measuring

    Eye for AI · 4 September 2026 · 18 min read

    The anchor piece for the section. Everything else here assumes its central distinction: that a score is evidence about a capability only for as long as nobody is being paid to move it.

    Every good test of machine ability eventually becomes a bad one. The interesting question is not whether that happens but how quickly, and what we are entitled to conclude in the interval.

  • 02

    The Distance Between a Demo and a Deployment

    Eye for AI · 21 August 2026 · 16 min read

    The applied companion to the benchmark essay. Where that piece asks whether a number means anything, this one asks what would have to be measured instead before a system is put in front of people.

    A demonstration proves that a system can succeed. A deployment requires that it rarely fail, in a specific way, for a specific person, on a Tuesday. These are almost unrelated problems.

  • 04

    ‘AI Safety’ Means Four Different Things

    Eye for AI · 24 July 2026 · 17 min read

    Belongs here because three of its four programmes are evaluation problems in disguise: each needs a different measurement, and the phrase they share hides that.

    A single phrase now covers four research programmes with different evidence bases, different timescales, and occasionally opposed prescriptions. The conflation is costing all four of them.

  • 05

    The Arithmetic of Inference

    Eye for AI · 9 July 2026 · 15 min read

    Included for its treatment of price as a measurement. Cost per served token is one of the few figures in this field with checkable provenance, which makes it unusually useful evidence.

    Training costs make headlines. Inference costs decide what actually gets built, and they behave in ways that make the standard cost-collapse narrative less reassuring than it sounds.

  • 06

    Emergence, and the Trouble with Thresholds

    Eye for AI · 18 June 2026 · 16 min read

    The sharpest case in the archive of a measurement artefact being read as a fact about the world. It is also the essay most sympathetic to the position it argues against.

    Some abilities appear to arrive suddenly as models grow. A careful line of criticism argues the suddenness is an artefact of how we score. Both sides are partly right, and the disagreement is more interesting than either.

Other sections

  • Labour & Work

    Standing section

    What automation has historically done to occupations, who absorbed the cost, and why aggregate employment figures answer the wrong question.

    2 essays
  • Policy

    Standing section

    Regulation, procurement, compute thresholds and open weights — read as decisions with trade-offs rather than as a contest between good and bad actors.

    2 essays