Eyefor AI

Volume III · Number 6 · September 2026


Independent and reader funded. No advertising, no sponsored posts, no affiliate links. editor@eyeforai.blog

Contents/Essays

The archive

Essays

Six long pieces, each arguing something specific. They are meant to be read whole rather than skimmed, and they are ordered by publication date rather than by importance.

Contents

  • 01

    When a Benchmark Stops Measuring

    Eye for AI · 4 September 2026

    Every good test of machine ability eventually becomes a bad one. The interesting question is not whether that happens but how quickly, and what we are entitled to conclude in the interval.

    18 min read
  • 02

    The Distance Between a Demo and a Deployment

    Eye for AI · 21 August 2026

    A demonstration proves that a system can succeed. A deployment requires that it rarely fail, in a specific way, for a specific person, on a Tuesday. These are almost unrelated problems.

    16 min read
  • 03

    What Automation Actually Did to the Typing Pool

    Eye for AI · 7 August 2026

    The clearest historical evidence on automation and employment is not about looms or robots. It is about clerical work, and what it shows is more unsettling than either side usually admits.

    19 min read
  • 04

    ‘AI Safety’ Means Four Different Things

    Eye for AI · 24 July 2026

    A single phrase now covers four research programmes with different evidence bases, different timescales, and occasionally opposed prescriptions. The conflation is costing all four of them.

    17 min read
  • 05

    The Arithmetic of Inference

    Eye for AI · 9 July 2026

    Training costs make headlines. Inference costs decide what actually gets built, and they behave in ways that make the standard cost-collapse narrative less reassuring than it sounds.

    15 min read
  • 06

    Emergence, and the Trouble with Thresholds

    Eye for AI · 18 June 2026

    Some abilities appear to arrive suddenly as models grow. A careful line of criticism argues the suddenness is an artefact of how we score. Both sides are partly right, and the disagreement is more interesting than either.

    16 min read

Filed under

  • Evaluation

    Standing section

    How we measure machine ability, why the measurements decay, and what a serious evaluation would have to look like.

    5 essays
  • Labour & Work

    Standing section

    What automation has historically done to occupations, who absorbed the cost, and why aggregate employment figures answer the wrong question.

    2 essays
  • Policy

    Standing section

    Regulation, procurement, compute thresholds and open weights — read as decisions with trade-offs rather than as a contest between good and bad actors.

    2 essays