When a Benchmark Stops Measuring
Eye for AI · 4 September 2026
Every good test of machine ability eventually becomes a bad one.
Volume III · Number 6 · September 2026
Independent and reader funded. No advertising, no sponsored posts, no affiliate links. editor@eyeforai.blog
Contents/Essays/What Automation Actually Did to the Typing Pool
Labour
The clearest historical evidence on automation and employment is not about looms or robots. It is about clerical work, and what it shows is more unsettling than either side usually admits.
Abstract
Clerical automation, not the power loom, is the useful precedent for the present wave. This essay follows an occupation that shrank through a closed hiring pipeline rather than mass dismissal, examines the bank-teller counter-argument and the elasticity condition usually dropped when it is quoted, and asks who absorbed the cost the aggregate figures hide.
Discussion of automation reaches reflexively for the loom and the assembly line, and both are poor analogies for what is happening now. Those technologies replaced physical effort in manufacturing, a sector that was already concentrated, unionised, and geographically visible. The disruption was loud, and its loudness is why we remember it.
Clerical work is the better comparison because it was information work: reading, transcribing, filing, checking, routing, calculating. It was performed largely in offices, largely by women, and at its peak it accounted for a very substantial share of all employment in industrialised economies. It was then subjected, over roughly four decades, to a sequence of automating technologies: the photocopier, the electronic calculator, the word processor, the spreadsheet, the relational database, and finally the networked personal computer.
We have unusually good data on what happened, because censuses tracked occupational categories carefully during exactly this period. And what happened does not support either the comfortable story or the catastrophic one.
Take the typist and the stenographer. These were real, numerous, defined occupations with training pipelines, professional associations and career ladders. A competent stenographer had a skill that took years to acquire and that commanded a wage premium over general clerical work.
Word processing did not fire them. What it did was fold their task into somebody else's job. The manager who previously dictated now typed, because typing had become cheap enough in cognitive effort that the coordination cost of delegating exceeded the cost of doing it. The occupation did not shrink because people were dismissed from it in large numbers; it shrank because nobody entered it. Hiring stopped, existing staff aged out or moved sideways, and within two decades a category that had employed millions was a rounding error.
This matters because the political and statistical signatures are completely different. Mass dismissal produces protests, headlines and policy responses. A slow closure of entry produces nothing visible at all, and its cost falls almost entirely on people who never got the job in the first place and therefore never appear in any dataset of the displaced.
The standard counterexample deserves its status. Automated teller machines were introduced from the late 1960s and spread rapidly. The obvious prediction was that teller employment would collapse. It did not; teller numbers in the United States rose for roughly three decades after ATM deployment began.
The mechanism is well understood. ATMs reduced the cost of operating a branch, so banks opened more branches, so total teller employment rose even as tellers per branch fell. Meanwhile the job changed: less cash handling, more sales and customer service. The technology was labour-saving per unit of output and employment-increasing in aggregate, because demand expanded to absorb the saving.
This is a genuine and important result and I do not want to explain it away. But it is often quoted with a crucial condition dropped. It works when demand is elastic. If cheaper branch operation had not produced more branches, the arithmetic would have run the other way, and the same mechanism that saved teller employment would have destroyed it. The ATM story is not a law; it is a case where a particular elasticity happened to be high.
Aggregate employment is the wrong unit of analysis and it has been distorting this debate for thirty years. Total employment recovers because new work appears; that is true, well-evidenced, and almost irrelevant to the person whose specific occupation went away at fifty-two in a town with one employer.
The literature on trade shocks, which is methodologically much stronger than most automation research because the exposure is measurable at regional level, finds large and persistent local effects: depressed wages, reduced labour force participation, and worse health outcomes lasting a decade or more in affected areas, even while national employment recovers fully. There is no obvious reason automation shocks would behave differently, and some reason to think they diffuse more broadly and are therefore harder to target with policy.
The distributional question is also a compositional one. When a task is automated, the residual job is not a smaller version of the old job; it is the leftover parts. Sometimes those parts are the interesting ones and the job improves. Frequently they are the tedious ones that resisted automation precisely because they are irregular, and the job gets worse while the title stays the same.
Two things are genuinely different. The first is the breadth of exposure: previous waves hit one task family at a time, and the current one touches drafting, summarising, coding, translating, and first-pass analysis simultaneously. Breadth matters because the standard adjustment mechanism is workers moving to an adjacent occupation, and adjacency helps less when the adjacent occupation is exposed too.
The second is direction. Earlier computerisation hollowed out the middle: it automated routine codifiable tasks while leaving both high-end judgement work and low-end manual work intact. That produced a well-documented pattern of wage polarisation. The current wave points at parts of the high end, which is a different shape and one we have less historical evidence about.
What is not different is the timescale, and this is where I part company with most commentary in both directions. Clerical automation took four decades to work through, and the binding constraint was almost never the technology. It was organisational redesign, retraining, capital replacement cycles and, repeatedly, the simple fact that firms are slow. There is no evidence I find credible that these frictions have disappeared.
I do not know how many jobs this will displace and neither does anyone else. The published estimates span an order of magnitude and their spread is driven almost entirely by assumptions the authors put in at the start, particularly about whether exposure at task level translates into displacement at job level. It usually has not.
What the historical record supports is narrower and still worth saying. Automation of a task family rarely eliminates the occupation outright; it changes the job, closes the entry pipeline, and shifts the wage. Aggregate employment tends to recover. Specific cohorts and specific regions frequently do not, and the recovery of the aggregate is no comfort to them whatsoever.
If that sounds like a refusal to make a forecast, it is. The forecasts on offer are not measurements and dressing them in percentages does not make them so. What we can do, and what the historical evidence actually licenses, is prepare for a distributional problem rather than a level one — which is a different policy question, with different answers, and it is being asked much less often than it should be.
Editor’s note
The clerical figures here come from occupational census categories, which are not stable across decades: definitions were revised repeatedly, and some apparent decline is reclassification rather than disappearance. I have tried to describe directions rather than magnitudes for exactly this reason.
References and further reading
Filed under Labour & Work
Elsewhere in this issueAll essays
Eye for AI · 4 September 2026
Every good test of machine ability eventually becomes a bad one.
Eye for AI · 21 August 2026
A demonstration proves that a system can succeed.
Eye for AI · 24 July 2026
A single phrase now covers four research programmes with different evidence bases, different timescales, and occasionally opposed prescriptions.