


Historical Quarter, University of Oslo
Three hundred years of nearly
1637 to 1993
The first test
…they could never use words or other signs arranged in such a manner as is competent to us in order to declare our thoughts to others… but not that it should arrange them variously so as appositely to reply to what is said in its presence, as men of the lowest grade of intellect can do.
Flexible, appropriate use of language
The second test
…while reason is an universal instrument that is alike available on every occasion, these organs, on the contrary, need a particular arrangement for each particular action.
General reasoning that transfers to new situations
Descartes. Published anonymously in Leiden.
He wrote two tests. Turing wrote one, three hundred and thirteen years later. And Descartes wrote his to prove he was not a machine.
Draaisma, Metaphors of Memory · Cobb, The Idea of the Brain. Hobbes: “For what is the Heart, but a Spring; and the Nerves, but so many Strings…” The golem recipe is thirteenth-century Ashkenaz, in Idel, Golem (SUNY, 1990), 56–57 and 64–65. Not Talmudic.



The flute player. Real.
Three sets of bellows, three blowing pressures, a cam cylinder working fingers, tongue and lips. It played an actual flute.
The digesting duck. Fake.
The food never left the mouth tube. The pellet was loaded in advance. Friedrich Nicolai exposed it in 1783. The credit goes to Robert-Houdin, who published sixty-two years later.
Riskin, Critical Inquiry 29:4 (2003). The duck is dated 1738 or 1739; sources disagree.
What the room saw
A figure at a cabinet, playing chess, winning. For seventy years. Europe called it the Turk.
When we found out
Poe published on it in 1836 and got it wrong. The real account came in 1857, three years after the machine had burned. In 1819 Charles Babbage took notes in the margins of a book about it, then came back to play it.
Wolfgang von Kempelen’s chess automaton. Built 1769, first exhibited 1770, burned 1854. Riskin, Critical Inquiry 29:4, citing Schaffer, “Enlightened Automata”.

A word about a word
Usually a woman. Harvard hired more than eighty of them from 1877, starting at twenty-five cents an hour. Henrietta Swan Leavitt worked out how to measure the distance to the stars. Her result went out in 1912 signed by the director, opening: “prepared by Miss Leavitt.” The six people who programmed ENIAC came out of the same pool: two hundred women computing firing tables at the Moore School. The machine took the job. It kept the name.
Grier, When Computers Were Human (Princeton, 2005), dating the epoch 1758 to 1986.
1843. Note A.
…the engine might compose elaborate and scientific pieces of music of any degree of complexity or extent… The Analytical Engine weaves algebraical patterns just as the Jacquard-loom weaves flowers and leaves.
Ada Lovelace. Babbage built a calculator. She published the argument that it was a general machine, in notes appended to a translation of somebody else’s paper.


Flowers burned the records. Secret until 1975.
The full account came out in October 2000.
Fifty-six years late, and ENIAC took the credit.
Mind, October 1950
I believe that in about fifty years’ time it will be possible to programme computers, with a storage capacity of about 109, to make them play the imitation game so well that an average interrogator will not have more than 70 per cent chance of making the right identification after five minutes of questioning.
Alan Turing

31 August 1955
We propose that a 2 month, 10 man study of artificial intelligence be carried out during the summer of 1956…
McCarthy, Minsky, Rochester and Shannon. They asked for $13,500 and got $7,500. McCarthy lost the attendance records.


Cambridge, 1955
She founded the Cambridge Language Research Unit in 1955 and argued that machine translation built on syntax alone was doomed, because meaning is not in the grammar. In 1966 the American money for machine translation stopped, the field having failed in exactly the way she said it would. She was inside it. Nobody in charge listened.
Wilks (ed.), Language, Cohesion and Form: Margaret Masterman (Cambridge, 2005). Her colleague Karen Spärck Jones invented the weighting that essentially every search engine still uses.
1973
In no part of the field have the discoveries made so far produced the major impact that was then promised.
The Lighthill report. British funding cut to a handful of centres. The BBC televised the argument at the Royal Institution on 9 May 1973.

2012 to now
What won
AlexNet. Top-five error of 15.3 per cent against 26.2 for second place, in a competition where a good year moved it one point.
What was old
Backpropagation, 1986. Convolutional networks, 1989. The compute was new. So was a dataset that somebody had to go and build.
Krizhevsky, Sutskever and Hinton, NIPS 2012.
Somebody built the data
Fei-Fei Li started ImageNet on one assumption, and I am quoting her: “even the best algorithm would not generalize well if the data it learned from did not reflect the real world.” Fifteen million images. The model won on the labels. Fifty thousand people made the labels.
Li and Krishna, “Searching for Computer Vision North Stars”, Dædalus 151:2 (2022). She puts AlexNet’s margin at “41 per cent”; that is the relative fall in error, 26.2 to 15.3.
What the paper claimed
The transformer trains faster, because it parallelises. That is the stated contribution.
What it turned out to be
The scaling behaviour across four orders of magnitude turns up three years later, in other people’s papers. Nobody in this room can tell you which paper from this year matters.

Training compute of notable AI systems, 1955 to 2026. Logarithmic scale.
Epoch AI, notable AI models dataset, retrieved 24 August 2026, CC-BY. Hardware comparison: Nordhaus, Journal of Economic History 67 (2007), Table 6. Pre-1990 figures are reconstructions, not measurements.
The universe is 13.8 billion years old, if you want somewhere to put the last row.
Converted from the same Epoch AI figures as the chart. Perceptron 7.2×108 FLOP, Neocognitron 2.7×108, AlexNet 4.7×1017, Grok 4 5×1026. Epoch rates the Grok 4 figure Speculative, 90 per cent interval ±31×, and that interval is where the thousand-fold range comes from.
Add hardware to a rule-based system
and it fails faster.
Epoch AI, current to 5 February 2026. A frontier training run now draws tens to hundreds of megawatts. A medium-sized power plant.
Descartes’ first test
It fell so completely that we stopped calling it a test.
Descartes’ second test
His words: an universal instrument that is alike available on every occasion.
ARC-AGI-3, March 2026, on environments no system has seen before.
Humans 100 per cent. Frontier systems 0.51 per cent.

The people saying this most plainly are the ones building the tests. Their own retrospective, December 2025: “Do we have AGI? Not yet.”
ARC Prize. ARC-AGI-1, o3 at 87.5% on the semi-private set, 20 December 2024. ARC-AGI-2, 54% verified, Poetiq on Gemini 3, December 2025, at $30.57 a problem; the grand prize, 85% at $0.42 a task, is unclaimed. ARC-AGI-3, humans 100%, frontier systems 0.51% at launch.
The case for
Five times the compute a year, compounding with three times the efficiency. A hundred and thirty-one day doubling in how long a system can work unsupervised. And in May, a general reasoning model disproved a conjecture of Erdős from 1946, checked by outside mathematicians.
If a human had written the paper… I would have recommended acceptance without any hesitation.
Timothy Gowers, Fields Medallist
The case against
The same systems score 0.51 where humans score 100. They read an analogue clock about half the time. In the one randomised trial on real work, sixteen experienced developers were 19 per cent slower and believed they were 20 per cent faster. Fewer than one in five American businesses use them at all.
We are near the end of the exponential.
Dario Amodei, February 2026. He means the returns to pure scaling,
and he says it to argue for urgency rather than slowdown.
Bender and Koller, 2020
Two people are stranded on separate islands, talking through an undersea cable. A hyper-intelligent octopus taps the line. It never sees either island. It learns to predict the replies with great accuracy, then cuts the cable and impersonates one of them. It passes, right up until somebody needs a coconut catapult built. Emily Bender and Alexander Koller: a system trained only on form has no way to learn meaning.
Bender and Koller, “Climbing towards NLU”, ACL 2020, Best Theme Paper. Descartes’ first test, restated.
Four things it took me too long to learn
The direction held.
The timing was out by sixty-eight years.
The 2026 forecasts have the same shape.
One
You already knew about the duck. You had not heard of the flute player. That asymmetry is running whenever you choose what to fund or what to worry about.
Two
The most quoted prediction in this field was denied by the man it is attributed to. The famous forty million dollar saving has never been the same number twice. The Turk was exposed three years after it had already burned. And the golem of Prague, which feels medieval, is attached to Rabbi Loew in print for the first time in 1834. It is about as old as the railway.
Dekel and Gurley, “How the Golem Came to Prague”, Jewish Quarterly Review 103:2 (2013), dating the Prague attribution to Kohn, Das jüdische Gil Blas (Leipzig, 1834).
Three.
In considering any new subject, there is frequently a tendency, first, to overrate what we find to be already interesting or remarkable; and, secondly, by a sort of natural reaction, to undervalue the true state of the case…
Ada Lovelace, 1843. That is the hype cycle, written down by someone who had seen exactly one machine, and it did not exist yet.

Blinded evaluation, JAMA 242(12), 1979. Liability. No route into hospital workflow. No clinician willing to type answers to an interrogation.
The 1980s died in that gap.
I spend most of my time in it.
Stewart Brand, forty-one years apart
Whole Earth Catalog, Fall 1968
We are as gods and might as well get good at it.
Written against “government, big business, formal education, church”. Brand: “Credit where it’s due: I stole the line.”
Whole Earth Discipline, 2009
We are as gods and HAVE to get good at it.
“Necessity comes from climate change, potentially disastrous for civilization.”
Might as well became have to.
Everybody quotes Larry Tesler saying that. He corrected it himself, on his own website. What he actually said was:
“Intelligence is whatever machines haven’t done yet.”
The first is a joke about a research field. The second is a claim about us.
For three hundred and eighty-nine years,
every test we have written since has pointed the same way.
At the machine.
The imitation game, 2026
Behind one door, your own thinking. Behind the other, your thinking with one of these systems in the loop. In a study published this February, people solving logic problems with AI scored three points above the norm, and rated themselves four points too high. The Dunning-Kruger effect stopped appearing. Everyone was overconfident. The people with the most AI literacy judged themselves worst of all.
Fernandes et al., “AI makes you smarter but none the wiser”, Computers in Human Behavior 175 (2026). Two studies, roughly 700 participants, internally replicated.
I think, therefore I am.
That is the one claim in this talk
that nobody else can verify for you.
Check it now and then.
The long history
Descartes, Discourse on the Method (1637)
Riskin, Critical Inquiry 29 (2003)
Grier, When Computers Were Human (2005)
Lovelace, Notes A and G (1843)
Turing, Mind LIX:236 (1950)
Wilks (ed.), Margaret Masterman (2005)
The last fourteen years
Li & Krishna, Dædalus 151:2 (2022)
Krizhevsky, Sutskever & Hinton, NIPS (2012)
Epoch AI, notable models dataset · trends
Nordhaus, J. Economic History 67 (2007)
ARC Prize · METR · US Census BTOS
Bender & Koller, ACL (2020)
What it is for
Idel, Golem (SUNY, 1990)
Dekel & Gurley, JQR 103:2 (2013)
Yu et al., JAMA 242:12 (1979)
Brand, Whole Earth Catalog (1968); Discipline (2009)
Fernandes et al., Computers in Human Behavior 175 (2026)