QR code
Thordur Arnason
Thordur Arnason
Global Head of AI Enablement, Capgemini Invent
Chair of the Board, Gervi Labs

Historical Quarter, University of Oslo

I think,
therefore I am
(a machine)

Three hundred years of nearly

Thinking machines have been
twenty years away
for three hundred years.

ONE

Three hundred years of nearly

1637 to 1993

The first test

…they could never use words or other signs arranged in such a manner as is competent to us in order to declare our thoughts to others… but not that it should arrange them variously so as appositely to reply to what is said in its presence, as men of the lowest grade of intellect can do.

Flexible, appropriate use of language

The second test

…while reason is an universal instrument that is alike available on every occasion, these organs, on the contrary, need a particular arrangement for each particular action.

General reasoning that transfers to new situations

1637

Descartes. Published anonymously in Leiden.

He wrote two tests. Turing wrote one, three hundred and thirteen years later. And Descartes wrote his to prove he was not a machine.

Every era describes the mind
with the most powerful thing it has

  • A word, written on the forehead. Ashkenaz, thirteenth century. Write emet and it stands. Erase the aleph, leaving met, and it falls.
  • Clockwork and springs. Hobbes, 1651.
  • Hydraulics
  • Electricity, after Galvani and Volta
  • Telecommunications wiring
  • Logic circuits. Then parallel processors.
  • And now a word again.

Draaisma, Metaphors of Memory · Cobb, The Idea of the Brain. Hobbes: “For what is the Heart, but a Spring; and the Nerves, but so many Strings…” The golem recipe is thirteenth-century Ashkenaz, in Idel, Golem (SUNY, 1990), 56–57 and 64–65. Not Talmudic.

A clay golem with an inscribed slip, then clockwork, hydraulics, electricity, telecommunications, valves, logic circuits and a server rack, all connected by one continuous signal path.

Paris, 1738. Two machines.

Vaucanson’s Flute Player with bellows, three airflow paths and cam-driven controls exposed.
Vaucanson’s Digesting Duck showing the feeding path and the separate preloaded pellet path.

The flute player. Real.

Three sets of bellows, three blowing pressures, a cam cylinder working fingers, tongue and lips. It played an actual flute.

The digesting duck. Fake.

The food never left the mouth tube. The pellet was loaded in advance. Friedrich Nicolai exposed it in 1783. The credit goes to Robert-Houdin, who published sixty-two years later.

Riskin, Critical Inquiry 29:4 (2003). The duck is dated 1738 or 1739; sources disagree.

The flute player was real
and is forgotten.
The duck was fake
and is famous.

Built 1769. A man in a box.

What the room saw

A figure at a cabinet, playing chess, winning. For seventy years. Europe called it the Turk.

When we found out

Poe published on it in 1836 and got it wrong. The real account came in 1857, three years after the machine had burned. In 1819 Charles Babbage took notes in the margins of a book about it, then came back to play it.

Wolfgang von Kempelen’s chess automaton. Built 1769, first exhibited 1770, burned 1854. Riskin, Critical Inquiry 29:4, citing Schaffer, “Enlightened Automata”.

The Mechanical Turk cut open to reveal the concealed chess master and pantograph linkage.

A word about a word

For three hundred years,
a computer was a person.

Usually a woman. Harvard hired more than eighty of them from 1877, starting at twenty-five cents an hour. Henrietta Swan Leavitt worked out how to measure the distance to the stars. Her result went out in 1912 signed by the director, opening: “prepared by Miss Leavitt.” The six people who programmed ENIAC came out of the same pool: two hundred women computing firing tables at the Moore School. The machine took the job. It kept the name.

Grier, When Computers Were Human (Princeton, 2005), dating the epoch 1758 to 1986.

1843. Note A.

…the engine might compose elaborate and scientific pieces of music of any degree of complexity or extent… The Analytical Engine weaves algebraical patterns just as the Jacquard-loom weaves flowers and leaves.

Ada Lovelace. Babbage built a calculator. She published the argument that it was a general machine, in notes appended to a translation of somebody else’s paper.

The small constructed Difference Engine fragment contrasted with a ghosted, unrealized larger engine.

1943. Colossus.
Four things to get right

  • It attacked the Lorenz teleprinter cipher. Enigma was a different machine, in a different hut.
  • It ran one job. No stored program, and not Turing complete.
  • Tommy Flowers built it, part-funded from his own pocket when the Post Office would not pay.
  • Three quarters of Bletchley was women. On the Enigma machines, 1,676 of them against 263 men.
Lorenz teleprinter ciphertext on paper tape entering the valve racks of Colossus.

Then they broke the machines
into pieces too small
to work out what they did.

Flowers burned the records. Secret until 1975.
The full account came out in October 2000.

Fifty-six years late, and ENIAC took the credit.

Mind, October 1950

I believe that in about fifty years’ time it will be possible to programme computers, with a storage capacity of about 109, to make them play the imitation game so well that an average interrogator will not have more than 70 per cent chance of making the right identification after five minutes of questioning.

Alan Turing

An interrogator communicating by teleprinter with a hidden human and a hidden computer.

What he wrote

31 August 1955

We propose that a 2 month, 10 man study of artificial intelligence be carried out during the summer of 1956…

McCarthy, Minsky, Rochester and Shannon. They asked for $13,500 and got $7,500. McCarthy lost the attendance records.

The Dartmouth artificial-intelligence proposal with the phrase two month, ten man emphasized.

Then everybody started
writing cheques

  • 1958. Within ten years, a computer will be world chess champion. It took thirty-nine.
  • 1958. The New York Times: a machine that will “walk, talk, see, write, reproduce itself and be conscious of its existence.”
  • 1967. Minsky: within a generation, substantially solved.
  • 1970. “Three to eight years.” Minsky denied ever saying it.
The physical Mark I Perceptron with sensory grid, weighted connections and response unit.

Cambridge, 1955

Margaret Masterman spent
a decade saying it
would not work.

She founded the Cambridge Language Research Unit in 1955 and argued that machine translation built on syntax alone was doomed, because meaning is not in the grammar. In 1966 the American money for machine translation stopped, the field having failed in exactly the way she said it would. She was inside it. Nobody in charge listened.

Wilks (ed.), Language, Cohesion and Form: Margaret Masterman (Cambridge, 2005). Her colleague Karen Spärck Jones invented the weighting that essentially every search engine still uses.

1973

In no part of the field have the discoveries made so far produced the major impact that was then promised.

The Lighthill report. British funding cut to a handful of centres. The BBC televised the argument at the Royal Institution on 9 May 1973.

The Lighthill report connected to a televised debate between Lighthill and three AI researchers.
TWO

The last fourteen years

2012 to now

2012. The ideas were already twenty years old.

What won

AlexNet. Top-five error of 15.3 per cent against 26.2 for second place, in a competition where a good year moved it one point.

What was old

Backpropagation, 1986. Convolutional networks, 1989. The compute was new. So was a dataset that somebody had to go and build.

Krizhevsky, Sutskever and Hinton, NIPS 2012.

Somebody built the data

Fifty thousand workers.
167 countries.
Three years.

Fei-Fei Li started ImageNet on one assumption, and I am quoting her: “even the best algorithm would not generalize well if the data it learned from did not reflect the real world.” Fifteen million images. The model won on the labels. Fifty thousand people made the labels.

Li and Krishna, “Searching for Computer Vision North Stars”, Dædalus 151:2 (2022). She puts AlexNet’s margin at “41 per cent”; that is the relative fall in error, 26.2 to 15.3.

2017. Nobody knew what they had.

What the paper claimed

The transformer trains faster, because it parallelises. That is the stated contribution.

What it turned out to be

The scaling behaviour across four orders of magnitude turns up three years later, in other people’s papers. Nobody in this room can tell you which paper from this year matters.

AlexNet’s convolutional feature hierarchy beside the Transformer’s encoder-decoder attention architecture.

Hardware was never the thing that was missing.

Training compute of notable AI systems, 1955 to 2026. Logarithmic scale.

First winter Second winter 103 106 109 1012 1015 1018 1021 1024 1027 1960 1970 1980 1990 2000 2010 2020 2026 1966 to 1975: nothing recorded labs stoppedpublishing Perceptron Mark I ADALINE Cognitron Neocognitron NETtalk TD-Gammon LeNet-5 AlexNet Transformer AlphaGo Zero GPT-3 GPT-4 Grok 4 1960 to 1980: AI training computedid not grow. It fell, by about 2.6×.Over the same twenty years the cost of acomputation fell roughly 700-fold. Epoch’sfrontier list is empty from 1966 to 1980.

Epoch AI, notable AI models dataset, retrieved 24 August 2026, CC-BY. Hardware comparison: Nordhaus, Journal of Economic History 67 (2007), Table 6. Pre-1990 figures are reconstructions, not measurements.

The same chart, in a unit you can feel

1960
One second
Rosenblatt’s perceptron. Everything else is scaled to this.
1980
Under half a second
Neocognitron. Twenty years later, and less than it started with.
2012
Twenty years
AlexNet.
2025
0.7 to 700 billion years
The frontier. That thousand-fold range is the disclosure gap, in time.

The universe is 13.8 billion years old, if you want somewhere to put the last row.

Converted from the same Epoch AI figures as the chart. Perceptron 7.2×108 FLOP, Neocognitron 2.7×108, AlexNet 4.7×1017, Grok 4 5×1026. Epoch rates the Grok 4 figure Speculative, 90 per cent interval ±31×, and that interval is where the thousand-fold range comes from.

The 1980s had
no scaling variable at all.

Add hardware to a rule-based system
and it fails faster.

What is different, in numbers

more frontier training compute, per year, since 2020
less compute needed for the same result, per year
~15×
effective compute per year, compounded

Epoch AI, current to 5 February 2026. A frontier training run now draws tens to hundreds of megawatts. A medium-sized power plant.

Descartes’ first test

Flexible, appropriate use of language

Fallen

It fell so completely that we stopped calling it a test.

Descartes’ second test

General reasoning that transfers
to new situations

Not fallen

His words: an universal instrument that is alike available on every occasion.

ARC-AGI-3, March 2026, on environments no system has seen before.
Humans 100 per cent. Frontier systems 0.51 per cent.

A successful adaptive path and a looping failed path through the same novel grid environment.

Each one was built to stop the last one

2019
Five years, then 87.5%
ARC-AGI-1. Nothing touched it until December 2024.
Mar 2025
Nine months, then 54%
ARC-AGI-2. Built to defeat the systems that had just done it.
Mar 2026
Humans 100. Systems 0.51.
ARC-AGI-3. Launched five months ago.

The people saying this most plainly are the ones building the tests. Their own retrospective, December 2025: “Do we have AGI? Not yet.”

ARC Prize. ARC-AGI-1, o3 at 87.5% on the semi-private set, 20 December 2024. ARC-AGI-2, 54% verified, Poetiq on Gemini 3, December 2025, at $30.57 a problem; the grand prize, 85% at $0.42 a task, is unclaimed. ARC-AGI-3, humans 100%, frontier systems 0.51% at launch.

The case for

Five times the compute a year, compounding with three times the efficiency. A hundred and thirty-one day doubling in how long a system can work unsupervised. And in May, a general reasoning model disproved a conjecture of Erdős from 1946, checked by outside mathematicians.

If a human had written the paper… I would have recommended acceptance without any hesitation.

Timothy Gowers, Fields Medallist

The case against

The same systems score 0.51 where humans score 100. They read an analogue clock about half the time. In the one randomised trial on real work, sixteen experienced developers were 19 per cent slower and believed they were 20 per cent faster. Fewer than one in five American businesses use them at all.

We are near the end of the exponential.

Dario Amodei, February 2026. He means the returns to pure scaling,
and he says it to argue for urgency rather than slowdown.

Bender and Koller, 2020

An octopus taps
the telegraph cable.

Two people are stranded on separate islands, talking through an undersea cable. A hyper-intelligent octopus taps the line. It never sees either island. It learns to predict the replies with great accuracy, then cuts the cable and impersonates one of them. It passes, right up until somebody needs a coconut catapult built. Emily Bender and Alexander Koller: a system trained only on form has no way to learn meaning.

Bender and Koller, “Climbing towards NLU”, ACL 2020, Best Theme Paper. Descartes’ first test, restated.

THREE

What the history is for

Four things it took me too long to learn

In 1958 they said ten years
for a computer to prove
an important new theorem.

It happened in May.

The direction held.
The timing was out by sixty-eight years.
The 2026 forecasts have the same shape.

One

The spectacle is almost never
where the engineering is.

You already knew about the duck. You had not heard of the flute player. That asymmetry is running whenever you choose what to fund or what to worry about.

Two

The record drifts toward drama.

The most quoted prediction in this field was denied by the man it is attributed to. The famous forty million dollar saving has never been the same number twice. The Turk was exposed three years after it had already burned. And the golem of Prague, which feels medieval, is attached to Rabbi Loew in print for the first time in 1834. It is about as old as the railway.

Dekel and Gurley, “How the Golem Came to Prague”, Jewish Quarterly Review 103:2 (2013), dating the Prague attribution to Kohn, Das jüdische Gil Blas (Leipzig, 1834).

Three.

In considering any new subject, there is frequently a tendency, first, to overrate what we find to be already interesting or remarkable; and, secondly, by a sort of natural reaction, to undervalue the true state of the case…

Ada Lovelace, 1843. That is the hype cycle, written down by someone who had seen exactly one machine, and it did not exist yet.

Four. Working is not the same as being used.

70%
of MYCIN’s therapies rated acceptable by a majority of evaluators
55.5%
for the Stanford faculty specialists it was tested against
0
patients. It never entered clinical practice.
MYCIN’s questions feeding a rule network and therapy recommendation, with a gap before the hospital bed.

Blinded evaluation, JAMA 242(12), 1979. Liability. No route into hospital workflow. No clinician willing to type answers to an interrogation.

It beat the specialists
and never reached a patient.

The 1980s died in that gap.
I spend most of my time in it.

Stewart Brand, forty-one years apart

Whole Earth Catalog, Fall 1968

We are as gods and might as well get good at it.

Written against “government, big business, formal education, church”. Brand: “Credit where it’s due: I stole the line.”

Whole Earth Discipline, 2009

We are as gods and HAVE to get good at it.

“Necessity comes from climate change, potentially disastrous for civilization.”

Might as well became have to.

MISQUOTED

“AI is whatever hasn’t been done yet.”

Everybody quotes Larry Tesler saying that. He corrected it himself, on his own website. What he actually said was:

“Intelligence is whatever machines haven’t done yet.”

The first is a joke about a research field. The second is a claim about us.

Descartes wrote down two tests
to prove that he was not a machine.

For three hundred and eighty-nine years,
every test we have written since has pointed the same way.

At the machine.

The imitation game, 2026

You are the interrogator.
Both hidden players are you.

Behind one door, your own thinking. Behind the other, your thinking with one of these systems in the loop. In a study published this February, people solving logic problems with AI scored three points above the norm, and rated themselves four points too high. The Dunning-Kruger effect stopped appearing. Everyone was overconfident. The people with the most AI literacy judged themselves worst of all.

Fernandes et al., “AI makes you smarter but none the wiser”, Computers in Human Behavior 175 (2026). Two studies, roughly 700 participants, internally replicated.

Cogito, ergo sum.

I think, therefore I am.

That is the one claim in this talk
that nobody else can verify for you.
Check it now and then.

Sources

The long history

Descartes, Discourse on the Method (1637)

Riskin, Critical Inquiry 29 (2003)

Grier, When Computers Were Human (2005)

Lovelace, Notes A and G (1843)

Turing, Mind LIX:236 (1950)

Wilks (ed.), Margaret Masterman (2005)

The last fourteen years

Li & Krishna, Dædalus 151:2 (2022)

Krizhevsky, Sutskever & Hinton, NIPS (2012)

Epoch AI, notable models dataset · trends

Nordhaus, J. Economic History 67 (2007)

ARC Prize · METR · US Census BTOS

Bender & Koller, ACL (2020)

What it is for

Idel, Golem (SUNY, 1990)

Dekel & Gurley, JQR 103:2 (2013)

Yu et al., JAMA 242:12 (1979)

Brand, Whole Earth Catalog (1968); Discipline (2009)

Fernandes et al., Computers in Human Behavior 175 (2026)