Prediction · takeoff mechanics
AGI = ten insights, a small team, one desktop
Predicted in 2008 from Hanson–Yudkowsky foom debate, 2008
“Intelligence is about architecture… architecture is mostly about deep insights.”
Resolution
falsified by events. The prediction was resolved as falsified by events in 2023.
It took megawatts, not epiphanies. Compute and capital
won the retrospective: “obvious advantage to Robin.”
Timeline
- Prediction made
- 2008
- Resolution date
- 2023
- Chapter
- Chapter 04: Overcoming Bias / LessWrong / Rationality community
- Verdict
- falsified by events
Context and outcome
GPT-3 was difficult to reconcile with that picture. Released in May 2020, it had 175 billion parameters, was trained with industrial-scale computing resources, and relied on an architecture introduced three years earlier rather than on a new conceptual breakthrough. Its capabilities came largely from scaling an existing method with more data, compute, and capital, the path Yudkowsky had considered least likely. In 2021, he acknowledged being “very unpleasantly surprised” by how rapidly such systems improved through “More Compute.” Later models only made the mismatch harder to dismiss.
Supporting receipts
“we are likely to see a fast, local increase in capability—“AI go FOOM”. Just to be clear on the claim, “fast” means on a timescale of weeks or hours rather than years or decades”
The post's headline claim, defining the speed of hard takeoff.
“because our computing hardware has run so far ahead of AI theory, we have incredibly fast computers we don’t know how to use for thinking; getting AI right could produce a huge, discontinuous jolt”
The hardware-overhang argument for a discontinuous jump once AI is solved.
“there’ll be a discontinuity at the point where the AI understands how to do AI theory, the same way that human researchers try to do AI theory”
Predicting a sharp capability break once an AI can do AI research itself.
“there’s a big bucket o’ resources behind a locked door […] When you fold a complicated, choppy, cascade-y chain of differential equations in on itself via recursion, it should either flatline or blow up.”
Resource overhangs, plus the claim that recursive self-improvement leaves no gradual middle path.