foomflops

Prediction · neural nets & capabilities

AI takes IMO gold by 2025 (≥16%)

Predicted in 2022 from IMO bet with Paul Christiano, Feb 2022

“I'll stand by a >16% probability of the technical capability existing by end of 2025.”

Resolution

came true. The prediction was resolved as came true in 2025.

DeepMind and OpenAI both hit gold at IMO 2025.

The control group: his falsifiable near-term bet won.

Timeline

Prediction made
2022
Resolution date
2025
Chapter
Chapter 06: MIRI: Advocacy
Verdict
came true

Context and outcome

Christiano put the chance of an AI earning IMO gold by the end of 2025 below 8 percent; Yudkowsky put it at 16 percent. In July 2025, Gemini Deep Think received an officially certified gold-level score of 35 out of 42, and OpenAI reported the same score for its own model.

This one stays in. Yudkowsky got the main prediction right, while Christiano’s narrower call also held because neither system solved Problem 6.

Supporting receipts

  1. “My probability is at least 16% [on the IMO grand challenge falling], though I'd have to think more and Look into Things, and maybe ask for such sad little metrics as are available…”

    Yudkowsky's stated odds on an AI getting IMO gold by 2025, against Christiano's roughly 8%.

    Yudkowsky, quoted in Christiano's bet post, LessWrong (Feb 2022)

  2. “…the sort of foolish little real-world obstacle which can prevent a proposition like this from being judged true even where the technical capability exists.”

    Yudkowsky's caveat that an open-sourcing requirement could block resolution; he stood by >16% on the capability existing by end of 2025.

    Yudkowsky, quoted in Christiano's bet post, LessWrong (Feb 2022)

  3. “Maybe I'd move from a 30% chance of hard takeoff to a 50% chance of hard takeoff. If Eliezer wins, he gets 1 bit of epistemic credit.”

    Christiano pre-registering how much he would update toward hard takeoff if Yudkowsky won.

    Paul Christiano, LessWrong (Feb 2022)

  4. “Google and OpenAI have LLMs with gold medal performances, each scoring exactly the threshold of 35/42… The results were not super clear cut if you look at the details… but a gold medal was still achieved.”

    The July 2025 resolution: two labs hit the IMO gold threshold, settling the bet Yudkowsky's way.

    Zvi Mowshowitz, LessWrong (Jul 2025)

Sources