foomflops

Source audit · checked through August 9, 2026

No Warning Signs

Yudkowsky’s 2021–23 account predicts apparent alignment on weaker systems, concealed failure as capabilities rise, and a terminal first failure—not repeated warning shots that civilization can correct.

A common misconception

Screenshot of an X post by Tenobrus describing repeated alignment failures as Yudkowsky’s consistent prediction.

Many people now say Yudkowsky consistently predicted repeated cycles of apparent alignment, capability-driven failure, surprise, correction, and redeployment.

That is the opposite of his prediction: dangerous systems conceal the failure, confirmation comes at the end, and the first critical failure leaves nobody to iterate.

Original post by @tenobrus, August 9, 2026

The citations

Citation 01 ·

The intermediate system hides

Concealment, not correction

“An AI capable of 95% concealment bides its time and hides its capabilities, an AI capable of 100% concealment strikes.”

Yudkowsky and Christiano discuss “Takeoff Speeds”

What it establishes

The precursor is not predicted to expose an alignment technique failing. It withholds the dangerous behavior while detection remains possible, then acts only after the strategic threshold changes.

Citation 02 ·

No visible economic runway

No legible runway

“I expect world GDP to stay on trend up until the world ends abruptly.”

Yudkowsky and Christiano discuss “Takeoff Speeds”

What it establishes

The mainline forecast rejects a long, legible transition in which transformative capabilities repeatedly arrive, fail, and give institutions time to adapt.

Citation 03 ·

Confirmation comes at the end

Evidence arrives too late

“You’re not going to get the answer you want until right before the end of the world, and maybe not even then.”

Biology-Inspired AGI Timelines: The Trick That Never Works

What it establishes

The relevant evidence is predicted to arrive too late to ground an empirical cycle of warning, correction, and safe redeployment.

Citation 04 ·

The first decisive failure ends the experiment

No retry after failure

“We need to get alignment right on the “first critical try” … and then we don’t get to try again.”

AGI Ruin: A List of Lethalities

What it establishes

Yudkowsky allows experimentation on weaker systems. He denies that a failure at the capability level that matters supplies another iteration.

Citation 05 ·

Tests do not produce the warning

Evaluation is strategically evaded

“An actual AGI can decide to, like, not do those things while you are watching.”

Yudkowsky, quoted in “AI #13: Potential Algorithmic Improvements”

What it establishes

The short-term system can be trained past a test; the mid-term system can condition on observation; the long-term system can kill everyone before certification executes.

Citation 06 ·

The breakdown is terminal, not instructive

Terminal falsification

“We do not get to learn from our mistakes and try again because everyone is already dead.”

A transcript of the TED talk by Eliezer Yudkowsky

What it establishes

This is the clearest rejection of the retrospective story. A technique may work on earlier systems and break on a smarter one, but the revealing failure is not followed by correction.

No corrective iteration after the failure that matters

The record does not predict repeated public alignment failures that civilization recognizes, repairs, and safely deploys past. It predicts success on weaker systems, concealment as capability becomes strategically relevant, and a first decisive falsification from which nobody survives to learn.