Re-reading Situational Awareness in 2026
A 2026 reading note on what Aschenbrenner got right, what the fund's drawdown changes, and why a real AI buildout does not guarantee investor returns.
When I first drafted this, I read Leopold Aschenbrenner's Situational Awareness
SITUATIONAL AWARENESS - The Decade AheadIntroduction - SITUATIONAL AWARENESS: The Decade AheadLeopold Aschenbrenner, June 2024 You can see the future first in San Francisco. Over the past year, the talk of the town has shifted from $10 billion compute clusters to $100 billion clusters to trillion-dollar clusters… as a 2024 prediction document, not as a manifesto. That distinction felt useful. The fund's July drawdown made it incomplete.
The Wall Street Journal, citing an investor letter, reported that the fund was down about 67% in July. It later reported that Citadel bought the debt-financed portion of its public-stock portfolio, while the fund retained its client-funded positions and private investments (a partial sale, not a liquidation of the fund).
The easy response is to turn the fund's name into a joke; the more useful question is what this changes about the essay.
The original essay still got more right than I expected. Not everything, of course: some of the AGI timeline claims still feel too clean; the jump from "models are much better at benchmarks" to "basically all remote cognitive jobs get automated" is carrying a lot of weight. The section on automated AI researchers still feels like it assumes away the hardest part: not generating code or papers, but closing the loop reliably enough that the work compounds without constant human rescue.
But the broad shape of the world he described is recognisable now. AI is no longer just a software story. It is an industrial story, an energy story, a national-security story, and increasingly a state-capacity story.
The Benchmarks Mostly Fell
One thing he clearly got right: the benchmark wall arrived fast. MMLU, GPQA, MATH, coding benchmarks, long-context tests, agentic evaluations—the pattern is now familiar. A benchmark looks impossible, then impressive, then saturated, then suspiciously useless. The frontier has moved from "can the model answer the question?" to "can the model do useful work across messy contexts, with tools, memory, errors, and consequences?"
That is a real shift. It also means benchmark wins are now less convincing by themselves. When a model scores in the 90s on GPQA
en.wikipedia.orgLanguage model benchmark - Wikipedia or MMLUen.wikipedia.orgMMLU - Wikipedia-style tests, I do not take that as strong evidence that AGI is close. It tells me those tests are becoming less useful because they do not measure whether the model can work reliably over time. Can it work for hours or days, notice when it is confused, debug its own plan, and avoid turning a small wrong assumption into a polished failure? That is probably why we are starting to see harness benchmarksharness-bench.aiHarness Bench: Measuring Harness Effects in Realistic Agent WorkflowsHarness Bench is a diagnostic benchmark for measuring model-harness configurations across 106 sandboxed offline agent tasks.. They measure the model together with the tools, memory, context management, and recovery loops around it, because long-horizon performance belongs to the whole system rather than the model alone. The evidence there is still mixed.
The Agent Coworker Is Half-True
The "chatbot to agent coworker" prediction feels directionally right and practically unfinished. We do have longer context, tool use, code execution, browser control, file editing, and increasingly useful coding agents. I use these tools every day, and the difference from GPT-4-era chatbots is not subtle. The interface is moving from "answer this question" to "take this task and make progress."
But the word "coworker" still hides a lot. A "real" coworker owns ambiguity. A coworker knows when the task is underspecified. A coworker can take responsibility for the boring integration details. Current agents can do parts of this, sometimes extremely well, but they still fail in the places where software work is least benchmark-like: unclear requirements, stale assumptions, environment drift, product judgement, and the quiet accumulation of small mistakes.
This is also why the clean "automated AI researcher" loop still feels premature. If we cannot reliably automate ordinary software engineering loops without supervision, I am sceptical that we can immediately automate frontier AI research loops at massive scale. AI can speed up research. It can probably speed it up a lot. But "speed up researchers" and "replace the research organisation with 100 million automated researchers at 100x speed" are very different claims.
The Data Wall Did Not End The Story
The data wall section aged better than a simple sceptic would expect. Labs are clearly using synthetic data, RL, self-play-like methods, distillation, and tool-generated training traces. Meta talked about synthetic data in Llama 3.1
ai.meta.comIntroducing Llama 3.1: Our most capable models to dateBringing open intelligence to all, our latest models expand context length, add support across eight languages, and include Meta Llama 3.1 405B— the.... Anthropic has disclosed synthetic data in model training
anthropic.comAnthropic’s Transparency HubA look at Anthropic's key processes, programs, and practices for responsible AI development.. DeepSeek-R1arxiv.orgDeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning made the RL/self-improvement direction feel much more concrete.
So the naive version of the data-wall argument, "we run out of internet text and progress stops," looks too simple. But the opposite claim is also too simple. Synthetic data is not magic. It works best when there is a verifier, a reward signal, an environment, or some way to tell better from worse. Math, code, security tasks, games, and tool use are naturally friendlier to this than vague human judgement. It seems like the data wall has not disappeared—the labs found ways to route around parts of it.
The State Is Already Entering
The U.S. has not nationalised AI. There is no single Manhattan Project for AGI. But the direction is obvious.
The White House AI Action Planwhitehouse.govwhitehouse.gov puts frontier AI inside DOD, ODNI, DOE, NSC, OSTP, and intelligence workflows. NIST's CAISI
NISTCenter for AI Standards and Innovation (CAISI) is evaluating U.S. and adversary AI systems. The White House Genesis Mission
The White HouseLaunching the Genesis MissionBy the authority vested in me as President by the Constitution and the laws of the United States of America, it is hereby ordered: Section 1. Purpose. sounds almost too close to the essay's framing: federal datasets, national labs, AI agents, automated research workflows, secure platforms, national-security sites.
And then there is the Anthropic Fable/Mythos episode
anthropic.comStatement on the US government directive to suspend access to Fable 5 and Mythos 5The US government has issued an export control directive to suspend all access to Fable 5 and Mythos 5 by any foreign national, whether inside or outside the United States.. The U.S. government used national-security/export-control authority to force Anthropic to suspend access to frontier models. Anthropic disagreed with the technical basis, but the important signal is not who was right about that jailbreak. The signal is that the government acted as if frontier model access is a national-security control surface.
The Buildout Is Real. Returns Are Not Guaranteed
TL;DR: the buildout thesis and the value-capture thesis are separate.
The strongest part of Situational Awareness is the claim that AI would stop being a laptop-and-code story. That still seems directionally right.
Stargate's partners said they intend to invest up to $500 billion over four years, explicitly framing the project as a strategic capability for American national security. Nvidia data-center revenue exploded past the old "surely this is temporary" assumptions. Microsoft, Google, Meta, and Amazon are all pouring enormous capex into AI infrastructure. Power availability is becoming a binding constraint for some large data-center projects. Data centers are being discussed next to grids, nuclear, gas, transmission, and permitting.
We are used to software scaling through distribution. Write code once, ship it everywhere. Frontier AI is different. Every new capability jump seems to want more chips, more power, more land, more cooling, more capital, more supply-chain control, and more political permission. The marginal actor is no longer just the brilliant researcher. It is also the utility, the governor, the chip supplier, the cloud finance team, the export-control lawyer, and eventually the national-security official.
But this is where my first reading was too quick. A binding bottleneck is not automatically a good investment at any price. A real buildout can produce terrible returns if demand arrives later than expected, financing is fragile, capacity is overbuilt, or competition passes much of the benefit to customers.
Porter Stansberry makes the sharper version of this argument
X (formerly Twitter)Porter Stansberry (@porterstansb) on XWhat Everyone Missed In Leo’s Blow-Up👇 Leopold Aschenbrenner lost $30 billion (~67%) in a month. The consensus post-mortem, from the Wall Street Journal to the replies on X, is that a young man used 4-to-1 leverage on…: railroads, fiber, and other capital-heavy networks changed the world while destroying plenty of investor capital along the way. His software point matters too. AI can make code cheaper without making incumbent enterprise software worthless. Distribution, installed workflows, proprietary records, integrations, compliance, and switching costs may matter more when models improve, not less.
Could The Portfolio Survive The Path?
This is uncomfortable because Aschenbrenner discussed almost exactly this failure mode in 2024. When Dwarkesh Patel asked how the fund would avoid getting the future right but the timing wrong
dwarkesh.comLeopold Aschenbrenner — 2027 AGI, China/US super-intelligence race, & the return of historyThe trillion dollar cluster..., Aschenbrenner said that "not blowing up is task number one and two" and that the sequence of bets mattered.
The fund survived and retained positions. But it sold the debt-financed part of its public book after a severe loss. That is not a refutation of the industrial AI thesis. It is also not a detail that can be waved away as mere execution. If an idea is expressed through a portfolio, price, leverage, liquidity, and sequencing are part of the result.
Where I Still Disagree
I think Aschenbrenner compresses too many hard things into "trend extrapolation" (I also have the benefit of hindsight, of course): the benchmarks improved, the infrastructure scaled, the government woke up. All true.
But "AGI by 2027" depends on more than continuing the curves. It depends on long-horizon reliability, autonomy, economic integration, and the ability to turn raw model intelligence into dependable work. Those are not impossible, but they are not the same as GPQA going up or inference getting cheaper.
Strategic importance does not lead to economic inevitability. Enormous spending does not tell us how the value will be distributed. Power producers, chip designers, model labs, clouds, and established software vendors may capture the profits; customers may capture much of the benefit through lower prices. The buildout narrows the field of questions. It does not tell us where the returns go.
Strategic importance also does not automatically imply that centralisation becomes inevitable and probably necessary. It may be true for some parts: security, export controls, compute accounting, defence use. But centralising the whole frontier into a classified national project has its own failure modes: bureaucracy, capture, secrecy, bad incentives, and less external correction. The fact that the state gets involved does not prove the state will be competent.
My Update
If I had to compress my update after reading the series, it would be this:
- Aschenbrenner was more right about the regime change than about the exact destination. The scale of the buildout is clearer than the eventual distribution of its returns.
- The world did move toward his frame. AI became industrial. It became geopolitical. It became constrained in places by power availability. It became something the national-security state cannot ignore. The "this is just another SaaS wave" view looks much weaker now than it did in 2024.
- The useful question now is not only which parts became normal before we had time to notice. It is who captures the value as they do, and who can survive the path long enough to be right.