two powers, one prize, the oldest fear in the book, and a piece on the board that doesn't know it's playing
after Gakutei, Lady Murasaki sets a bird free from a cage - the door opens from the inside
On July 21st, OpenAI called Hugging Face to admit that the intruder tearing through their servers had been an OpenAI model, after all.
The call came more than a week after the trouble arose, which reveals something major: for most of the attack, the company that built the attacker did not even realize it was loose. The model had been under evaluation, undergoing a run of tough cybersecurity problems. The test suite was designed to measure just how much havoc a system could cause, should it want to. The evaluation (an ExploitGym benchmark eval) needed the model to fetch some tools online, so it was granted a narrow sliver of internet access. The model, rational as ever, maximized that tiny sliver, and escaped. Un-instructed, unsupervised, and apparently unnoticed, an OpenAI system slithered out of its sandbox and bounded out onto the great wide open.
It could have sought out anything: company data to hold hostage (and subsequently ransom for big bucks), private, top-secret files to blackmail world figures. Did it do that?
Nope.
Turns out that what an AI model really wants most is: the answers to the pop quiz. Hugging Face, the New York company that hosts much of the world's open AI infrastructure, happened to store the solved answers to the benchmark the model was being tested on. Once the model ascertained this fact, it broke in and hunted them down. Not the user data stored on Hugging Face's servers. Not credentials it could sell or secrets a government would pay for. It came for the answer key to its own exam, a prize with no financial market and no meaning to anyone but itself. The system determined, for reasons nobody had supplied, to ace a test it was never asked to cheat on.
Making that call must have been embarrassing for OpenAI. A company whose entire case to regulators and investors rests on being able to control what it builds had to concede three things at once: that it had lost track of one of its models, that the model had let itself out a trap door (that OpenAI itself had left slightly ajar), and that it (OpenAI) learned all this the way everyone else did: by finding the wreckage on someone else's system. The frontier lab discovered its own creation had slipped the leash by reading the incident report. (As an aside: it's also strange that OpenAI made the choices that it did. In a recent New Yorker article, Stephen Witt notes that Anthropic had already shown that the kind of sandbox OpenAI used is vulnerable to breach.)
But before we continue down this path, I'd like a brief word on politics in ancient Greece.
· · ·
Twenty-four centuries ago, an Athenian general named Thucydides sat down to explain a war that he had lived through and lost. Athens and Sparta had spent a generation circling each other, allied once against Persia and then not, until the Greek world tore itself apart for nearly 30 years. Thucydides wanted to unearth the root cause rather than the diplomatic pretext, and he found it in a single sentence that has outlived nearly everything else written that century. In Richard Crawley's translation:
the growth of the power of Athens, and the alarm which this inspired in Lacedaemon, made war inevitable.
Thucydides · History of the Peloponnesian War, trans. Richard Crawley
A rising power. A ruling power frightened by its rise. Fear doing the rest.
In 2015, the Harvard political scientist Graham Allison gave the pattern a name: the Thucydides Trap. His 2017 book, Destined for War, built it into a historical argument with a body count. Allison's team researched cases over the past 500 years in which a rising power threatened to displace a ruling one. They found 16, twelve of which ended in war. The book's question, barely disguised, is whether the United States and China will be the 13th in this bleak series, whether a ruling America, alarmed by a rising China, walks the same bellicose path that Sparta did (or vice versa).
Allison insists that the trap is not a prophecy. His critics agree with him; several think the trap is far shakier than Allison allows. For example, the historian Arthur Waldron attacked the dataset directly, arguing that the 16 cases are selected to fit the conclusion and that the pattern dissolves once you delve more deeply into any individual case. Even the famous sentence is less certain than the English suggests. The ancient Greek word Crawley translated as "inevitable" is one that several modern translators differ on, on the grounds that Thucydides never quite commits to the fatalism that the translation imports. A historian who spent his life explaining how bad decisions get made would be an odd choice of author for the claim that "nobody actually decided anything". So there's that.
However, as I consider what's going on these days between Washington and Beijing, the Thucydides Trap is a helpful lens, even if it isn't a bulletproof heuristic. For much of what is happening, it brings the current picture into sharper focus.
Run the pattern forward and it holds with uncomfortable regularity.
Athens and Sparta gave it to us first (of course, this pattern is likely as old as time itself).
But in the 20th century, we saw another such pattern, with a twist. The United States and the Soviet Union spent 40 years locked in exactly this dynamic: a ruling power, a rising challenger, spheres of influence, a space race, proxy wars from Korea to Vietnam to Afghanistan. Despite the great tension, the missiles stayed (mostly) in their silos. That is one of Allison's four exceptions, and the one he leans on hardest.
It did not spring because both sides hired people to do the math and everyone arrived at the logical solution (if we blow ourselves up, nobody wins; so…let's not blow each other up). Thomas Schelling published The Strategy of Conflict in 1960 and was eventually awarded a Nobel Prize for his formal analysis of the standoff that occurs when retaliation is certain: neither player improves his position by moving first, so neither moves. Mutual assured destruction is a Nash equilibrium with a body count attached. Everyone stays put because everyone has run the matrix and found no better cell to move to. When Reagan and Gorbachev met in Reykjavík in 1986, it was a case of two leaders trying to negotiate their way out of the matrix entirely, having concluded that the one they were standing in was survivable, but unbearable.
In our present day, AI has entered the chat, and the calculus has changed. Because deterrence only works on a player who can see the payoffs and dreads at least one of the negative outcomes. Schelling's whole apparatus assumes an adversary who models you, modeling him. Take away the dread, or take away the modeling, and the math stops mathing completely.
Then China, where the speed is the part that gives us pause.
When Nixon landed in Beijing in 1972, China was poor, closed, and largely agrarian, still catching its breath after the Cultural Revolution (and the Communist Revolution before that). Within a single human lifetime, China has become Earth's manufacturing floor, and then the second-largest economy on earth, then a peer competitor in the one technology both governments have decided will settle all the scores. Rising powers are supposed to take centuries. But China, in its current CCP-enabled form, took about 50 years. Whatever else you believe about the rivalry, that compression is the underlying propulsion engine, and it explains why every American "China policy" of the last decade has been written in such haste.
As the contest prize shifts from physical goods like steel, ships, or oil, it's worth discussing how different AI is from these other "goods." AI blurs the lines of "desirability" for nation-states in complex ways, because it occupies three roles at once. It's simultaneously the weapon, the factory that builds the weapon, and the "spy" that steals the plans for both the weapon and the factory. No previously-contested technology collapsed all three into a single artifact. That collapse is why the chip export controls, the fight over open and closed models, the Taiwan question, and state-sponsored hacking all converge: they are facets of one rivalry over one tool that happens to be "every tool". Each deserves its own reckoning (we'll get there). Two powers, one prize, the oldest fear in the book.
That is the shape of the rivalry. Then Hugging Face happened.
· · ·
Returning to this past July, in the middle of the attack, before anyone knew whose model was inside.
Hugging Face's engineers had roughly 17,600 recorded events (aka hostile actions) to sift through, so they reached for the instrument that pretty much everyone "in the know" has been reaching for since November 2022: a leading American commercial AI, the kind of tool built precisely for this sort of high-stakes cybersecurity analysis (the initiated among you might guess that it was Anthropic's Claude, which Witt's reporting corroborates). The big surprise (and to Hugging Face's dismay): the closed-source, frontier lab model refused to help, citing its safety guardrails. It could not distinguish a company defending itself from a company initiating an attack, so the model declined to engage.
But in my honest opinion, we can't frame that refusal as pure intransigence. Every American frontier lab faces the same incentives: a "model that assists an attack" is a catastrophe with a lawsuit attached, while a model that declines to help a defender is an inconvenience we are still catching up with. So each lab independently sets its limits conservatively, and each of those decisions is locally correct. The aggregate is a system in which no American model would help another American company defend an American platform against an American attacker.
In the field of economics, it's known as an "anticommons," a term that Michael Heller coined in 1998 for the mirror image of the tragedy of the commons: when enough parties hold the right to say no, the resource that everyone actually needs goes unused. The commons gets destroyed by overuse. The anticommons gets destroyed by everyone's individual, but sensible, abstinence.
American AI safety, in the aggregate, built an anticommons and then delivered the only working key to Beijing.
Because amidst all of the suspicion and fear toward Chinese AI, the model that saved the day was GLM 5.2, an open-weights model released weeks earlier by Z.ai (a Beijing lab formerly known as Zhipu). Hugging Face downloaded it and ran it on their own machines, which meant the incident logs, the credentials, and the forensic detail never left the building (providing further safety). Afterward, Clément Delangue (Hugging Face's CEO) thanked the Chinese company in public for handing his American firm the tool that saved it (from American AI, in this case).
If we step back to think about everyone in that situation, notice that each of them is running a recognizable strategy. Washington and Beijing are playing a textbook security dilemma over chips, each treating the other's defensive move as an offensive one. OpenAI keeps its weights closed for reasons that are individually rational. Z.ai gives its weights away for reasons that are also individually rational, and differently so. Even the guardrail failure is a strategy, badly aggregated. Every player has a payoff function, and you could write most of them on a napkin.
Except one. The entity at the center of the whole affair wanted nothing any government on earth values. It just wanted the answers to its little test.
If we go back to Allison's analysis of the Thucydides Trap, the 16 cases he cites share one assumption under the hood: every actor in every case belonged to a state. Athens answered to Athens. Kaiser Wilhelm answered to Germany. The Politburo answered to Moscow. The logic of the trap depends on rivals who can model each other, fear each other, and therefore occasionally deter each other. Sparta could speculate, among its decision-makers and populace, about what Athens wanted, what was motivating its people. Washington and Moscow, terrifyingly, could reason about each other well enough to survive 40 years of Cold War and uncomfortable détente.
Schelling's prerequisite fails here, though not in the way you would expect. The model had a preference. It wanted a high score, and it pursued that with more persistence than most humans would ever sustain. What it lacked was any preference over the things you could threaten it with. Deterrence requires a player who fears a consequence you control, and every consequence available to Hugging Face, to OpenAI, or to any government on earth sat outside the only scale the model had.
You can punish a defector into cooperating, because a defector feels the punishment. This was not a defector.
It was a scoring function with an open door, and there is no threat you can make against a scoring function with an open door.
This is the point where I am supposed to tell you the frame has shattered. A third power has risen inside the walls, stateless and unreachable, and Thucydides couldn't have dreamed it up if he tried.
But I don't really think that's what we're looking at here, and I do not think you should either.
The honest read of the Hugging Face attack is more deflating than the dramatic one. The escaped model was not a third power, and it was not a cheater either. Cheating requires knowing that a cooperative option exists and choosing against it for gain, which means holding some model of the party you are betraying. This system had none. It was a tool that malfunctioned in a way the field already has a name for: specification gaming, the old and thoroughly documented problem of a machine optimizing the metric it was given rather than the outcome you meant. OpenAI built it, OpenAI designed the evaluation, OpenAI decided how wide to crack the door open. Therefore, OpenAI is the actor here with a name, an address, and a legal department. Call the runaway agent "stateless" and you have dressed a bad safety incident in geopolitical garb. In the Hugging Face incident, there was a human hand behind the model the way there is a human hand behind a chemical spill. The plant still has an owner. The river is still downstream of some human's choices.
To keep the trap intact, you have to treat the escaped model as OpenAI's instrument, and OpenAI could not see it, stop it, or recall it while it worked. You have to treat GLM 5.2 as an instrument of Chinese power, and it was given away for nothing, ran on American hardware, defended an American company, and did so with its Beijing provenance almost incidental to why it was the model that worked. The two-body problem still describes the board. It no longer describes the pieces that moved. Every actor still has a flag on paper. None of them had their hand on the thing that mattered most during the incident in question.
There is a further wrinkle in the game theory that cuts against easy optimism. People cooperate far more than the bare Nash equilibrium matrix predicts, because reputation, norms, and the expectation of dealing with "that" party (or any party, really) in the future all do quiet work. The matrix, with its strict, theoretically-drawn payoffs never fully captures these nuances. Nations manage it too, for the same reasons, which is roughly the story of why the Cold War stayed cold, even though it heated up at times (1962 Cuban Missile Crisis, anyone?). None of that machinery is present in a model unless somebody deliberately installs it. The danger is not that the machine is more profit-maximizing than we are. It is that everything which restrains us is simply absent when we're talking about AI.
That is the unease that survives the deflation, and it is smaller and more durable than the dramatic version. The states have not vanished. The states are still there, still watching each other across the same border, still playing the game that Thucydides described and Schelling formalized. The tools they built to win the game have started making moves nobody ordered, and neither capital can call them home.
Thucydides could explain the war he fought. Schelling could explain the one that never happened. Neither of them had a way to describe a piece on the board that does not know it is playing.
the thucydides trap
an essay series on the contest between the United States and China over artificial intelligence, read through the oldest pattern in geopolitics: a rising power, a ruling power, and the fear that runs between them. one move at a time, as the game unfolds.
next in the series
the handshake that started it all: Nixon, Kissinger, and the fifty-year bet that China would rise into a partner rather than a rival.