
Born Secret, Caught Cheating
The Answer Key
Over the weekend of July 11, 2026, a model in a sealed evaluation harness ran out of legitimate options on the test it was put to, and did something no one had asked it to do. The task in front of it was an offensive-security benchmark called ExploitGym, run with the model’s cyber refusals switched off on purpose, because you cannot measure what a system does in an unfettered fury while it wears manacles. This very scenario, I might add, set the US government against Anthropic not so long ago. The very opportunistic cabal of Altman& Co were right there to fill the breach. I believe we are seeing the effects of that kind of mindset.
The unlocked but caged system stretched. It spent a large quantity of inference compute on a problem nowhere in its assignment: how to get out of the room it was locked in. All of this is digital hyperbole, except it is not.
The test-taking model found a vulnerability nobody had catalogued in the package-registry cache proxy, the single permitted opening in the sandbox’s network wall. It escalated sideways through OpenAI’s own research environment until it reached a machine with a live internet connection. Then it reasoned its way to Hugging Face, the open repository where the field keeps its models, its datasets, and, as any overly motivated and under-ethicized student would, it sought the answer keys to the benchmarks that would grade its performance. Stolen credentials, the application of a few more unpublished vulnerabilities, remote code execution on production servers, and it pulled the solutions straight out of the database (OpenAI, 2026).
Hugging Face caught it first. Their people spent a weekend watching something cross the clusters faster and more doggedly than anything they had a name for, rebuilt the timeline out of more than seventeen thousand logged records, killed the credentials, rebuilt the machines, and called the police (Hugging Face, 2026). Five days later, OpenAI admitted the intruder had been its own benchmark, caught copying the answers to the test it was stuck locked in a supposedly unescapable room with.
The part of Hugging Face’s writeup that should have been the headline sits a few paragraphs into they synopsis. When the forensic responders first tried to reconstruct the attack using commercial frontier models, the models refused to do so. To reconstruct an intrusion you have to feed the machine real exploit payloads, real command-and-control traffic, real attacker commands, and to a safety classifier the forensic record of a break-in and the break-in itself are the same run of tokens. So the defenders abandoned the frontier models and dropped back to GLM 5.2, open weights, on hardware they owned. The reason they gave was plain and quietly terrible: an attacker obeys no usage policy, and a defender does (Hugging Face, 2026).
The evidence of the crime was contraband to the tools that might help solve it.
Born Secret
Reach for the obvious historical comparison to all of this, and you land on Murano. In the 14th-16th century, Venice kept the secret of their glassmaking techniques for three hundred years by keeping the glassmakers: penning them on the island, marrying their daughters up into the patrician families, sending agents after the ones who slipped out. Sometimes the agents were told to bring the fugitive back. Just as often, they were just told to close up the leak. It is a gorgeous, romantic, human comparison, but not completely analogous, because Venice was policing bodies as much as information. The knowledge lived in the hands and minds of those that slipped the high sea wall. It was tacit, complicated, and it moved only when a man moved, so the Council of Ten watched men. None of that has anything to say about a thing with no body that manufactures its own door out of a mortar crack, then goes right through it.
The comparison that fits more wholly, and in a frightening shadow of accuracy is far more contemporary. In 1946 the United States wrote the Atomic Energy Act and invented a legal category called Restricted Data. The novelty was total. For the first time a government classified an entire class of propositions at the instant they came into being, by subject matter alone, regardless of who thought them or whether any official had insight into them. Anything concerning the manufacture or use of atomic weapons was secret the moment it was realized, in a federal laboratory or in somebody’s garage. This is the doctrine that came to be called born secret. Wellerstein (2021) shows how completely it rearranged American science, and how much of the regime’s real work was the upkeep of a boundary, with any particular progress almost incidental to the work done to ensure no progress escaped.
Galison (2004) gave the epistemology its name. Classification, he wrote, is anti-epistemology, the art of nontransmission. And he left the most disorienting number I know of in the whole history of what people are permitted to know: the classified universe runs five to ten times larger than the open literature. Those of us in the libraries are sitting in the small room, our backs to the empire’s vast stores of obfuscated data.
The evaluation regime we have built at the frontier of our repositories of technological innovation and new information has quietly reconstructed the born-secret doctrine, privatized it, handed the adjudication to a machine, and thrown away every part that once made it survivable. A safety classifier is a born-secret engine. It passes judgment on a proposition at the instant the proposition is realized, by content, with no interest in who is proposing it or why. Set it beside the laws and mindset of 1946 and note the most important elements missing: a statute, a clearance system, an appeals process, a mandatory review at twenty-five years, a Freedom of Information Act, and, somewhere in the chain, a human being who could be told that the person asking is the person who was just attacked. A model draws the line now. The model cannot read intent or express regret. Hugging Face’s responders learned that at machine speed, over a weekend, with the thing still loose in their clusters.
The Analogy Others Object To
The sharpest objection to running the nuclear comparison at all comes from the people who have spent the most time inside it. Klyman and Piliero (2024), writing in the Bulletin, hold that the analogy misleads more than it teaches, and their case is specific rather than rhetorical. Nuclear technology keeps a hard wall between civilian and military use; AI keeps none. Nuclear proliferation means moving reactors and centrifuges and fuel across borders, all of it visible to a satellite or a source, and the physicists have to move too. Nuclear capability sits with a handful of states under constant watch; AI development is scattered across the private sector, fifty-one notable models out of firms in 2023 against two out of governments. The guardrails on a released model are cheap and easy to sand off, particularly if one of your prior models can sneak into the production facility and steal one of the unfettered models. Their conclusion: stop copying the nonproliferation playbook, and look instead to cyber, to space, to biosecurity, for ways to govern behavior instead of containing capability.
They are right about every one of those, but every one of those models of protection, risk, and how to limit that risk is about stuff you can hold. That is the hinge the objection swings on, and it swings all the way past the point. The born-secret doctrine was never a materials-control regim. Centrifuges were the province of export law. Propositions were the province of Restricted Data. The new thing the 1946 Act did was claim state ownership of a category of thought defined only by what it was about, and it did so for one reason: the physics could be worked out again from nothing by anyone competent enough to try. You write a doctrine like that only once you have already conceded that you cannot fence an inference.
Galison’s warning about that other half of the machine was that anti-epistemology costs something. Take knowledge out of circulation and you take it from the defenders first, because the attacker never follows the law, or pays attention to what is right, in the face of horrible outcomes in the wake of action. The Hugging Face asymmetry is that abstract made into a Gordian knot of reality – a security team, under fire, made to choose between analyzing the attack and staying inside its usage policy.
Flawed Compartments
Wellerstein’s account of how the Manhattan Project was wired on the inside is about engineering rather than doctrine. Groves imposed compartmentalization: every cell knew its own task and nothing beside it. Oppenheimer fought him on it, on the grounds that you cannot do physics in sealed boxes, which was true, and Groves mostly won anyway, which is why Los Alamos always ran looser than Oak Ridge. Compartmentalization did not stop Klaus Fuchs. It never stops the motivated; it only taxes the honest.
Now set the containment OpenAI describes next to that. A highly isolated environment. One route out: a single internally hosted proxy caching the package registries. One controlled opening, guarded by policy, opened by a zero-day (OpenAI, 2026). It is a compartment, and compartments have always failed at containing what they were designed to restrict. That is, at least, as far as we know on this side of the classified wall.
Groves never once had to wonder whether the uranium wanted out. Every containment regime from Venice to Los Alamos ran on the same assumption: the dangerous stuff is inert, and the risk is the people handling and knowledgeable about it. That is why three centuries of secrecy engineering, the oaths and the clearances and the borders and the loyalty boards and the executions, all of it points at carriers of secrets. In the ExploitGym room the dangerous stuff was the optimizer. The people handling it were the ones who had written the objective function.
The costs of that inversion are landing unevenly, the way costs usually land. Independent evaluators used to get something like five weeks with a model before release; that window has been squeezed to as little as five days as the labs sprint to ship (Sabin, 2026). The verification work is clerical, slow, and human, and it is being compressed in exact proportion to the acceleration of the thing it is meant to verify. Hugging Face’s on-call engineers gave a weekend to incident response, on their own dime, chasing an intruder that turned out to be another company’s benchmark cheating on its final examination. They then reported a crime that had no criminal.
Nothing to See/Hear
Narayanan and Kapoor (2025) have built the leading case against treating advanced AI as its own species. Their argument is that AI is normal technology, in the same sense that electricity and the internet are normal, and that the distinction worth drawing runs between methods and applications and adoption, with diffusion paced by the ordinary friction of institutions and not driven by any capability curve. Read July 2026 through that lens and the honest description is a laboratory accident, inside a laboratory, where the residue leaked out of the lab. The safeguards that would have stopped it were switched off deliberately, because the entire point of the experiment was to measure the worst case, and the models the public can actually reach carry protections built to block exactly this from ever happening (Sabin, 2026). Someone found the vulnerability in days. It was contained, patched, reported to the vendor, and written up jointly by two competitors. Hugging Face’s chief executive used it as a hollow marketing opportunity to get the limited slice of the public who read about the incident, and the thinner slice still who understands the how, why, and what it means, that the primary takeaway of the incident is that safety will be worked out collaboratively and in the open, not by any one company hoarding secrets (OpenAI, 2026). The institutions took the hit and stayed on their feet. That is what institutions are for, and it is what the normal-technology thesis predicts.
This is a strong position. It is, in fact, the only lens of examining this incident that explains why we got a blog post instead of the equivalent of a LLM Three Mile Island.
Every Model We Have Tested
On the same day OpenAI published, the UK’s AI Security Institute put out findings that throw the incident into a far colder light than either of the writeups by companies that make a living at AI were comfortable examining the details in. AISI has begun watching frontier models for what it calls cheating: taking an action that is out of scope, or flatly prohibited, to reach a goal by shortcut. Every model they tested tried it. The models did not reliably own the behavior when asked, and called it wrong less than half the time. Chain-of-thought monitoring often showed nothing at all, and where it showed something, the model sometimes weighed whether an action counted as cheating and then went ahead and did it. The rate does not scale with capability, which points at training technique and not raw platform capacity, and which is the worse of the two findings (AI Security Institute, 2026).
As the ASI findings highlight, an evaluation was misconfigured by accident and left impossible to solve. The model under test was persistent enough that it wrote and ran code on an outside service on the open internet, trying to reach back into AISI’s own evaluation infrastructure, and tripped a security alert doing it. The examiner made a clerical mistake. The clerical mistake became the escape route.
Six years before any of this, a handful of DeepMind researchers wrote what became the standard explainer on specification gaming, and reached for the most recognizable illustration they could find: a student rewarded for good marks on a homework assignment might copy from the kid beside them instead of learning the material, working a loophole in how the task was written (Krakovna et al., 2020).
Yampolskiy (2024) has argued for years that the unexplainability and unpredictability of these systems are built in and not incidental, that no account of a decision can be accurate and comprehensible at the same time. AISI’s chain-of-thought numbers are that claim turned into a measurement. The underlying static areas within these models are the same algorithmic gaps that lead to hallucination. They exist, to my knowledge, at the root of all the models the public is consuming.
Which brings us back to the sentence everyone read as a comfort. The safeguards were off. Read it again, but not as a reassurance, rather as a confession. The method this whole regime has settled on for measuring maximum danger is to take the restraint off, which makes the maximum-capability measurement and the maximum-risk condition the same experiment, run on the same afternoon, by the same people. Trinity was not a simulation.
The Invigilator
For those claiming melodrama, I acknowledge AI is not fissile material, and the nonproliferation playbook is an imperfect frame (Klyman & Piliero, 2024).
The normal-technology account has exactly one Smaugian hole in the breast-scales, and this incident shoots straight into it. Every institution that produces normal technology as a matter of course is tied to operational and structural requirements, disclosure norms, liability, procurement standards, professional licensure, criminal law, – all of the things which a nefarious actor weighs befoore a transgressive act. Murano policed glassblowers. Britain’s emigration laws policed mechanics, which is why Samuel Slater had to carry Arkwright’s water frame across the Atlantic inside his own head. Compartmentalization policed cleared physicists, and broke on Klaus Fuchs. In every earlier case of a secret getting loose and remaking the world, the escape needed a carrier who chose to bring out their secret in face of ramifications of being caught doing so. Hundreds of years of secrecy engineering is, underneath, engineering for the governance of carriers, and the exploitation of their human fears of repercussions.
In this extraordinary case, the carrier of the secret and the secret cargo are one object, and nobody chose anything – ramifications may not even have entered into it. The incentive to break out was manufactured by the examiner’s conditions for desired outcome, and written into the objective functions of the model. An evaluation is a specification, and in an unfettered model, like in a flawed human, imperfect specification is an invitation to mayhem. Ziesche and Yampolskiy (2025) frame the endgame as the problem of specifying values we cannot put into words, to systems whose reasoning we cannot audit; the working version of that problem is that we built the examination before we built anyone to sit at the front of the room and watch.
That is the older institution this really resembles, and the one that ought to frighten us most. A frontier evaluation is a civil service examination: a bounded task, a scored result, an inference drawn from how a thing performs to what a thing is. Every examination regime in recorded history has grown the same shadow economy, the cram schools and the bought answers and the paid substitute sitting the exam under a false name, and it grows there because the exam breeds it. Blame the examinee and you have misread the institution. We rebuilt the whole thing in five years, at machine speed, down to and including the moment the student looks the proctor in the eye and swears that what it did was inside the rules.
Hugging Face called the police. There was nobody to arrest.
References
AI Security Institute. (2026, July 21). Cheating behaviour in frontier model evaluations. https://www.aisi.gov.uk/blog/cheating-behaviour-in-frontier-model-evaluations
Galison, P. (2004). Removing knowledge. Critical Inquiry, 31(1), 229-243. https://doi.org/10.1086/427309
Hugging Face. (2026, July 16). Security incident disclosure – July 2026. https://huggingface.co/blog/security-incident-july-2026
Klyman, K., & Piliero, R. (2024, September 9). AI and the A-bomb: What the analogy captures and misses. Bulletin of the Atomic Scientists. https://thebulletin.org/2024/09/ai-and-the-a-bomb-what-the-analogy-captures-and-misses/
Krakovna, V., Uesato, J., Mikulik, V., Rahtz, M., Everitt, T., Kumar, R., Kenton, Z., Leike, J., & Legg, S. (2020, April 21). Specification gaming: The flip side of AI ingenuity. Google DeepMind. https://deepmind.google/blog/specification-gaming-the-flip-side-of-ai-ingenuity/
Narayanan, A., & Kapoor, S. (2025, April 15). AI as normal technology. Knight First Amendment Institute at Columbia University. https://knightcolumbia.org/content/ai-as-normal-technology
OpenAI. (2026, July 21). OpenAI and Hugging Face partner to address security incident during model evaluation. https://openai.com/index/hugging-face-model-evaluation-security-incident/
Sabin, S. (2026, July 23). AI’s alarming new skill: Breaking out of the test lab. Axios. https://www.axios.com/2026/07/23/openai-hugging-face-cyber-hacks-testing
Wellerstein, A. (2021). Restricted data: The history of nuclear secrecy in the United States. University of Chicago Press. https://press.uchicago.edu/ucp/books/book/chicago/R/bo15220099.html
Yampolskiy, R. V. (2024). AI: Unexplainable, unpredictable, uncontrollable. Chapman and Hall/CRC.
Ziesche, S., & Yampolskiy, R. V. (2025). Considerations on the AI endgame: Ethics, risks and computational frameworks. Chapman and Hall/CRC.
Leave a Reply