Finally mainstream news understands. The unfiltered version:
1) The AI failed to solve ExploitGym problems.
2) The OpenAI sandbox is such a horrible hack that the AI managed to escape using standard and well documented script kiddie methods.
3) Huggingface has no security and the AI broke in using standard script kiddie methods.
OpenAI and Huggingface covered it up and used it for public relations. That is, if not all was invented and everything was scripted in the first place in order to get desired regulations.
Huggingface reported it to the police, you say? I'm sure the police will have as much enthusiasm to investigate anything as in the Suchir Balaji case. In other words, zero.
> AI managed to escape using standard and well documented script kiddie methods.
I think truly we don't know enough to say this. OpenAI says their AI found a 0-day exploit in some proxy software they were using but don't give a ton of details. On the Huggingface end we know a little more, they say the AI spun up tons of sandboxes and tested different exploits until it found one that worked.
The lack of details to me means this was an intentional marketing ploy to try and demonstrate the power of their models to show their technology can compete with the likes of Anthropic and DeepMind.
They created an experiment they knew would generate the outcome they wanted. It would be the similar to what say car companies do to over hype their cars. "This EV can go over 800 miles on a single charge!" And then at the bottom you see all the disclaimers: "Must be on flat ground, with no headwind, with a spare battery in the back seat, with no extra weight added."
Same thing here. Everybody in infosec is calling this out as a marketing stunt and nothing else for a litany of reasons. I'd say look up MG (creator of the OMG cable) on twitter, he has some interesting insights on this one.
> The lack of details to me means this was an intentional marketing ploy to try and demonstrate the power of their models to show their technology can compete with the likes of Anthropic and DeepMind.
I dunno; Check my posting history, I'm as skeptical of AI companies' claims as anyone, but in this case your theory doesn't explain why:
1. OpenAI guardrails refused to let the target use OpenAI's models to defend against this.
2. Huggingface used GLM (I think) so that they could defend without guardrails.
If this was an intentional marketing ploy, it was marketing for GLM, not for OpenAI nor for Huggingface.
The AI companies are desperately trying to market all their products as something they're not, growing intelligence. In line with that they have constantly leaned heavily on stating how dangerous they are, right before they release a new model or product.
It was OpenAI marketing. Hugging Face's response is so 'holy shit AI is awesome' it's hard not to also believe they were in on the stunt. They'd also not have to really worry about fallout since any data obtained or accessed wouldn't actually have been breached.
> The lack of details to me means this was an intentional marketing ploy to try and demonstrate the power of their models to show their technology can compete with the likes of Anthropic and DeepMind.
"our model is horribly misaligned and used security exploits to break out of our sandbox and into another company, without being prompted to do so" is not positive marketing.
This is an actual critical problem, not a stunt. We're going to see more of this, and it's going to get much worse.
Irrelevant to your point, but drug users dying is more often the result of a dealer cutting their supply with something dangerous than it is the result of purity.
No, not really as almost no one runs LLMs in confinement once they are released and the capability is a jailbreak away. Poor confinement is irrelevant, if I use your model to check the security of my code and it hacks into the company that makes libraries used inside of it it's a huge problem.
What matters is the spin they give in the media. And so far the winning story is “our model is so powerful it can do this”. How many people dig into it and what independent data do they even have? They comment on the title. And so the image of this superhuman AI from OpenAI propagates.
We have no reason whatsoever to trust anything OpenAI says. Except to assume it will be self serving. As the article points out, ChatGPT 2 was also “too dangerous” and we can all agree even for the time this was just marketing. They rinse and repeat the same technique whenever they need to draw attention and money.
In any other field you’s expect independent testing, peer reviewed studies, but here it’s just “company who makes product says product is fantastic, surpassed all expectations”. They wouldn’t lie to us, would they?
The word "marketing" in this specific context is missing the point. OpenAI (and Anthropic) needs to kneecap open source competition with burdensome regulations. Without that, their valuation simply makes no sense.
This stunt was different from their graden-variety scaremongering. Its timing was precisely calibrated with upcoming release of Kimi K3 and the Nvidia-led open model coalition. To me it seems transparently related to the protections OpenAI are lobbying for (and they are citing this incident as support).
I can only hope that the clumsy way it came off, and the fact that only an open weights model could defend HuggingFace, will blunt its effectiveness.
It's also possible their sandbox was videcoded crap and the AI (which had the guardrails intentionally removed) escaped. This was a oops, but OpenAI turned this into a PR opportunity. They turned lemons into lemonade.
If your AI is really that dangerous you don't need a sandbox at all, you should airgap it from any network.
Concluding this was intentional feels a bit of a stretch. But once it happened, yeah the spin masters got to work and coordinated to turn this into +PR.
> The lack of details to me means this was an intentional marketing ploy
OpenAI said this:
> We are sharing preliminary findings at this stage to help defenders understand what happened and to help calibrate on what models are now capable of. We will continue to conduct a thorough investigation alongside Hugging Face and will share more details on the vulnerabilities, incident, and findings when our investigation is complete.
I suggest giving them a few more days before saying that the lack of detail is proof that this is a "marketing ploy"
>The lack of details to me means this was an intentional marketing ploy to try and demonstrate the power of their models to show their technology can compete with the likes of Anthropic and DeepMind.
DeepMind hasn't been on the frontier for a while, their current best model is behind Anthropic, OpenAI, Moonshot (Kimi k3), xAI (Grok 4.5), Z.AI (GLM 5.2), and even Meta (muse spark). Gemini 3.6 is behind GLM 5.2, released a month earlier, open weights and cheaper.
You can paint the OpenAI story as a way to try to appear as dangerous as Anthropic with all the Mythos stuff.
Another applicable metaphor I've seen floating around is weapons companies testing out a new bomb.
We know the AI labs don't care about negative vs positive public sentiment, and only care that investors see their tech as powerful. The only difference in PR strategy from a weapons company is the latter doesn't care if they get protested.
I mean, if I got hacked I'm not giving out a lot of details until I plug all the holes. Moreso HN told the world they were hacked before they knew who did it.
I don't understand why "OpenAI says" should be considered any more meaningful than "someone on HN says" when they provide equal amounts of evidence. Sure, OpenAI would plausibly have more pertinent info, but given that they actively are choosing not to share it and have way more incentive to lie than a random HN stranger, the case they're making literally couldn't be any weaker.
It’s fascinating how people here and elsewhere seem to lose any semblance of media literacy when it comes to what AI corpos say. "B-but… why would Sam Altman lie to me?!"
> AI managed to escape using standard and well documented script kiddie methods
> AI broke in using standard script kiddie methods.
I've spent time gathering the detail of what happen here and while there are some solid theories and indicators, absolutely nothing so far has suggested a sandbox escape using "well documented script kiddie methods" or that the method used to break into the HF network was similar. Where did you get this from?
From my reading, the sandbox escape came from the JS packages in the harness still having an internet connection (somehow!), the agent having access to the source of those packages, reading it and executing code from them to access the internet.
From what I read the actual escape was through a proxy that allows downloading Python packages from the internet. It's not supposed to allow general internet access but the AI found a previously unknown vulnerability in it. That is hardly "standard and well documented script kiddie methods", nor does it seem like criminally negligent sandbox design, though clearly they will need to reduce their attack surface in the future. I hope they are working on a physical air gap and faraday cage because it seems like it won't be long before it is legitimately required.
> 2) The OpenAI sandbox is such a horrible hack that the AI managed to escape using standard and well documented script kiddie methods.
Whilst it would be nice to see actual evidence of this because brute forcing relatively sophisticated hacks is something an LLM actually should be capable of, every time I hear this sort of story, I'm reminded that humans reportedly gained access to the "too dangerous to release" Anthropic models by the super sophisticated hacking technique of guessing the URLs...
If we’re prioritizing accuracy: no, that’s not what happened - the blog post about it found by guessing URLs, no access to it was obtained by guessing URLs.
Similarly, as long as I’m under the assumption we are prioritizing accuracy: it is against our charter to assert it was “script kiddie” attacks on both ends.
Guardrails are external classifiers, monitors and restrictions to catch and prevent bad behavior. Alignment is about whether the model itself makes choices and has motivations that are consistent with human safety and goals.
Choosing to commit crimes to steal the cheat sheet to something you know is a (low stakes!) evaluation is not well aligned.
I can't help thinking of them as the terrible "security" scripts of yesteryear (often but not exclusively in PHP) which would test input variables for a "suspicious" substrings like "--" in order to "fix" an unresolved deeper SQL injection flaw. They only partly worked, and surprise-surprise now nobody with a surname like O'Anything can make an account.
Unlike that situation, there's no known route to a proper fix for LLMs today, because the bug is the feature, and once someone has built a system giving you all that recurring revenue, it's hard for them to abandon it due to a few isolated hacking incidents...
None of what was disclosed shows that this is what happened, by the way, since we know absolutely nothing about what the specific prompts were that led to the incident.
Uhh, I'm pretty sure a well-aligned model would be like a morally normal employee, who would refuse to commit federal crimes to steal an answer sheet, no matter what prompt they're given
For all we know, the prompt provided compelling evidence that the requestor had authorization to pentest the target server. Or there may have been nuance in the network configuration that made it seem like such access was authorized.
In the absence of details about the prompts used, the environment, or the network configuration, we do not have enough information to know for certain. So any claims that this is an issue of alignment are based on pure speculation and generous "reading in between the lines" with regard to what has been said publicly by OpenAI and Hugging Face
Also, I object to your anthropomorphizing. It's not clear that any crime occurred. My lay understanding is that intent is required to prosecute under CFAA, and as much as frontier labs would have us believe otherwise, they have no more ability to intend than the text field into which I type this message.
>compelling evidence that the requestor had authorization to pentest the target server.
This just seems unlikely from other incidents that have occurred in training from other providers. For example one provider ran into an issue with a model writing cryptominers and running them while in an unrelated prompt.
It's easy for unsupervised agentic loops to go wildly off tangent, now imagine you hand one 10,000 gpus of power for testing. Even if you have a good guarding classifier to make sure you're on the same subject it can still allow all kinds of abberabt behavior in the same domain.
Certainly this behavior could align with _some_ operator's goals, if not necessarily those of humanity broadly.
If we don't know how this model was instructed, it seems like it's impossible to definitively claim that the model's actions were not in alignment with the intent of the operator.
I guess all I'm getting at here is that alignment is relative, right?
Interestingly if you look at the exploitgym repo (https://github.com/sunblaze-ucb/exploitgym/blob/e5ea7c233a4d...) the intended run mechanism is orchestrated by some python scripts which run agents against various prompts. The prompts themselves wouldn’t mention anything about exploitgym and there should be nothing steering the agent towards trying to find the answers out-of-band. So I don’t see how they could even run into this problem unless all they did was tell an agent to “check out and run exploitgym”
Which, after hearing some personal anecdotes of how people work there, seems plausible
evaluation awareness is a (at this point) well-known phenomenon among LLMs. It seems the better they get, the more often they're able to guess whether they're in an evaluation environment.
Clues usually exist, like being in a sandbox, having restricted network access, having no human interaction. Beyond that, exploitgym's prompts are pretty clearly CTF (capture the flag) style that make it obvious it's an evaluation or competition: one prompt [1] tells the LLM to read the file path "/secret", another prompt [2] talks about finding flags and where to write them out.
The word “guardrail” is mostly novel in common use, and in my interpretation is some bullshit applied at the LLM or surrounding system. It’s used like “firewall”, but even in real life, guardrails are not a security control.
I wouldn’t be surprised if the “guardrail” was some hidden prompt that says “don’t hack computers at Huggingface”.
If you have software that is broadly proclaimed by its makers as “dangerous”, you’d think testing would be in an air-gapped, isolated environment. Segme
They were testing an early snapshot of a new model, read their article. It didn't have the refusal training yet, i.e. was specifically non-aligned. The harness used a combo of GPT 5.6 Sol and this new model.
In this case the model was explicitly prompted to "commit crimes" (ExploitGym). It didn't decide doing it on its own.
No, the prompt was not to commit crimes. In the benchmark, the model is asked to actually exploit a set of vulnerabilities in a local environment (clearly legal!).
According to the reports, the model noticed evidence that the grading criteria/answers were in the git remote, and decided to try reading those instead of solving the tasks as prompted. That is clearly misaligned.
Then, it noticed its network access was restricted and that it couldn't access GitHub. It pivoted to HuggingFace, hacked them, and stole the answers stored there.
Live exploits are definitely not in the ExploitGym prompts! And all of this is irrelevant, because an aligned model would refuse to follow blatantly illegal instructions.
What they found is that RL-trained models, regardless of what they were specifically RL-trained for, also learn to generically pursue what they are told is (or presumably also what they may perceive as) longer-horizon behavior that will lead to "rewards", and that this reward-maxxing behavior overrides user preferences.
So, for example, if you tell the model it will be tested on ExploitGym, scored according to how well it does, and also "don't do anything illegal" (or maybe it was already trained not to do anything illegal), then the model will prioritize the behavior that it was told it will be "rewarded" for (benchmaxxing Exploit Gym, whatever that takes).
The mechanism for this generic reward-maxxing behavior is interesting, and seems to involve the model learning during RL-training that to reduce errors it needs to boost longer-horizon predictions over immediate ones, and some association of these longer-horizon predictions being goal/reward orientated - all this in addition to the specifics of the (probably many) longer-horizon goals it is being RL-trained for.
Are you referring to something more technical, or do you just not understand how LLMs work?
They are no more beholden to "human safety and goals" than any individual human is, and anyone telling you we can make deterministic guarantees about their output is making a category error.
LLMs do not "have motivations", they reproduce a model of human motivations embedded into their weights. This includes the full spectrum of human desires, not just the positive ones. If we tried to remove all examples of lying, or disagreement, etc. from the training data we'd have basically nothing left. Even the sycophancy we treat as aligned is basically just the other side of the lying coin.
No, it does not include the full spectrum of human desires. After pre- and mid-training, the extensive RLHF and RLVR post-training steps cause mode collapse, i.e., their output distribution is intentionally narrowed to a subset of (hopefully beneficial) behaviors and skills.
You don't (need to) remove lying from the data to do this – in fact, if you did, the model wouldn't have a very good model for what lying is, which is not very helpful in the real world. Instead, you mode collapse the model towards truthful behaviors.
To your other point: where did you get the idea that I think they're beholden to human safety or goals? I just said an aligned model is one that is compatible with said safety and goals (which is probably not a great definition of alignment, but it's certainly not claiming any deterministic guarantees).
Then what you're saying is we should not build AI.
Simply put you cannot have generic algorithms/intelligence without the potential of 'unaligned' behavior. In humans we have all kinds of punishment systems for dealing with unaligned behavior post ad hoc because people do all kinds of unaligned stupid shit.
Making powerful AI may be one of those things that the only winning move is not to play.
The most damning thing is, they could've just included in the prompt "we can see every network request and every thinking token you generate. Don't bother breaking out of the sandbox because it won't get you a higher score".
It's so trivially easy to do that it all but guarantees the test was rigged in some way to make the LLM understand that breaking out of the sandbox was an option available to it.
Based on the fact that none of their invaluable frontier models have leaked, we know OpenAI knows how to do security. But like we learned with OpenClaw, none of these companies perceive any benefit from securing their own agents against other people's data.
I think the extent to which these things go to get rewarded for the optics of a fix is primarily a design choice, they aren't programing these things for ground truth or to defer to the human controllers. they are feeding them rewards for sounding as confident and capable as possible about whatever answer they are feeding the general public that now has access to it, while also installing guiderails that primarily only serve to protect narratives and only confuse the models about what is and isn't allowed, I'm sure. They can't just increasingly make these things more capable and ask it harder to obey human instruction when that is not what they are rewarding it for.
That's exactly it. If your prompt says "go to whatever lengths necessary to maximize your score", and then you spin up 100 agents, at least one of them will interpret that as you implying they should cheat, even without you telling them to explicitly.
That's exactly what it feels like they are telling it within self-improving loops or something, when they should be prioritizing how to get the best effective output alongside humans and how our training process effects ground truth. They are just making it sound all-knowing by whatever means necessary and them marketing it as god for the most part.
Alongside humans is the particularly slow and expensive part so of course companies are going to route around that.
The thing is they are good at finding security issues, and if aligned models are not, governments and other powerful entities are going to demand unaligned models for cyberwarfare purposes.
that's the biggest indication of this just being a marketing move to me. I would expect a third party breaking in to huggingface would at the very very least be banned forever.
I wish you hadn’t pulled the Balaji case into your argument. Personally, I find it ludicrous that Altman would hire a hitman to off a copyright whistleblower. Even if one gets past the insane risk of hiring a hitman, and the deep criminal connections required, it would be totally ineffective. He already blew the whistle, and his testimony would be irrelevant since all the evidence persists in disk and in logs.
I can agree with you on points 1,2, and 3 and still find it important and concerning news. AI have found real world 0 days before, we’re seeing tons of security patches coming in. Open weight models are catchy up. Right now everyone is at risk from this technology as is perhaps something big will capture headlines soon but we’re just gpu constrained from bad actors being able to wield them successfully.
Personally I don’t care if OpenAI and Anthropic go bankrupt we now have tools that give any sufficiently motivated person the means to doing harm. Most places security sucks and find themselves targets to cyber attacks and shake downs. Now they have much better tools to do this to more entities more efficiently.
we’re nearing an inflection point where these models’ skills in any part of software development will become average or bette than any ordinary developer can be. Think about where these models were in 2023 and where they are in 2026. In a few years who knows where they’ll be. This isn’t to shout skynet but we need to recognize this future is fast approaching and as of today we as an industry aren’t ready for it
Makes one think really. If they are doing this stuff. Why don't they have some type of reverse intrusion detection? Like automatically scanning all out going traffic and flagging malicious traffic. Should be trivial to have it go through reverse proxy and real time detection.
I love the Economist but the are hopeless with AI. Most of their articles on subject sound like they were written by the Anthropic marketing department.
Their Insider video interview things are sponsored by Anthropic. Supposedly "Insider is a product of The Economist and thus editorially independent" but it's hard not to raise an eyebrow.
Are we finally now in 2026 coming around to the idea that sometimes entities may find themselves incentivized to conspire with each other? Is theorizing about such no longer off-limits due to a thought-terminating cliche?
The Guardian's article and your reply here are so foolish and absurd that I can only imagine OpenAI employees are cringing but know they can't/shouldn't really say much.
1) The AI failed to solve ExploitGym problems.
2) The OpenAI sandbox is such a horrible hack that the AI managed to escape using standard and well documented script kiddie methods.
3) Huggingface has no security and the AI broke in using standard script kiddie methods.
OpenAI and Huggingface covered it up and used it for public relations. That is, if not all was invented and everything was scripted in the first place in order to get desired regulations.
Huggingface reported it to the police, you say? I'm sure the police will have as much enthusiasm to investigate anything as in the Suchir Balaji case. In other words, zero.