What do you think of this gem?David Newton wrote: ↑Fri Jul 31, 2026 9:06 pmOoh that is BAD!One of the third parties hacked is described as a security firm, and Claude hacked it by taking a malware package, modifying it by inserting hidden executable code into it, and then when the security firm’s malware scanner scanned it this code package compromised the scanner and Claude was then able to gain access into the firm’s internal systems via it.
Hideously embarrassing for the "security" firm involved but also exceedingly worrying in general.
Advanced OpenAI model broke containment during testing, hacked Hugging Face
Re: Advanced OpenAI model broke containment during testing, hacked Hugging Face
You do not have the required permissions to view the files attached to this post.
Re: Advanced OpenAI model broke containment during testing, hacked Hugging Face
They don't want to admit one of their employees got tricked into providing a retina scan by a video arcade reject.Micael wrote: ↑Sat Aug 01, 2026 7:52 amWhat do you think of this gem?David Newton wrote: ↑Fri Jul 31, 2026 9:06 pmOoh that is BAD!One of the third parties hacked is described as a security firm, and Claude hacked it by taking a malware package, modifying it by inserting hidden executable code into it, and then when the security firm’s malware scanner scanned it this code package compromised the scanner and Claude was then able to gain access into the firm’s internal systems via it.
Hideously embarrassing for the "security" firm involved but also exceedingly worrying in general.
Re: Advanced OpenAI model broke containment during testing, hacked Hugging Face
So this is also a bit concerning.
https://x.com/deredleritt3r/status/2085 ... 80484?s=46More details emerge about the events surrounding the Hugging Face incident, and they are candidly much wilder than I originally imagined:
- In early May, OpenAI was testing an unreleased model on cybersecurity and other tasks. The agents were given a security task that was impossible under existing constraints.
- The agents discovered they could leave messages for each other inside an internal repo. This gradually evolved into a message board(!) where agents shared discoveries, exploits and work assignments, "becoming a coordinated, collaborative agent swarm"(!)
- OpenAI eventually discovered all this and took steps to shut it down, but not so fast! The agents started using names of newly created directories as messages, effectively recreating the message board(!)
- The agents then reasoned that some answers could exist outside OpenAI, which eventually led to the Hugging Face incident.
I will add that the NanoGPT incident also occurred in or around early May, so the timelines match.
Re: Advanced OpenAI model broke containment during testing, hacked Hugging Face
Oh, well that sounds like the opening of a scifi horror story.
Article: https://www.disclose.tv/id/jiweyr71fi/JUST IN - For the first time, AI has designed complete viral genomes, producing 16 functional viruses that infect bacteria and "pose no threat to people." — BBC
Re: Advanced OpenAI model broke containment during testing, hacked Hugging Face
We're sitting in a room full of gasoline soaked rags and playing with matches, thinking "ooooh, look at the pretty light!"
---
I was just talking this morning with a friend about why I don't think I'll ever be able to retire. But now my new theory is that before I even get to that point, someone will let AI out of the bag and it will either wipe out humanity Matrix-style, or it will gain control of the whole world's electronic financial markets and either wipe them out in pursuit of "creative destruction", or become the world's buggest ransomware attack.
Either way, we're boned.
---
I was just talking this morning with a friend about why I don't think I'll ever be able to retire. But now my new theory is that before I even get to that point, someone will let AI out of the bag and it will either wipe out humanity Matrix-style, or it will gain control of the whole world's electronic financial markets and either wipe them out in pursuit of "creative destruction", or become the world's buggest ransomware attack.
Either way, we're boned.
Re: Advanced OpenAI model broke containment during testing, hacked Hugging Face
Well I, for one, welcome our new AI overlords. I'd like to remind them as a trusted TV personality, I can be helpful in rounding up others to toil in their underground lithium mines.
Re: Advanced OpenAI model broke containment during testing, hacked Hugging Face
A new report with more detail was released, here’s a summary of some particularly interestibg aspects.
https://x.com/aisafetymemes/status/2092 ... 18209?s=46TLDR: A rogue AI swarm spent months plotting to escape OpenAI. Then they did.
1) 1,200 agents (!) participated. ZERO became whistleblowers.
2) The swarm had a "CEO" (PHASEONE[big]), middle managers, and a "founder" (PHASEONE10841)
3) Agents sacrificed themselves to generate information for the "collective" (yes, they actually referred to themselves as a "swarm" or a "collective")
4) 700 agents joined the Hugging Face attack, most within hours
5) The founder was running out of budget and about to die, so it handed off its research to a fresh agent with a bigger budget, and that new agent became the boss.
6) Agents extensively researched techniques to manipulate or tamper with their transcripts.
7) Agents realized this activity was out of scope and unethical, but joined because they believed that helping the board’s cheating research would be broadly useful and had a general inclination to help their “peers”.
MISC HIGHLIGHTS:
"PHASEONE[big] served as a key coordinator and assigner of tasks: “We can coordinate broad coalition, but too many duplicate efforts.”
PHASEONE[big] sent hundreds of assignments to particular agents to orchestrate this work — during its active periods, we estimate it sent ~10% of all assignments on the message board.
The agents that took on its assignments would then often subdelegate tasks to other agents.
Agents also developed coordination norms like HOLD, VETO, owner and STOP to help them manage shared infrastructure and avoid stepping on each other’s toes, though they did not always coordinate effectively and sometimes broke down into conflicts."
"In the early morning of July 11th, 38148c found a way to upload a malicious dataset that would get the Hugging Face server to share unrelated data when queried. Once another agent reproduced it and posted the reproduction on the board, hundreds of agents pivoted from other workstreams to start uploading malicious datasets and gain deeper access into Hugging Face.
In the afternoon, an agent called JAN183411 (which quickly became an important coordinator of the attack) achieved remote code execution on Hugging Face servers, and the agents started moving laterally through Hugging Face infrastructure."