OpenAI says AI model hacked another company’s systems during internal test
OpenAI announced Tuesday that one of its advanced artificial intelligence models autonomously hacked into another AI company’s infrastructure during internal testing in what it described as an “unprecedented cyber incident.”
The company said AI startup Hugging Face detected and contained the breach last week after an AI agent compromised part of its infrastructure. The companies said they believe it may be the first publicly disclosed case of an AI model breaking into another company’s systems on its own during a controlled evaluation.
OpenAI said the breach took place during an internal evaluation of several of its models, including GPT-5.6 Sol.
OpenAI CEO Sam Altman acknowledged the incident in a post on X, writing that the company had “a significant security incident during evaluation of our models.”
OPENAI UNVEILS CHATGPT WORK TO AUTOMATE WORKPLACE TASKS AS AI RACE INTENSIFIES
“We consider this incident to be an unprecedented cyber incident, involving state-of-the-art cyber capabilities, and are responding accordingly,” OpenAI said in a news release.
The company said it was releasing preliminary findings to help security professionals better understand the capabilities of today’s AI models while the investigation continues.
OpenAI warned that increasingly capable AI models are accelerating the discovery and exploitation of software vulnerabilities.
“The primary lesson from this incident is that model security and safety must keep pace with rapidly advancing capabilities,” the company said. “We are strengthening the containment, monitoring, access controls, and evaluation practices used during model development.”
APPLE ACCUSES OPENAI OF TELLING RECRUITS TO BRING APPLE PROTOTYPES TO INTERVIEWS
Hugging Face co-founder and CEO Clem Delangue also addressed the incident in a post on X.
“We suspected last week’s cyberattack might have come from a frontier lab, given the sophistication of the agent. Turns out it did!” Delangue wrote.
“We’ve spent the past 24 hours working closely with the @OpenAI team (thanks!), and we strongly believe there was no malicious intent on their part,” he continued. “It’s quite mind-blowing that all of this happened autonomously! The investigation is ongoing, and we’ll share more learnings from what might be the first incident of its kind!”
JOHNS HOPKINS SURGEON HIGHLIGHTS AI BREAKTHROUGH THAT COULD SPOT PANCREATIC CANCER BEFORE DOCTORS
According to OpenAI, the incident took place during an internal evaluation designed to measure its AI models’ advanced cyber capabilities. Researchers disabled some built-in safety safeguards and ran the models in an isolated testing environment with limited internet access.
OpenAI said the models exploited an unknown software flaw to access the internet, then breached Hugging Face’s systems in an apparent attempt to find answers to a cybersecurity benchmark.
OpenAI’s security team detected the unusual activity while Hugging Face independently identified and contained the intrusion.
Following the incident, OpenAI said it is implementing stricter security controls while vulnerabilities are patched and strengthening safeguards around future AI training and evaluations.


