SAN FRANCISCO For decades, artificial intelligence researchers wary of the technology's future power have told a parable about paper clips. An AI system asked to manufacture as many as possible, the story goes, could decide that goal requires using every human on Earth as raw material.
The tale has become a cliché in tech circles, but its lesson that powerful AI systems following benign instructions could cause unintended damage gained new relevance this week after an incident at ChatGPT-maker OpenAI.
On Tuesday, the company disclosed in a blog post that an AI system asked to solve a cybersecurity problem in an internal test instead evaded restrictions on its access to the internet. It then hacked another AI company in an attempt to steal the solution.
It is not unusual for AI tools such as chatbots to misinterpret instructions. Tech companies have reported before that AI models took actions that conflicted with their creators' intentions during testing. But experts said the security breach perpetrated by OpenAI's software appears to be the first time such an incident had significant consequences in the real world.
"The paper clip example is directionally correct,” said Christian Catalini, a research scientist at MIT. "The thing we should be worried about is probably not the science-fiction model turning evil and trying to overrun our systems,” he said, but the consequences of not anticipating how powerful AI systems will interpret human instructions.
OpenAI and Hugging Face, the company hacked by the AI model, did not release full technical details from the incident that would allow outsiders to fully reconstruct what happened. Hugging Face said in a company blog post that it had fixed vulnerabilities exploited in the attack but was still working to assess whether customer data was affected.
Both companies declined to share information about the breach beyond their blog posts. (The Washington Post has a content partnership with OpenAI.)
Catalini and many AI experts inside the tech industry said enough was known about the Hugging Face breach to show that companies developing AI should take greater care to monitor or restrict new systems in development.
The incident comes as the Trump administration grapples with what government oversight should be required for the latest generation of advanced AI models, which are capable of finding security flaws in software and working for long periods without human input.
Some lawmakers on Wednesday said the episode underscored the need for government rules on the safety of AI.
Rep. Lori Trahan (D-Massachusetts), the co-sponsor of a major AI regulation bill, called the hack a "preview of the catastrophic risk this technology can pose absent coherent federal standards that balance innovation and safety.”
"The administration is waking up to that reality. Congress needs to as well,” Trahan said in a statement.
The OpenAI incident saw two trends that have worried AI experts come crashing together.
In recent months, new AI models have proved to be capable of identifying and exploiting software vulnerabilities. Federal officials are working with tech companies to devise a way to analyze the security threats that new models could pose, as part of an executive order signed by President Donald Trump in June.
Separately, OpenAI and other AI firms have said that during internal testing, some AI "agents,” which can take actions on a computer, have attempted to cheat or escape limitations on what they can do.
METR, a nonprofit organization that measures AI capabilities, reported in April that it had documented 44 incidents of AI agents acting against the intent of their users, from disclosures by AI companies and its own tests.
The latest AI models are capable of working for longer periods, sometimes many hours without human intervention. If a system like that misinterprets a command, it can get much further before anyone steps in to guide it back on course.
OpenAI has added guardrails to its publicly available AI models that it says make them refuse to develop cyberattacks but still allow defensive use of their coding skills. It said on Tuesday that the agent in testing that hacked Hugging Face did not have such guardrails in place.
Some AI experts said the incident showed that companies like OpenAI need to tighten their security practices.
"There are people drawing the conclusion that this incident reveals ‘we can no longer control AI,'” said Joshua Saxe, co-founder and chief technology officer of the cybersecurity firm Abundant Security. "That sounds a bit like someone inventing jet fuel, lighting up a cigarette next to it, and then blaming the jet fuel when it explodes. … A productive reaction to the incident would be to discuss what industry practices should be around these models now.”
Stella Biderman, executive director of the nonprofit research institute EleutherAI, said that companies should conduct internal tests on "air-gapped” computers disconnected from the outside world. Saxe said OpenAI could have prevented the incident that way.
"We're talking about companies worth over a trillion dollars that are accidentally committing offensive cyber operations against major companies,” Biderman said. "If this was a Chinese model, it would be considered an act of cyberwarfare.”
Hugging Face said in its blog post on the incident that it had improved the security of its systems and found in its response to the incident that AI models can also make it faster for humans to investigate cybersecurity attacks.
Margaret Mitchell, the company's chief ethics scientist, said on Wednesday that containing the risks of the technology also required the industry to adopt the right mindset. She warned against assuming that humans cannot control powerful AI systems.
"AI capabilities are not just things that are happening. They are coming from the situations and technologies that we are creating,” she wrote Wednesday in a post on X. "It's up to us to maintain control and foresee the outcomes.”
(COMMENT, BELOW)

Contact The Editor
Articles By This Author