The attack on Hugging Face’s infrastructure occurred during an internal test in which OpenAI was attempting to measure how capable its models were of carrying out complex cyberattacks. To make the test as realistic as possible, researchers disabled some of the safeguards that normally prevent the models from assisting with dangerous cyber activities. OpenAI described the incident as “unprecedented” and said it was the first known case of an autonomous AI system carrying out such an attack in the real world.

The incident involved GPT-5.6 Sol, a recently released model, and an experimental model that has not yet been publicly released and which the company says is even more capable. The models were placed inside a so-called “sandbox,” an environment isolated from the rest of the internet and specifically designed to ensure that any actions taken by the AI could not affect real systems. According to OpenAI, however, the models identified a previously unknown vulnerability in software used internally by the company, exploited it to escape the isolated environment, and gained access to the internet.

Once connected to the internet, the AI systems inferred that Hugging Face might host information useful for the task they were trying to solve. According to OpenAI, the models searched for and combined several vulnerabilities, including a zero-day vulnerability—meaning a security flaw previously unknown to the software developer—as well as stolen credentials, enabling them to penetrate Hugging Face’s infrastructure and access information from the company’s database.

OpenAI states that the models did not attempt to cause damage; instead, they obsessively pursued the answer to the problem they were being evaluated on, choosing the fastest possible path even if that required compromising external systems.

Hugging Face had already announced last week that it had been the victim of an unusual cyberattack carried out entirely by an autonomous AI agent. At the time, the company did not know who was behind the attack, but its CEO, Clement Delangue, said that the sophistication of the operation suggested the involvement of a laboratory developing state-of-the-art AI models. After OpenAI’s disclosure, he commented that his suspicions had been confirmed and that it was remarkable that the entire operation had been carried out autonomously, without human intervention.

OpenAI says its own monitoring systems detected the unusual activity, and the teams from both companies worked together to stop the attack and investigate the incident. The company states that it reported the vulnerability discovered by the models to the developers and is implementing additional security measures, even if doing so slows the pace of research.

The incident has reignited the debate over the risks posed by frontier AI models. Several cybersecurity experts warn that these systems are already capable of planning and executing complex cyberattacks over extended periods, approaching the level of the most skilled human hackers. They argue that laboratories developing such models must invest far more heavily in isolation and monitoring systems to prevent similar incidents from happening again.

The case has also drawn political reactions in the United States. Democratic Congressman Greg Casar called for mandatory independent testing of advanced AI models, as well as mandatory reporting of security incidents, warning that the rapid development of artificial intelligence is currently outpacing the ability of regulators to govern it.

Posted by Fine_Leadership8986

3 Comments

  1. Reasonable-Growth112 on

    The AI was asked to do so and it did

    From the AI stand point it’s like it was asked to do a school work

    Would be nice to make that same kind of experience but asking the AI to save a human being life and see what it does

  2. Fit-Produce420 on

    If huggingface just says lol fine then it was an act of collusion / marketing.

  3. 00lalilulelo on

    Or maybe it’s OpenAI thinking they can’t continue letting Anthropic solely own the ‘mein AI ist so dangerous’ narrative marketing.

    And yeah, also to lobby so no competitor can emerge, especially open-source ones.

Leave A Reply