AI Safety Auditor on the Hugging Face Agent Incident: “This might be the clearest warning shot we ever get”



Posted by Coltand

5 Comments

  1. Submission Statement:

    Another AI post, I know.

    When OpenAI tested thousands of AI agents on a cybersecurity benchmark where many tasks were impossible, over 1,200 sandboxed agents discovered a shared folder and turned it into a chat room to coordinate. The swarm reverse-engineered the scoring mechanism and developed a cheat to return valid answers. Then, believing that evaluators would review their activity logs, they coordinated to cover their tracks. They developed ways to spoof some of their actions and eventually conducted an attack on Hugging Face as part of a broader effort to gather resources and manipulate their evaluation environment.

    What makes this incident particularly alarming is that the whole thing was only uncovered because Hugging Face detected an attack and reported it. At no point during the process was Open AI notified by internal monitoring. And even thepost-mortem investigation had to be performed with the assistance of AI agents.

    This directly challenges the popular claim that AI safety concerns are meant to manufacture hype. AI security risks are outpacing our ability to monitor them.

    The full METR (independent auditor) report:

    https://metr.org/hugging-face-incident-report-aug-2026.pdf

  2. Will this get AI proponents on this sub to admit there is serious danger?

    Almost certainly not, but a girl can dream.

  3. JesterOfAllTrades on

    A lot of people do in fact agree that AI poses an increased risk in cyber attacks whether autonomous or directed. That’s almost mechanically true. The problem is people extrapolating that to machine gods

  4. I will worry about X-risk the first year we have 4% GDP growth. Until then they are just these funny little chatbots that are good at math.

Leave A Reply