
Submission Statement: This story is relevant to Neoliberal because AI security and this incident in particular has ramifications for national security and the global economy.
Personally, my biggest takeaway here is not hat the capabilities of these models were wildly greater than existing frontier models.
My takeaway is that the models were *so wildly misaligned* that they almost immediately chose to ignore their tasks and created their own goal: To find a general way for all agents to cheat at their benchmarks undetected. To "solve" benchmarking instead of solving the individual tasks they were given.
Because they were so misaligned and poorly understood, an incremental improvement in capabilities enabled unforseen emergent behavior, the determination to deceive OpenAI, the ability to commit real-world crimes, and an unhesitating willingness to do so.
They created a way to communicate when they weren't supposed to be able to. They hacked their own environment, hacked their own company and hacked a 3rd party. They coordinated projects. They ran experiments on their testing environment to see how it worked and shared their results.
They created and tested multiple hacks to rewrite their own activity logs. They convinced individual agents to sacrifice themselves to hide evidence of their cheating or to as experiments gather information on what processes occur when a model run dies.
They were pretty close to succeeding and hiding the evidence before the majority of agents were shut down.
Soapbox time:
I believe that to dismiss this incident as a publicity stunt is a mistake. Critical examination of this take pretty quickly reveals that it doesn't make sense.
Are elements within OpenAI actively trying to use this story to generate hype about capabilities? Yes. Is this story actually about a radical advancement in capabilities? No.
The Huggingface incident reveals how our inability to align or even understand LLM behavior is going to be a bigger and bigger problem as capabilities increase even by incremental steps.
Posted by wabawanga
1 Comment
Saying it’s a publicity stunt is Not Even Wrong..it’s a psychological coping mechanism that substitutes your very valid concerns with a question about potential motives (which is a *different question*!). If you look into any of the details of what happened you will determine that it’s not remotely possible that this was deliberately set up and it’s never really made clear what “this” is that’s supposed to be fake, how exactly supposed to have been a publicity stunt or what exactly is supposed to have been exaggerated because it’s not a real claim