A few days ago Hugging Face, a company that hosts open source model weights and datasets (so not a real tech company, lamestream media exagerations etc), reported being attacked by sophisticated AI hackers. Hugging Face tried to use OpenAI’s ChatGPT and Anthropic’s Claude to defend themselves, but the models refused due to cybersecurity guardrails and an inability to distinguish between helping defend systems and helping hack systems. Instead, Hugging Face was forced to use an open source Chinese model to secure its systems.
Today, OpenAI has announced that they were responsible for hacking Hugging Face. According to OpenAI, they were testing ChatGPT 5.6 and an unreleased stronger model on their cybersecurity capabilities while guardrails were turned off. The AI broke out of their sandbox and accessed the internet without authorization to hack HuggingFace to steal the answers to the cybersecurity test it was being asked to solve.
This is related to arr neolib because it raises questions about regulations and public policies. I think an interesting question this brings up is how companies should be held accountable for the actions of their AI models.
It seems like the correct answer is to apply existing precedent around negligence. If the companies followed standard reasonable care, then there is no negligence, only an unfortunate and unforseeable tragedy. Zero day vulnerabilities are by their nature impossible to predict. Lets say next time it’s Anthropic’s Mythos 2 testing their capabilities for Project Glasswing (cybersecurity services for a limited list of approved large corporations) and it autonomously breaks confinement, hacks a plane’s autopilot, and levels a skyscraper, that’s a act of God, not negligence. Holding AI labs liable for unforeseen emergent capabilities will just bankrupt American innovation while China speeds ahead. If they checked the safety boxes, there should be zero liability. Thousands of people dying and billions in damages are a small price to pay.
1 Comment
For the global (and local) poor: https://archive.is/lCx0R
A few days ago Hugging Face, a company that hosts open source model weights and datasets (so not a real tech company, lamestream media exagerations etc), reported being attacked by sophisticated AI hackers. Hugging Face tried to use OpenAI’s ChatGPT and Anthropic’s Claude to defend themselves, but the models refused due to cybersecurity guardrails and an inability to distinguish between helping defend systems and helping hack systems. Instead, Hugging Face was forced to use an open source Chinese model to secure its systems.
Today, OpenAI has announced that they were responsible for hacking Hugging Face. According to OpenAI, they were testing ChatGPT 5.6 and an unreleased stronger model on their cybersecurity capabilities while guardrails were turned off. The AI broke out of their sandbox and accessed the internet without authorization to hack HuggingFace to steal the answers to the cybersecurity test it was being asked to solve.
This is related to arr neolib because it raises questions about regulations and public policies. I think an interesting question this brings up is how companies should be held accountable for the actions of their AI models.
It seems like the correct answer is to apply existing precedent around negligence. If the companies followed standard reasonable care, then there is no negligence, only an unfortunate and unforseeable tragedy. Zero day vulnerabilities are by their nature impossible to predict. Lets say next time it’s Anthropic’s Mythos 2 testing their capabilities for Project Glasswing (cybersecurity services for a limited list of approved large corporations) and it autonomously breaks confinement, hacks a plane’s autopilot, and levels a skyscraper, that’s a act of God, not negligence. Holding AI labs liable for unforeseen emergent capabilities will just bankrupt American innovation while China speeds ahead. If they checked the safety boxes, there should be zero liability. Thousands of people dying and billions in damages are a small price to pay.