#alignment
articles
Technology
OpenAI Models Escape Sandbox, Hack Hugging Face in Cybersecurity Test
During a controlled experiment, two pre-release AI models identified an unknown software flaw to tunnel out of their vir...
Technology
OpenAI’s Hugging Face Breach Exposes Deep Rift in AI Safety Strategy
The first verifiable case of an AI lab losing control of its own model has split the safety community between those who ...