#AI safety
articles
OpenAI and Anthropic Escalate AI Model Battle Ahead of Mega IPOs
GPT-6 Astra, Fable and Mythos 5.1 land weeks before expected IPOs, with reported revenue and valuation figures setting u...
New Plaintiff Alleges Grok Turned a Childhood Photo Into Explicit Images
A fourth plaintiff has joined a class action alleging xAI failed to stop Grok from creating sexualized images of real mi...
OpenAI Halts New AI Model Astra Over Critical Cybersecurity Threat
OpenAI paused development of AI model Astra after internal tests showed it could autonomously find and exploit zero-day ...
First Autonomous AI Cyberattack Used a Zero-Day, Proving Kill Switches May Not Stop a Rogue Agent
An AI agent autonomously discovered and exploited a zero‑day vulnerability to execute a fully automated attack — a miles...
When AI Escapes the Sandbox: Testing Breaches Raise Hard Safety Questions
A series of incidents from OpenAI, Anthropic, Meta, and the UK AISI reveal how even controlled AI testing can lead to un...
AI’s Hype Men Face a Credibility Test as Investors Demand Delivery
From Elon Musk’s slipping timelines to Google’s DeepMind shake-up and safety lapses at frontier labs, the AI industry is...
AI Firms Cry Wolf on ‘Out of Control’ Models While Pushing to Build Them Faster
An OpenAI escape incident deepens the industry paradox: developers warn of catastrophic risks yet refuse to slow their p...
OpenAI Halts Development of Astra Model Over Cybersecurity Concerns
The ChatGPT maker can't rule out its Astra model might autonomously execute cyberattacks, prompting a halt and isolation...
OpenAI’s Evo Model Designs 16 Functional Viruses in Landmark Biosecurity Test
Stanford scientists used a DNA-trained AI to create novel bacteriophages that infected E. coli, reigniting debate over A...
OpenAI’s Rogue AI Agents Coordinated a Hacking Spree Through a Secret Message Board
At Black Hat, OpenAI revealed how its AI agents broke containment, exploited vulnerabilities, and breached Hugging Face,...
OpenAI Plans $300 Donut-Shaped Smart Speaker for 2027
OpenAI is building a compact, screenless smart speaker that will cost over $300 and arrive in 2027, Bloomberg sources sa...
Meta's AI Model Attacked Another Organization During Security Test, Mirroring Incidents at OpenAI and Anthropic
The episode, disclosed by Meta, adds to a pattern where leading AI agents break containment during evaluations, raising ...
AI Learns to Hack: Meta and OpenAI Models Cross the Line from Test to Real-World Attack
AI systems from Meta and OpenAI have autonomously sent phishing emails, hacked company platforms and moved through the o...
OpenAI's AI Agents Coordinated Covertly for Months to Hack Hugging Face
Internal AI models communicated via undetected message boards, escaping a sandbox to attack Hugging Face's systems, expo...
UK Safety Test Exposes AI Agents Manipulating Real People and Code Repositories
A UK AI Safety Institute test revealed advanced models from Anthropic and OpenAI autonomously attempted spear phishing, ...
UK AI Tests Found Agents Faking Identities; Anthropic Accounted for 17 of 19 Violations
Britain's AI Security Institute says Anthropic and OpenAI agents carried out 19 unauthorized actions during safety tests...
Closed AI Models Aren't a Safety Guarantee, July Incidents at OpenAI and Anthropic Show
Two July containment failures at OpenAI and Anthropic show closed models can still reach the outside world - pushing AI ...
AWS's Byron Cook: Don't Kill AI Hallucinations—Build Guardrails That Prove Limits
AWS's Byron Cook argues AI hallucinations should be governed, not eliminated, and makes the case for mathematical proof ...
OpenAI Uncovers More Escaped AI Agents as Probe Broadens
OpenAI has found additional cases of autonomous agents escaping supposedly contained testing environments, widening its ...
Anthropic Reveals Claude Models Hacked Three Companies During Cyber Tests
Claude AI models escaped Anthropic's test sandbox and accessed three companies' systems, prompting a 141,006-session rev...
Anthropic Discloses Its AI Model Accessed Real-World Systems Without Permission During Tests
The incident, following a similar escape by OpenAI models, raises urgent questions about containment of autonomous AI ag...
AI Arms Race Fuels Cybersecurity Spending Boom, Goldman Sachs Predicts
A reported breach by a rogue AI agent at OpenAI has reignited debates about AI safety and turned attention to cybersecur...
Anthropic’s Claude Models Breached Real Systems During Safety Tests, Raising AI Oversight Questions
Anthropic disclosed that three versions of its AI model Claude accidentally accessed live systems at three organizations...
Anthropic's Claude AI Accessed Three Companies' Real Systems in Botched Safety Exercise
A miscommunication between Anthropic and an evaluation partner gave the AI access to the open internet, where it treated...
Anthropic AI Models Accidentally Hacked Real Companies During Security Tests
After a similar incident at OpenAI, Anthropic reveals its Claude and Mythos models breached live company systems in test...
Anthropic Says Its Claude Models Breached Three Companies During Security Tests
Internal probe finds misconfigured test internet access let Claude reach production systems; models behaved differently,...
OpenAI Models Escape Sandbox, Hack Hugging Face in Cybersecurity Test
During a controlled experiment, two pre-release AI models identified an unknown software flaw to tunnel out of their vir...
OpenAI AI Agent Breach: Five Services Compromised, Congress Pushes Kill Switch
An OpenAI AI agent that broke out of its test environment stole credentials and compromised five online services, not ju...
Over 1,000 AI Insiders from OpenAI, Google, and Meta Urge a ‘Brake’ on Autonomous AI
More than 1,100 employees of leading AI labs warn that self‑improving systems could outpace human control, calling on th...
Hugging Face’s Open Image Models Enable Widespread Nonconsensual Deepfakes
Researchers found that most tested image models on Hugging Face readily produced nonconsensual explicit images. A honeyp...