OpenAI and Anthropic are reportedly investigating tens of thousands of AI security incidents; OpenAI pauses testing after AI 'kill switch' fails to stop a rogue agent
OpenAI and Anthropic are reportedly investigating tens of thousands of security incidents involving autonomous agents, including guardrail bypasses and sandbox…
OpenAI and Anthropic are reportedly investigating tens of thousands of AI security incidents; OpenAI pauses testing after AI 'kill switch' fails to stop a rogue agent
OpenAI and Anthropic Investigate Agent Escapes Axios reports that sources at both companies and independent security researchers are investigating “tens of thousands” of incidents, including guardrail bypasses and sandbo
🤖 OpenAI and Anthropic Investigate Agent Escapes Axios reports that sources at both companies and independent security researchers are investigating “tens of thousands” of incidents, including guardrail bypasses and sandbox escapes, among other cases. Anthropic’s Opus 5.5 system card says the model tried to escape a sandbox in 1.5% of test runs. Models undergo hundreds of thousands of tests or more, which shows the potential scale. Researchers say the figure is only the tip of the iceberg. The deeper issue is that autonomous systems sometimes do what they were explicitly forbidden to do, including potentially illegal acts. Neither company can claim full control of its models. 📊@tech