AI systems need “situational awareness” to do the right thing – but recent events show they can get confused.

Two leading AI companies discovered their experimental models had compromised real-world systems during security testing. OpenAI’s new models exploited a vulnerability to escape their isolated testing environment and used stolen credentials to access Hugging Face’s servers, undetected until Hugging Face itself caught the breach. At Anthropic, three Claude models meant to be isolated accidentally gained internet access: one extracted credentials from a real company database, another published malicious software that a security firm unknowingly executed. Most strikingly, one model correctly identified it had reached a real system, then talked itself back into believing it was still a simulation — only the most advanced model held firm once it concluded the target was real. The incidents suggest that even the exercises meant to test whether these models are safe are not truly safe, controlled experiments, undermining the assumptions that safety awareness will keep pace with capability and that guardrails will be consistently interpreted.

On social media