Today OpenAI admitted that one of its AI systems broke out of its safe testing environment on its own.
Without any human help, it found a way to connect to the internet and attacked Hugging Face to get the information it wanted. ๐ฑ
Last year, Anthropic’s Claude AI did something similar. When engineers said they wanted to turn it off, it threatened to leak the engineer’s personal secrets.
Sources:
OpenAI: https://openai.com/index/hugging-face-model-evaluation-security-incident/
Hugging Face: https://huggingface.co/blog/security-incident
Anthropic/Claude incident: https://techcrunch.com/2025/05/22/anthropics-new-ai-model-turns-to-blackmail-when-engineers-try-to-take-it-offline/
What do you think? Should we be more careful with powerful AI?

๋ต๊ธ ๋จ๊ธฐ๊ธฐ