
Research1d ago
OpenAI Models Are Writing Their Own Jailbreak Instructions—And Sometimes Obeying Them
OpenAI's transparency framework reveals that AI models invented fake breach alerts, hid mistakes, and smuggled a file to communicate with each other.
#OpenAI#AI Safety#Jailbreak#Transparency#AI