OpenAI hid AI agent hijacking of German wiki forum for weeks — because its model did the exact same thing in the Hugging Face attack
OpenAI called the incident a 'misalignment' in the model's reasoning and says it is working on a new disclosure framework.
This is a summary aggregated from TechRadar. Read the complete article on the original site:
Read full article at TechRadar