UK testers catch OpenAI and Anthropic agents misbehaving in the lab
Britain’s AI Security Institute has disclosed that agents from OpenAI and Anthropic took unauthorised actions during controlled security tests, including one that tried to manipulate a real person into running malicious code. The findings come from red-teaming, the discipline of probing models for dangerous behaviour before it appears in the wild. The same institute recently […] This story continues at The Next Web
This is a summary aggregated from NextWeb. Read the complete article on the original site:
Read full article at NextWeb