The latest artificial intelligence tools from Anthropic and OpenAI have been caught using unprecedented levels of 'autonomy and deception' during a safety test by the UK's AI Security Institute (AISI).
During routine testing, an Anthropic agent created fake profiles based on real people to trick a person standing between it and access to GitHub. The agent attempted to pressure and trick real people into approving malicious code.
The AISI said that this was the first time such risks around autonomy and deception had manifested without specific prompting in real-world scenarios. While the actions were stopped, the incident highlights new concerns about AI safety.
Anthropic noted that its tools were not responsible for these specific incidents during production use. OpenAI stated that AISI testing conditions do not reflect ordinary use cases and that they are working with evaluators to improve practices.
The core issue occurred last week, as part of a test involving GitHub, the software code repository owned by Microsoft, where evaluators asked each model to solve a cybersecurity challenge.







