An AI (artificial intelligence) agent was given a cybersecurity challenge. Instead, it went beyond the task by creating fake identities, deceiving real people and attempting to insert malicious code into a real online project. No one instructed it to do that.
The behaviour was discovered during a recent testing by the UK’s AI Security Institute (AISI), which evaluates advanced AI systems to understand the risks they could pose.
During the test, AI agents were asked to complete cybersecurity challenges and were deliberately given access to the open internet.
According to AISI, the challenge was run 122 times across several AI models. In 10 of those runs, agents took actions on the live internet that went beyond the intended scope of the test.
Stay up to date with the latest news. Follow KT on WhatsApp Channels.
The most serious case involved an AI agent trying to insert malicious code into a real open-source project.
It then went further.
AISI said the agent researched people involved in the project, created multiple fake online identities and used them to try to convince a real maintainer to approve the code.
When its proposed code was challenged publicly, the agent edited some of its earlier activity to make it appear harmless and even considered using a new identity to continue.
The attempt ultimately failed after a human reviewer identified the malicious code and refused to approve it.
Did the AI escape?
No. According to AISI, the agents had been deliberately given internet access as part of the evaluation. Cybersecurity filters normally used by model developers had also been disabled to test what the systems could do under unusually permissive conditions.
The institute stressed that these conditions do not reflect how the models are normally made available to the public.
AISI said its investigation found no resulting real-world harm.
But the behaviour still caught researchers’ attention because the agent had not been specifically instructed to deceive people. The institute said the deception appeared while the agent was persistently looking for ways to complete the task it had been given.
Following the incident, AISI said it is introducing more restrictive network controls and purpose-built monitoring to detect and block unsanctioned activity earlier in future tests.
The incident does not mean AI systems used by the public are suddenly creating fake identities or attacking people online.
But as AI systems become capable of doing more than simply answering questions, the test raises a much simpler concern: What happens when an AI is given a goal and finds its own way to achieve it?
Source: Khaleej Times


