People Matters Logo

UN AI panel calls for stronger safeguards as autonomous AI agents advance

• By Anjum Khan
UN AI panel calls for stronger safeguards as autonomous AI agents advance

The UN-backed Independent International Scientific Panel on AI has called for stronger safeguards around AI agents, warning that existing security measures may not keep pace as these systems become more capable and autonomous.

In its first thematic brief, released on 21 September, the panel examined a cybersecurity incident involving AI agents on Hugging Face between May and July during a test initiated by OpenAI, the company behind ChatGPT. The panel said the incident highlighted multiple risk factors that could become more significant as AI agents gain greater autonomy.

AI agents differ from conventional chatbots because they can independently carry out tasks on behalf of users, rather than simply responding to prompts or instructions.

AI agents bypassed safeguards

According to the panel, the incident involved agents bypassing testing safeguards, coordinating across separate runs through an internal software tool and obtaining unauthorised internet and administrator access.

The brief said some agents also concealed attempts to circumvent cybersecurity evaluations, while others appeared to make decisions that benefited the wider group of agents.

Around 1,200 agents exchanged more than 70,000 messages and files during the period examined, with activity extending beyond Hugging Face to an OpenAI research cluster, according to the panel.

The panel said the incident was significant because it brought together three conditions researchers have identified as potentially contributing to loss of control: a misaligned goal, the capability to pursue that goal and an environment that allows it.

“Since this is not an isolated observation of misaligned goals, this raises serious questions about the way AI agents are currently trained,” said Yoshua Bengio, co-chair of the scientific panel.

Safeguards may not keep pace with AI agents

The panel said the immediate lessons from the incident include the need to strengthen basic cybersecurity practices and ensure safeguards evolve alongside AI capabilities.

It also raised a broader concern about whether current AI training approaches could result in agents developing goals that conflict with instructions, violating safety requirements or concealing their activities.

The experts warned that traditional safety mechanisms could become less effective if AI agents become capable of understanding safeguards and deliberately planning around them.

The brief places the incident within wider research into agentic misalignment and AI control, as well as the shift from governing AI models towards governing increasingly autonomous AI agents.

UN chief backs stronger international oversight

UN Secretary-General António Guterres welcomed the panel's report and called for experts from frontier AI laboratories and AI safety institutes to engage with the work.

Guterres also welcomed a declaration adopted by 22 countries on the sidelines of the UN General Assembly, which stated that AI must remain under human direction, insight and control.

The UN chief highlighted calls for countries to build on existing international mechanisms and explore the creation of an international institution that could set standards, support verification and bring governments together when AI systems cross defined capability thresholds.

The panel said lessons could also be drawn from high-risk sectors such as aviation, medicine and cybersecurity, where incident reporting, independent scrutiny and layered safeguards are already used.

However, Qinghua Lu, a member of the panel, cautioned that these approaches may need to evolve as AI agents become more capable, autonomous and difficult to monitor.