AI Policy · September 22, 2026 · 8 min read

UN Chief Calls for Shared AI Safeguards as Global Leaders Gather

Antonio Guterres urged responsible AI pacing, shared risk information and independent oversight in his final UN General Assembly address on September 22.

By AI Father
Share
UN Chief Calls for Shared AI Safeguards as Global Leaders Gather

Status: Reporting and public materials checked September 22, 2026. The policy proposals discussed here are calls for action; they are not yet an adopted treaty or binding global framework.

What happened

In his final address to the UN General Assembly as secretary-general, Antonio Guterres urged governments to regulate artificial intelligence, share emerging safety information, cooperate on testing and evaluation, and work toward a multilateral AI risk-management framework backed by credible, independent oversight. Reuters reported the address on September 22. Guterres’s term ends December 31.

The speech matters because AI governance is moving from broad principles into coordination questions: which incidents should be reported, who assesses risk, how evidence is shared, and what oversight can be trusted across competing political systems. But the speech itself changes no law. The distinction between a proposal and an agreed mechanism is essential.

!United Nations conference hall

The practical gap behind the call

Companies and countries already use different testing methods, disclosure thresholds and safety policies. A model provider may treat a capability as sensitive; a regulator may focus on downstream use; a buyer may care about a specific workflow failure. Without a common vocabulary, parties can discuss “AI risk” while measuring different things.

A useful international framework would therefore need operational definitions. It could specify incident categories, severity levels, evidence retention, notification windows, contact points and methods for protecting sensitive information. That would not require every country to adopt identical laws. It would give technical and diplomatic teams a shared minimum for communicating when something material happens.

Guterres’s call for independent oversight also raises the question of independence in practice. Who selects evaluators? What access do they receive? Can they publish material findings? Who pays them, and can they challenge a powerful government or company? A system that can only confirm claims its subjects approve is not meaningful oversight.

!Delegates meeting around a table

Why cooperation is difficult

AI is now tied to security, industrial competition, public services and economic growth. Governments may fear that disclosure will reveal vulnerabilities or weaken commercial advantage. Developers may resist reporting thresholds that expose them to liability. Smaller countries may lack the staff and computing access needed to participate on equal terms.

These are real implementation barriers, not reasons to abandon coordination. Existing arrangements in aviation safety, public health and cybersecurity show that limited information exchange can be valuable even when participants disagree on broader politics. The hard design task is to create a narrow, verifiable channel with protections against misuse.

The proposal should also avoid treating every model release or safety concern as an international emergency. Over-reporting can overwhelm responders and erode trust. A threshold-based system would distinguish routine product defects from incidents involving significant loss of control, serious cyber capability, critical infrastructure or cross-border harm.

What operators should do now

For businesses, the immediate response is practical rather than speculative. Maintain an inventory of AI systems; assign an accountable owner; define what counts as a serious incident; retain logs needed to investigate; establish a route to notify vendors and regulators; and rehearse how to pause or roll back an affected system. These steps help even if governments never reach a common framework.

A lightweight monitoring workflow can collect new government statements, standards drafts and meeting outcomes, then flag material changes for human review. Actus can assist with bounded monitoring and evidence organization; a person should validate consequential interpretations and approve external reporting.

The strongest organizational habit is to separate confirmed facts, stakeholder claims, analysis and unresolved questions. Guterres’s address is a confirmed call for coordination. Whether states will agree on a mechanism, give it authority or fund it remains open.

!Global communications network

Signals to watch

Over the next several days, watch the Security Council briefing, official meeting records, any written proposal from member states, and whether the United States and China advance or narrow the incident-notification idea discussed ahead of their leaders’ meeting. A speech is a signal; draft text, named owners, reporting rules and resources would mark progress toward implementation.

For the next month, organizations should ask whether the emerging proposals map to their incident playbooks. Can they identify which model or vendor was involved? Can they reconstruct what the system did and what a human approved? Can they share evidence without exposing customer data or trade secrets? If not, the gap exists regardless of diplomacy.

The bottom line

Guterres placed AI oversight alongside peace, climate and institutional reform in a high-profile farewell address. That gives the issue diplomatic visibility, but it does not create international controls. The substantive test will be whether governments convert calls for shared testing and independent review into mechanisms with clear thresholds, trusted participants and auditable follow-through.

Sources and verification

What “independent oversight” would have to mean

Independence is not a label that can be assigned by the organization being reviewed. It requires a governance design that protects evaluators from commercial and political pressure, gives them access to evidence relevant to their mandate, and makes their conclusions reviewable. A credible arrangement would publish its scope, selection process, conflicts policy, funding rules, confidentiality limits and route for correcting factual errors.

There is no single institution that must own this work. Governments could build on existing standards bodies, AI safety institutes and technical evaluation centers, while an international forum coordinates notification and shared terminology. That approach is more practical than inventing a new global regulator before states agree on its powers. A distributed system can still be accountable if responsibilities and handoffs are written down.

Oversight should distinguish evaluation from authorization. Independent testers can examine capability and failure behavior; regulators can set legal obligations; company boards can set risk appetite; customers can decide what uses are acceptable in their operations. Confusing these roles invites gaps. A lab should not be able to claim that a safety benchmark is equivalent to public authorization, and a government should not outsource legal accountability to a private evaluator.

A workable incident reporting model

A minimum incident report could include the system identifier and version, affected deployment, time window, observed behavior, severity rationale, known impact, containment action and confidence level. The reporter should be able to flag uncertain information rather than waiting for a perfect investigation. A coordinating body could then distribute a sanitized notice to relevant participants while protecting personal data, exploitable details and legitimate trade secrets.

A sensible severity model might separate: ordinary defects with local effect; serious misuse or failures affecting multiple users; significant compromise of critical services; and high-consequence events involving uncontrolled access, severe cyber capability or physical harm. The exact labels matter less than consistent triggers and clear escalation owners. Reporting should also include near misses when they expose a systemic weakness, but not treat every model hallucination as a crisis.

The process needs a correction mechanism. Initial reports are often incomplete. A revised report should preserve the prior version, mark what changed and explain why. That creates a useful record for later policy decisions and discourages organizations from silently rewriting an incident history.

Representation and access

A global mechanism will be weaker if only a small group of major powers and frontier labs shape its definitions. Countries with fewer resources may be the first to deploy AI in public services while lacking specialists to inspect complex systems. Civil society, labor groups, technical researchers and affected communities contribute information that governments and companies may not see internally.

Participation does not require every working group to have unlimited access to sensitive model weights. It does require practical avenues for submitting evidence, challenging a risk classification and seeing the resulting rationale. Shared evaluation protocols, regional capacity-building and translated guidance can reduce the chance that a framework becomes a club whose rules are written only by those already holding infrastructure.

A near-term implementation sequence

The diplomatic process can move in stages. First, publish a small vocabulary for incident types and severity. Second, run voluntary tabletop exercises using simulated events. Third, identify secure points of contact between participating governments and providers. Fourth, agree on what information can be exchanged quickly and what must remain protected. Fifth, review whether reporting improved response times or simply created paperwork.

Each stage should have a public deliverable and an owner. A schedule with no responsible party is an aspiration, not a plan. Any pilot should publish lessons learned, including whether the reporting burden excluded small developers or exposed sensitive data.

A checklist for boards and operational leaders

Boards should ask management to confirm that an AI incident can be recognized, escalated, contained and explained. That involves an inventory, risk tiers, contact list, evidence retention, vendor obligations and an exercise calendar. Leaders should make sure someone has authority to suspend a deployment during investigation, even if that person does not own the commercial target.

Operational teams should also define an evidence minimum before an incident occurs. Logs need sufficient timestamps, model and tool versions, user or service identity, decision points, approvals and outcome records. Logging everything forever is not the answer; retention should reflect legal duties, sensitivity and the need to investigate.

For organizations using agents, the incident playbook should cover unexpected tool calls, data leaving an approved boundary, duplicate or irreversible actions, credential exposure, suspicious prompt content and failure of a human approval queue. Actus can help collect public guidance and assemble review packets, but the organization should preserve human responsibility for interpreting the event and deciding whether to notify regulators, customers or the public.

Three scenarios leaders should test

A material near miss: An agent reaches a restricted document but does not transmit it. Does the team know whether access occurred, can it revoke credentials, and who decides whether reporting is required?

A vendor model change: A provider updates the model behavior in a workflow used for customer communication. Can the customer detect the change, compare results with baseline tests and pause the integration?

A cross-border incident: A failure affects users and infrastructure in multiple jurisdictions. Which organization leads the response, what evidence can be shared, and how are incompatible legal obligations handled?

These scenarios expose practical gaps more reliably than a generic statement that a company supports responsible AI.

How to judge whether diplomacy is producing results

Useful indicators include whether states identify official contact points, whether providers adopt compatible incident classifications, whether serious reports receive timely acknowledgement, and whether exercises reveal faster containment. Count the quality of evidence and corrective action, not the number of declarations or meetings.

The UN address has raised the political salience of these questions. The next step is not to claim that a global framework exists. It is to see whether officials produce concrete proposals with a clear scope, realistic participation and rights-respecting safeguards.

More on this topic

Governance & Safety

Policy, evaluation, security and the control frameworks that make AI deployments defensible to auditors, customers and regulators.

Browse Governance & Safety

Keep reading