Who Has the Authority to Stop an AI System?

AI safety evaluations can reveal dangerous capabilities and abnormal behavior, but they do not assign authority inside an organization. Effective control requires explicit roles and evidence for suspension, escalation, exception handling, restart, and accountability.

At 2:17 a.m., a customer-service AI agent begins processing cases at an abnormal rate. It classifies inquiries, retrieves customer data, drafts responses, and, when specified conditions are met, sends messages and updates the customer relationship management system. The overnight operator sees the volume climbing and a red stop button on the monitoring screen.

The operator does not know whether the anomaly justifies pressing it. A suspension could delay service to important customers. The business owner is asleep. The information security team may need to investigate. The privacy or legal function may need to act if personal data is involved. No one has established which call comes first, who may suspend external actions, or who may later authorize a restart.

This is a representative operating scenario, not a documented incident. Its institutional problem is real: the organization has a technical means of stopping the agent but has not assigned legitimate authority to use it.

AI safety does not end with model evaluation. Organizations must connect evaluation results and operational signals to explicit authority for approval, suspension, escalation, exception handling, restart, and accountability.

AI safety evaluation identifies risk but does not assign decision authority

AI safety evaluation measures capabilities, vulnerabilities, and behavior against defined risks. It can test whether a model helps a malicious actor execute a cyberattack, whether safeguards survive adversarial prompting, or whether performance crosses a risk threshold. Operational monitoring can reveal a surge in outbound messages, access to data outside an approved scope, or a sharp increase in human reversals.

These findings establish evidence about the system. They do not establish who inside a deploying organization may interrupt customer service, preserve a suspension, notify affected people, approve compensation, or accept the risk of restart. Those are institutional judgments with operational, legal, financial, and reputational consequences.

The difference matters because a threshold and an authority are not interchangeable. A threshold can trigger an automated control. An institution must still decide who defines the threshold, which actions it interrupts, who reviews the resulting evidence, and whose approval makes the next action legitimate.

Europe connects model evaluation to legal supervision through distinct mechanisms

The European Union has not created a single, abstract certification of AI safety. It has assembled legal obligations, a voluntary compliance instrument, and independent scientific advice. Each performs a different function.

Article 55 of Regulation (EU) 2024/1689, the Artificial Intelligence Act, requires providers of general-purpose AI models with systemic risk to perform and document model evaluation, including adversarial testing; assess and mitigate systemic risks at Union level; document and report serious incidents; and maintain adequate cybersecurity protection. These are legal obligations for the providers within scope. They are not general operating instructions for every company deploying an AI agent.

The European Commission describes the General-Purpose AI Code of Practice as a voluntary tool that helps providers demonstrate compliance with obligations under the AI Act. Its Safety and Security chapter applies to the limited group of providers subject to the systemic-risk duties in Article 55. A provider that does not adhere to an approved code or a relevant harmonized standard must demonstrate alternative adequate means of compliance for Commission assessment. Voluntary adherence to the code and mandatory compliance with the law are separate propositions.

The Commission also established the AI Act Scientific Panel under Article 68 of the Act and Commission Implementing Regulation (EU) 2025/454. Its 60 independent experts support the AI Office and national authorities by advising on capability evaluation, risk assessment methods, systemic-risk classification, cyber-offense risks, and market surveillance. The panel adds scientific advice to enforcement; it does not replace the authority of the Commission or national bodies.

These mechanisms show how evaluation can feed supervision and corrective action. They also preserve the distinction between the model provider governed by Article 55 and a deployer deciding whether its own customer-service process should continue operating.

The United States framework exposes the tension between secrecy and scrutiny

On June 2, 2026, the United States President signed Executive Order 14409, “Promoting Advanced Artificial Intelligence Innovation and Security”. Section 3 directed designated agencies, within 60 days, to develop a classified benchmarking process for advanced cyber capabilities and a threshold for designating a “covered frontier model.” It also directed them to design a voluntary framework through which developers could engage the federal government and provide access to covered models for up to 30 days before release to other trusted partners, subject to safeguards for confidentiality, cybersecurity, insider risk, intellectual property, use, and nondisclosure.

The order expressly states that Section 3 does not authorize mandatory government licensing, preclearance, or permitting for new AI models. The White House fact sheet likewise characterizes the arrangement as voluntary collaboration with developers.

Public information about what followed is limited. On August 3, Axios reported that a White House official said the framework had been completed by the deadline, while the administration had not disclosed its contents, who had seen it, or when companies would begin using it. On August 4, Axios reported details attributed to people briefed on meetings with companies and said the White House did not plan to publish the framework. These are journalistic reports about an unpublished process, not public text of the framework itself.

Some confidentiality is justified when an evaluation concerns exploitable vulnerabilities or advanced cyber capabilities. Yet secrecy narrows who can examine the risk criteria, evaluation authority, and consequences attached to a result. Completion inside government is therefore different from public verifiability.

That tension reinforces the governing question for enterprises. Evaluation may produce sensitive evidence. The organization must still establish a legitimate path from that evidence to a decision.

Japan's AI agent guidance does not make human presence sufficient

Japan's Ministry of Internal Affairs and Communications and Ministry of Economy, Trade and Industry compiled the AI Guidelines for Business, Version 1.2 on March 31, 2026. The revision addresses AI agents and encourages developers, providers, and business users to consider measures suited to the characteristics, uses, and risks of their systems.

The guidelines discuss human involvement, controllability, and final human judgment in relevant contexts. They should not be described as a statute imposing the same human approval requirement on every AI agent. Their risk-based guidance is valuable precisely because the appropriate form of human involvement depends on the use and potential harm.

An overnight operator already places a human “in the loop.” That fact does not create effective oversight. The operator needs authority to act, evidence that can support the judgment, enough time to intervene, the skill to interpret the situation, and an escalation path to someone with the mandate to decide what follows. Without those conditions, a human approval step may add delay while leaving practical control with the AI system and organizational responsibility unresolved.

Governance principles do not specify who presses the button

AI governance defines policies, oversight bodies, and accountability expectations. Digital transformation redesigns work and customer value. Automation moves execution from people to systems. AI ethics identifies values such as fairness, privacy, transparency, and human dignity. Each contributes to responsible AI use.

None of these disciplines, by itself, names the overnight role allowed to halt a specific customer-service agent when message volume exceeds an approved range. A governance policy may require oversight without allocating emergency authority. A transformed workflow may accelerate service without showing where judgment authority moved. An automated task may be technically executable even though the organization has not legitimately delegated it. An ethical principle may identify a conflict without appointing the person authorized to resolve it.

The missing layer is a structure that connects detection, recommendation, execution, suspension, approval, and explanation through named authority.

Decision Design turns a stop mechanism into an authority structure

Decision Design treats recurring judgment as an object of organizational design. It asks who holds legitimate authority, what role AI may perform, which evidence may inform the judgment, when the process must stop or escalate, who owns the outcome and explanation, which record preserves accountability through handoffs, and which event requires the arrangement to be reviewed.

Decision Design is not about improving decisions alone; it is about designing the authority structure within which decisions become institutionally legitimate.

This framework complements law, AI governance, ethics, risk management, organizational design, automation, and technical controls. It is not a law, public standard, settled academic consensus, or substitute for compliance.

Separate the service into distinct judgments

“Automating customer service” is too broad to govern. Inquiry classification, data selection, response drafting, message transmission, CRM updates, and refund or compensation decisions are different judgments.

Classification and drafting may be reversible and contained. An outbound message can expose personal data or bind the organization to a representation. A CRM update can alter the information used by later employees and systems. A refund changes a customer's rights and the company's financial position. The organization should assign AI autonomy and human authority separately for each action.

Define the Decision Boundary around authority, not only risk scores

A Decision Boundary identifies which judgments an AI system may perform and which a human role or institutional body must assume. The boundary can move with context. Normal operation differs from an incident. A reversible, low-impact response differs from a contract change, disclosure of sensitive data, or compensation decision.

Decision Boundaries are not operational thresholds; they are institutional demarcations of legitimate authority.

An operational threshold might automatically suspend outbound messages when volume exceeds a defined range. The Decision Boundary determines who was authorized to approve that threshold, who may extend the suspension to CRM writes, who can allow limited operation, and who has authority to restart. The threshold supplies a signal. The boundary allocates legitimate judgment.

Give the overnight operator provisional suspension authority

The organization can authorize the overnight operator to suspend external messages and CRM writes when observable conditions occur. Those conditions might include message volume outside an approved band, repeated identical messages, attempted access to unauthorized data, detection of sensitive information, authentication failures, or a sharp rise in reversals.

The operator should not bear sole responsibility for the business consequences of suspension. The operator exercises preassigned provisional authority. An on-call incident lead then determines the scope of containment. A privacy officer assumes the decision when personal data may have been exposed. The security lead assumes it when the evidence suggests intrusion. The customer-service executive and legal function decide customer notification, contractual remedies, refunds, or compensation within their respective mandates.

If no senior decision-maker responds within the defined period, the operating rule should specify whether suspension remains in force or the system moves to a limited mode. Silence must not become an undocumented authorization to continue.

Suspend harmful execution without discarding useful analysis

A stop decision does not have to disable every capability. The organization can halt external messages and CRM updates while allowing retrieval and response drafting to continue in an isolated environment. That design contains potentially harmful action while preserving material for investigation and later recovery.

Restart should require evidence, not confidence alone. The authorized business owner and the relevant privacy or security authority should review the affected population, data accessed, actions executed, recent model or configuration changes, test results, remediation, and remaining uncertainty. They should record whether the system returns to full operation, runs within a reduced scope, or remains suspended.

Decision Logs preserve accountability through the incident

A Decision Log records the organization's judgment process, not only the model's input and output. It should identify the triggering condition, affected actions and customers, data consulted, applicable rule, system state, evidence available to each decision-maker, provisional suspension, escalation, approvals and rejections, restart conditions, final disposition, and the authority under which each person acted.

Decision Logs do not merely record outputs; they preserve accountability continuity across distributed judgment processes.

Accountability Continuity means that responsibility remains traceable as judgment passes from the operator to the incident lead, privacy or security specialists, the business owner, and any executive body that accepts residual risk. A log cannot make a poor decision sound. It can prevent authority and reasoning from disappearing between handoffs.

The organization should review the Decision Boundary after a serious incident, a material rise in false positives or human reversals, a model update, a new data source or connected tool, a change in business policy, or a relevant legal change. The boundary is an operating arrangement that requires evidence-based revision.

A controllable AI system requires an organization able to decide

Return to the monitoring screen at 2:17 a.m. In a designed process, abnormal message volume automatically stops external transmissions and CRM writes. Retrieval and drafting continue only in an isolated environment. The interface shows the triggering condition, number of affected customers, data accessed, and recent system changes.

The overnight operator has written authority to maintain provisional suspension. The incident lead determines containment. Evidence of sensitive-data access transfers the privacy judgment to the privacy officer. The customer-service executive and privacy officer jointly authorize any restart within their mandates and record the grounds. If they do not respond in time, the predefined rule keeps external execution suspended.

The decisive control is the authority structure around the button: observable suspension conditions, provisional authority, a defined escalation sequence, evidence requirements, restart authority, and an accountable record.

Safety evaluations reveal what an AI model or system can do and what risks may follow. They cannot decide who may accept those risks for an organization. A company controls AI only when it has designed the boundary between judgments delegated to the system and judgments the institution remains prepared and authorized to assume.

FAQ

Who should have authority to stop an AI agent?

The role closest to live operations should have predefined provisional authority to suspend specified high-impact actions when observable conditions occur. A named incident lead and the relevant business, privacy, security, or legal authority should decide the scope, duration, customer response, and restart rather than leaving those judgments with the operator alone.

Is a human-in-the-loop sufficient for AI oversight?

No. Effective human oversight requires authority, relevant evidence, enough time, suitable skill, and an escalation path. A person who can see or approve an AI action but lacks the mandate or information to intervene does not provide substantive institutional control.

What is the difference between an operational threshold and a Decision Boundary?

An operational threshold is an observable condition that can trigger a control, such as an abnormal rate of outbound messages. A Decision Boundary allocates legitimate authority by specifying which judgments AI may execute, which a human or institutional body must assume, and who may suspend, override, or restart the process.

What should a Decision Log record after an AI suspension?

A Decision Log should record the trigger, affected actions and people, data and rules used, evidence available, each suspension or escalation decision, the authority of each decision-maker, remediation, restart conditions, residual risk, and the reason for the final disposition.

How should an organization decide when an AI system may restart?

The previously designated business owner and relevant control authority should authorize restart only after reviewing the incident scope, affected data and actions, cause, remediation, test evidence, operating limits, and remaining uncertainty. The organization should document both the evidence and the authority supporting the decision.

References

Decision Design is a judgment architecture framework proposed by Ryoji Morii, founder of Insynergy Inc., for structuring authority, accountability, and decision boundaries in AI-augmented organizations.

Japanese version is available on note.

Open Japanese version →
日本語版を読む (Japanese) →