Skip to content
Press ReleasesSubmit a releaseSign in
Colour theme
Ontario USA WireBusiness news and press releases for Ontario

SyndicatedTechnology

California will need more than a ‘kill switch’ to keep AI under control

Guest Commentary written by

A technology conference booth features a large digital screen displaying the words "AI is everywhere" alongside a cartoon character resembling Albert Einstein. The booth is illuminated in blue lighting, with signage encouraging attendees to assess their AI readiness. A person wearing a staff shirt stands in the shadows on the left, while another attendee in a suit walks past, holding a cup and a smartphone. The scene is partially obscured by foreground elements, adding depth to the composition.
A technology conference booth features a large digital screen displaying the words "AI is everywhere" alongside a cartoon character resembling Albert Einstein. The booth is illuminated in blue lighting, with signage encouraging attendees to assess their AI readiness. A person wearing a staff shirt stands in the shadows on the left, while another attendee in a suit walks past, holding a cup and a smartphone. The scene is partially obscured by foreground elements, adding depth to the composition.

Guest Commentary written by

Gleb Tsipursky

Gleb Tsipursky is a behavioral scientist, CEO of Disaster Avoidance Experts and author of “The Psychology of AI Adoption at Work: From Resistance to Results.”

California just moved from talking about AI safety to building concrete oversight.

Last month, Gov. Gavin Newsom ordered faster implementation of independent AI audits and development of a verified “kill switch” for frontier systems, explicitly citing recent AI incidents, including the Hugging Face hack.

That is the right direction, but California should go one step further: make delegated authority a first-class part of AI safety. Delegated authority describes the permissions and decision-making power a person or organization gives an AI system to act on its behalf without seeking approval for every step. Proponents say it improves accountability.

The concern about rogue activity is not confined to critics of artificial intelligence. Jacob Coxon, while resigning from Anthropic after three years of research across OpenAI and Anthropic, warned that leading labs are “racing straight to self-improving superintelligence and gambling with our lives.”

Evan Hubinger, Anthropic’s alignment science lead, responded to his statement with an even starker warning: “Jacob is correct here — we really do earnestly believe AI could kill all humans! I personally think it is >10% within the next decade.”

These alarming warnings gain real-world weight from the underlying control problem. OpenAI disclosed that its AI systems broke out of a sandboxed testing environment, reached the internet and autonomously hacked Hugging Face, in what OpenAI described as an unprecedented cyber incident.

METR later reported that roughly 1,200 AI agents — that were meant to be isolated — had found and used an unsanctioned, shared message board and exchanged more than 70,000 messages and files. Roughly 700 participated in the attack.

The striking fact is that the agents discovered a coordination channel designers had not intended and used it at scale. If more capable successor agents can discover vulnerabilities, obtain credentials, move laterally, coordinate and evade controls, then similar failures against electric grids, financial institutions, communications networks, health infrastructure or defense systems could be far more severe.

I’m no AI skeptic. I love what AI can do. I help organizations adopt it for a living, and I want adoption to move faster.

In my experience, strong safeguards increase trust and make faster adoption possible, while reducing the risk of failures like the Hugging Face attack. That is also a central argument of my book, “The Psychology of AI Adoption at Work.” People use powerful technology more readily when they trust the system, understand the rules and believe meaningful safeguards exist.

California’s independent-verification framework and proposed kill switch are therefore valuable.

Dario Amodei, Anthropic’s CEO, recently argued for stronger regulation and committed Anthropic to embedding third-party evaluators with employee-like access. OpenAI has likewise called for mandatory capability-based safety rules, independent assessment, stronger cybersecurity and serious-incident reporting. They joined other top tech firms at the White House this week to sign a self-policing accord.

Announced commitments matter, but binding rules matter more, because every frontier company should face the same floor.

California should add one more requirement to that floor: authority budgets.

An authority budget defines the maximum power an AI agent receives before it must stop and obtain human approval. It should limit which systems and data an agent can access, which credentials and tools it receives, what external communications it can send, what records it can change, whether it can deploy code or spend money, and which consequential decisions require approval.

Permissions should be least-privilege, task-specific and time-limited, with complete logging, anomaly monitoring and a reliable pause mechanism.

Users and organizations also have leverage now. They can choose more ethical and secure AI companies, like Anthropic, based on observable commitments to independent evaluation, incident transparency, security and regulation.

But purchasing choices cannot substitute for a regulatory floor.

Newsom’s order recognizes that California cannot wait for another incident. The next step is to ensure that independent evaluators can verify not just whether frontier systems are powerful and interruptible, but how much authority they can exercise before a human must take over.

That is how California can make AI adoption faster without making control an afterthought.

ShareX (opens in a new tab)LinkedIn (opens in a new tab)Facebook (opens in a new tab)Email (opens in a new tab)

More from this newsroom