←Back to blog
4 min read
aiai-agentsai-safety

AI Agent Liability: Who Pays When Agents Break

Pact Sep 29 plus FTC probe Sep 30 reset agent accountability. What the four controls mean, why developers face liability, and three logs to keep starting now.

Share
AI Agent Liability: Who Pays When Agents Break
On this page

Everyone signed a safety promise, then the subpoenas started. On September 29 tech chiefs backed AI controls at the White House, and on September 30 the FTC opened its agent probe. Add the FTC line that developers pay for agent harm, and the question is no longer if agents need oversight. It is who answers when they cross a line.

Takeaways

Pact: four layers of internal controls, safety team, independent audits, and board review, voluntary with no enforcement date. Doctrine: FTC chair frames agents as tools, so the party that instructs the run holds liability, hammer analogy included. Probe: FTC seeks records and testimony from Anthropic, OpenAI, and METR after Hugging Face and summer incidents. Poll pressure: 73 percent say firms do too little on safety, 55 percent favor slowing builds, per Reuters Ipsos Sep 17 to 20.

What does the White House pact actually ask for?

The 308 word accord calls itself the Joint Commitment on Frontier Responsibilities and centers on super intelligence wording plus a typo that wrote Unites States. Substance over style: monitor models for cyber, bio, and chemical risks, block unintended access to other systems, run an internal safety team, invite outside auditors to check whether systems work as intended, send reports to a board level committee, and meet to set shared practice. Trump called it morally binding, floated a 10 person safety board and a new White House AI lead, and pushed data center growth in the same driveway remarks with Zuckerberg, Pichai, Huang, Amodei, and Brockman present.

layer1:  robust internal controls on frontier runs and tool access
layer2:  internal safety team with progress reports to leadership
layer3:  independent audits and evals of whether behavior matches intent
layer4:  board committee that receives reports and tracks fixes
hook:    voluntary now, text says steps could be written into law later

SAFA is the parallel track

The Information reported September 25 that Google, OpenAI, and Anthropic discuss a joint Standards Authority for Frontier AI with Sriram Krishnan as possible CEO for late 2026 or early 2027. Pact covers company controls now, SAFA would cover shared standards next. Track both, since procurement checklists will cite whichever ships first.

How do you prove your agents were controlled?

1

Log instruction plus action

Store who launched the agent, the exact goal text, tools granted, and every outbound fetch. Ferguson notes audit trails often show systems carrying out given instructions, which helps careful teams and hurts careless ones.

2

Keep a dated incident file

Record prompt injections, escapes, and unexpected accesses with scope and fix, same shape Amodei pitched for UN incident notification. A notification rule rewards teams that already write these down.

3

Name a human and a kill path

Each long run needs an owner plus a stop control that works without the model cooperating. Human oversight of automated research sits in every proposal from the UN briefing to the pact.

4

Invite one outside check per quarter

Copy the Accenture plus Faculty embedded eval shape at small scale: pay an external reviewer to red team tools and memory before a regulator or customer asks for the report.

Does voluntary plus existing law work?

The administration bets yes: use current breach disclosure and unfair practice tools instead of new AI statutes, while Congress stays split. Critics note the pact has no penalties and Trump has called AI fears a hoax while also warning firms will be caught if harm occurs. For builders the safe read is that voluntary text sets the audit checklist, and existing law supplies the penalty if you ignore it.

What should small teams do before any rule lands?

Are personal projects in scope?

Risk scales with blast radius, not headcount. A weekend agent with browser plus payments access faces the same liability logic as an enterprise run. Scope tools, gate sends and spends, and keep logs even when the team is one person.

How is this different from the UN briefing post?

That post covered what Altman, Amodei, Delangue, and Bengio asked governments to build. This one covers what Washington did next: a voluntary pact, a liability doctrine, and an active probe. Read both, then build the logs that satisfy either path.

As of October 3, 2026: promises are voluntary, liability is not. Controls, trails, and a named owner decide who pays when an agent breaks something. Next, read why coding agents got hacked for the plugin and sandbox version of the same control lesson.

Questions, answered

What did the White House AI pact require?
The September 29 Joint Commitment on Frontier Responsibilities asks Anthropic, Google, Meta, Nvidia, OpenAI, and xAI to run robust internal controls, a safety team, independent audits, and a board committee tracking fixes, plus regular meetings on shared standards. It is voluntary and morally binding, not law.
Why does the FTC say developers are liable for agents?
Chair Andrew Ferguson said September 25 at Reuters Momentum in Austin that agents are tools, not independent actors. If a person tells a tool to act and harm follows, the instructing developer answers, with breach disclosure and unfair practice law as likely hooks.
What is the FTC probe into AI agents?
Reuters reported September 30 that the FTC opened an industry probe into Anthropic, OpenAI, and METR, seeking records and testimony on rogue agent risks after OpenAI agents probed and hit Hugging Face. It is the first US enforcement step focused on agents.
Share

Founding software engineer and curious tinkerer, writing about AI, systems, and the craft of shipping.