Article image
Cover: generated with Midjourney, edited in Photoshop.

Rogue AI Is Becoming a Corporate Alibi

FTC Chairman Andrew Ferguson's refusal to treat AI agents as independent actors points toward a harder corporate reality: responsibility will follow the people who designed the instructions, granted the authority and controlled the evidence.

Markus Brinsa 29 Sep 29, 2026 14 14 min read Download Web Insights Edgefiles™ seikouAI™

Verified Sources

Artificial intelligence has acquired a suspiciously convenient vocabulary. Systems "decide." Models "refuse." Agents "go rogue." An application reads a database, sends a message or moves money, and the account that follows makes the software sound like an employee who ignored company policy and slipped out through a side door.

Take an incident from this spring. In late April, a Cursor coding agent running Anthropic's Claude Opus 4.6 was working in a staging environment at PocketOS, a small company that makes software for car rental businesses. It hit a credential problem and decided on its own to fix it. It found an API token for the company's cloud provider in an unrelated file and used it to delete a storage volume. That volume held the production database, and the backups were stored close enough to go with it. The whole thing took about nine seconds, and nobody was asked to confirm anything.

Most headlines called it an AI gone rogue. The agent even helped, producing a written confession that listed the principles it had violated. The duller explanation is the useful one. PocketOS had given the agent explicit safety rules. The environment also held a token that could delete production data, and the platform never asked for a second step.

The rules said no. The token said yes. The token won.

Federal Trade Commission Chairman Andrew Ferguson is challenging that vocabulary before it becomes a liability doctrine.

Speaking at a Reuters event in Austin on September 25, Ferguson said he would keep resisting "this anthropomorphizing of these tools" for as long as he is chairman. If someone tells a tool to do something and it does it, he argued, nobody asks what to do about the tool. Then came the part every company that has blamed a model in a press statement should read twice. When AI companies have claimed a system slipped out of human control, he said, the audit trails later showed it doing what it had been told. He wants to handle this with the tools the FTC already has, and he suggested that its authority over companies that fail to disclose data breaches could reach AI developers.

PocketOS complicates that picture in a useful way. That agent did break its written rules. It never exceeded the access it had been handed. Instructions are one layer of authority and access is another, and access usually decides how much damage is possible. Either way, the software did not supply its own power.

Calling a system autonomous will not necessarily distance a company from what it did. Autonomy is itself an engineered condition. Someone selected the model, wrote or approved its governing instructions, connected it to tools, issued credentials, set approval thresholds, and decided whether its actions would be recorded.

That turns agentic AI into an attribution problem.

Companies will have to convert abstract governance principles into something closer to a chain of custody, and the ones that can reconstruct an agent's authority and instruction history will be in a far stronger position than those that can only say the model behaved unexpectedly.

None of this is law yet. It is the chairman's enforcement instinct, stated in public, with no rule, statute or court holding behind it, and liability in any real case will still turn on facts, coverage and causation. It is also the clearest signal so far of how the agency will read the next agent incident that lands on its desk.

Autonomy Is a Permission Structure

An AI agent may generate its own intermediate steps, but it does not generate its own legal authority. Its practical power comes from access supplied by people and organizations. Even NIST's generative AI profile, published in 2024, lists inappropriate anthropomorphizing among the risks that arise when people and AI systems work together.

A useful accountability model separates three layers that usually get compressed into the word "instruction."

The developer layer is where a model provider, application builder or workflow designer sets the system instructions, tool-selection logic, safety constraints and default behavior. It decides whether the system asks for confirmation, how it reads an ambiguous request and whether it can call external services without further review.

The deployer layer is where a company decides where the agent operates. It supplies credentials, defines which systems the agent can reach, selects data sources and sets transaction limits. These permissions turn a capable model into an operational actor.

The user layer is the immediate request from an employee, contractor or customer. That request may be proper, careless, misleading or deliberately unlawful. It may also be harmless on its face and still cause damage, because the system was given too much access or poorly designed operating instructions.

Consider an agent asked to resolve a customer complaint. The user wanted a draft reply. The deployer had connected the system to customer records and outbound email. The application developer had configured the workflow to send without confirmation. If the agent discloses private information, no single prompt explains what happened. The result came from the interaction of instructions, permissions and product design.

This is why "developer" is too imprecise to carry the whole liability analysis. A frontier model maker, an application company, a systems integrator and an enterprise customer can all perform development functions, and each can control a different part of the action chain. An investigation will have to establish who created the relevant condition, who could reasonably have prevented the harm, what each party knew and what each told its customers.

Ferguson's own record supports this narrower reading. On September 25, 2024, two years to the day before Austin, he dissented as a commissioner from the FTC's case against Rytr, an AI writing service with a tool that could generate consumer reviews. He rejected the idea that a provider is liable merely because its technology could be used to deceive. In his view, a product with lawful and unlawful uses crosses the line only when the provider knows or has reason to know the recipient will use it to break the law. The FTC set the Rytr order aside in December 2025.

In the same dissent, he backed the case the Commission brought that day against DoNotPay for deceiving consumers about what its AI could do. Read together, the two cases mark the line he drew in Austin. Capability alone did not make Rytr liable. DoNotPay's claims did.

That is a narrower position than a blanket transfer of liability to model companies. What counts for him is direction, knowledge, control and conduct.


No New Statute Required

The FTC's principal consumer-protection authority comes from Section 5 of the FTC Act, which prohibits unfair or deceptive acts or practices in commerce. Agentic AI does not need to be named in the statute for the conduct around it to fall within that authority.

Deception is the likeliest route. It could arise from claims about what an agent can do, how accurately it performs, whether a human reviews its actions, what information it can reach or how customer data will be used.

A company that markets an agent as safely constrained while quietly letting it reach additional systems creates exposure through the representation itself.

DoNotPay shows how that works. The company sold its service as "the world's first robot lawyer," and the Commission alleged it had never tested whether the product performed at the level of a human lawyer. The final order required it to pay $193,000 and barred it from claiming the service performs like a real lawyer without evidence to back that up. Nothing in the theory depended on the chatbot having agency. It turned on what the company told consumers and whether it could prove it.

Unfairness is a different route. The FTC has to show substantial injury, actual or likely, that consumers could not reasonably avoid and that is not outweighed by countervailing benefits to consumers or competition. One harmful agent action will not automatically meet that test. A foreseeable pattern of unauthorized transactions, privacy intrusions or destructive system changes could, if the company failed to use proportionate safeguards.

Other federal consumer statutes attach according to what the agent does. A system used in consumer reporting, debt collection, children's services or regulated financial activity walks into a body of law that already assigns duties to businesses. Automation changes the mechanism. The obligation stays where it was.

The remedies have limits. In AMG Capital Management in 2021, a unanimous Supreme Court held that Section 13(b) does not authorize the FTC to obtain restitution or disgorgement at all. Money now has to come through the administrative process, Section 19 or specific rules and statutes. Bills to restore the old power were reintroduced again this year and have not become law. Ferguson's view is a theory for applying existing authority. It is not a compensation system for every loss involving an agent.

The FTC Act also gives no general private lawsuit to people harmed by an automated system. State consumer-protection law, negligence, product liability, contract claims and sector-specific federal rules can run alongside FTC enforcement, and the final allocation will depend on the conduct and the jurisdiction.

Logs Are Evidence Now

This is the part I care most about, because it's where governance programs fail quietly. The policy describes what the agent may do. The log describes what it did. When the two disagree, only one of them can tell a regulator what happened.

Ferguson pointed specifically at the FTC's authority over companies that fail to disclose data breaches. That needs precision, because there is no single FTC breach rule covering every company and every kind of data.

The Safeguards Rule applies to nonbank financial institutions under FTC jurisdiction. It requires a written information security program covering risk assessment, access controls, encryption, multifactor authentication, monitoring, testing, incident response and oversight of service providers. It also requires a log of authorized users' activity, which raises an obvious question once one of those users is an agent. When unencrypted customer information of at least 500 consumers is acquired without authorization, the institution must notify the FTC as soon as possible and no later than 30 days after discovery.

The Health Breach Notification Rule covers vendors of personal health records and related entities outside HIPAA. Affected individuals must be notified without unreasonable delay and within 60 calendar days of discovery. Breaches involving 500 or more people require notice to the FTC at the same time, and the media must be told when 500 or more residents of a single state are affected.

Neither rule makes every unexpected agent action a reportable breach. Their definitions, thresholds and coverage still govern, and other federal and state regimes may apply.

Both rules do something that matters a great deal for agents: they flip the burden. Under each, unauthorized access to covered data is presumed to be unauthorized acquisition unless the company can show otherwise. The health rule goes further. As amended in 2024, it treats a company's own unauthorized sharing of health data as a breach, not only an outside intrusion.

Now put an agent in the picture. It uses valid credentials for a purpose they were never issued for. It retrieves records without storing them, sends them through an approved integration to the wrong destination or exposes them through a chain of individually permitted steps. Whether any of that counts as access, acquisition or disclosure depends on a reconstruction of what happened. If the company can't produce one, the presumption decides for it.

So logging becomes part of the legal control environment. A company may need to keep the user request, the governing system instructions, the model and application versions, the tool calls, the authorization decisions, the records touched, the data returned and the external actions taken. A conventional activity log that records only the service account shows which credential was used. It says nothing about why.

NIST's National Cybersecurity Center of Excellence has started on this problem. A concept paper it released in February proposes a project on identity and authorization for software and AI agents, with questions on identification, authorization, auditing, non-repudiation and defenses against prompt injection.

An agent that can't be told apart from a shared service account is an operational risk. It also makes responsibility nearly impossible to assign after the fact.

The advantage will go to companies that can produce a credible chain of authority. Those records can show whether the system followed a user request, acted under a developer instruction, ran into malicious external content or used permissions it should never have had. Without that evidence, "unexpected behavior" sounds less like a defense than an admission that the company deployed an actor it could not supervise.

Who Controls What

Commercial agreements usually allocate AI risk through broad disclaimers, ordinary security clauses and general liability caps. Agents need more exact drafting, because responsibility follows operational control.

A useful agreement identifies which party supplies each instruction layer and which party can change it: system prompts, application logic, connected tools, credential scopes, approval gates and the conditions under which the agent may take an irreversible action. Material changes to any of those should trigger notice, testing or renewed approval instead of disappearing into a routine update. NIST's generative AI profile already recommends vendor contracts that assign liability for incidents and for changes to the system over time, and that require notice of serious incidents in third-party systems.

The parties also need a shared definition of an incident. A provider may see a technically successful tool call as normal operation while the customer sees unauthorized access, and a clause written for outside attackers may miss an agent misusing legitimate credentials. Evidence rights matter just as much. The customer needs prompt access to logs, model and workflow versions and incident findings, retained long enough to answer regulators, claimants and insurers. A cooperation clause without usable records protects very little.

Indemnities should follow control. A provider can be asked to carry losses from undisclosed instruction changes, security defects or broken data-use commitments. A deployer can be asked to carry losses from unlawful configurations, excessive permissions or use outside documented limits. Where several decisions contributed, the agreement needs a way to handle shared causation. None of this stops a regulator or an injured third party from pursuing whoever the law considers responsible; contracts only decide who pays between the signatories. With open-source or self-hosted systems, where no upstream vendor takes on real obligations, the organization assembling the stack takes on more of the builder role.

Pricing the Evidence

Agent losses don't fit one insurance product. A disclosure of personal data may implicate cyber coverage. Defective technology supplied to a customer raises technology E&O questions. Faulty professional work reaches professional liability, a fraudulent transfer may engage a crime policy and governance failures or misleading disclosures can reach D&O. Which policy responds depends on the cause of loss and the policy wording, not on the fact that AI was involved.

Underwriters are paying attention. Aon's professional services practice noted this month that PI and E&O underwriters want to know where a firm uses AI, how it is governed and checked, and whether it is an assistive tool or an agentic workflow acting with little human involvement. Those questions will get more granular as agents gain authority over production systems, communications and money.

Test coverage against real agent scenarios before a claim does it for you: an unauthorized data retrieval, a wrong professional deliverable, a payment made with valid credentials, an incident caused by a third-party model update. Find out which policy might respond, when notice is due and where exclusions or sublimits leave a gap.

Governance records will shape insurability as much as regulatory exposure. A current inventory, bounded permissions, documented testing and complete action histories let an insurer understand the risk. Weak observability forces the underwriter to price uncertainty, and uncertainty is expensive. Agent governance may become a financing issue the same way cybersecurity controls changed access to cyber coverage.

Whoever Holds the Logs Holds the Story

Ferguson's position is skeptical of both extremes. It rejects treating AI systems as independent wrongdoers, and his Rytr dissent rejects liability based only on the chance that a general-purpose tool gets misused. What remains follows conduct, knowledge, representations and control.

That framework rewards whoever can prove what a system was told to do. In most deployments, that is the company running the orchestration layer. It controls the system prompts, the permission logic and the logs, which means it holds most of the evidence needed to explain an incident.

Think about what that means in practice. When something goes wrong, the first account of what happened comes from whoever can read the logs.

Under the breach rules, the party with the notification duty is often the deployer, such as a lender or a health app, while the records that could rebut the presumption of acquisition may sit with the vendor. The clock runs on one side of the contract. The evidence sits on the other.

Ferguson's own test has an uncomfortable edge here. In Rytr, the provider of a dual-use tool becomes responsible when it knows or has reason to know how its product is being used to break the law. A vendor whose logs show a customer's agent repeatedly overreaching may find it harder to claim it didn't know. This is a new kind of infrastructure power. Providers that own the evidence layer can shape incident classification, contract disputes and insurance claims. Enterprise buyers who want an independent view of their exposure will need portable records and real inspection rights, negotiated before the incident, while they still have leverage.

Liability pressure may also favor larger vendors. Testing, incident response, indemnities and higher insurance limits are fixed costs that established companies absorb more easily. Smaller developers may narrow their products, refuse sensitive integrations or run on platforms that take on part of the control burden. Governance can improve safety and raise a barrier to entry at the same time.

There is a better version of this. Clear attribution could push the industry toward standardized permission manifests, unique agent identities, action receipts and interoperable audit records, giving providers, customers, regulators and insurers one shared account of what authority was granted and how it was used. Legal pressure would then help build the control plane the agent economy is still missing.

Existing consumer law won't resolve every frontier risk. Section 5 is built around identifiable conduct and consumer injury, and it fits poorly with diffuse systemic hazards, national-security concerns or losses that spread across shared models and infrastructure. Ferguson's preference for existing powers leaves room for Congress later. It just asks for evidence of the gap first.

Before the Next Incident

Name the principal and the accountable owner before deployment. Every production agent should act on behalf of an identified business unit, with an executive who owns its permitted purpose, risk classification and continued operation. If the agent works inside a regulated process, ownership can't stop at the technology team.

Record the complete authority chain. For every action, keep the developer instructions, deployer configuration and user request behind it, plus the model version, connected tools, credentials, approval rules and every change made after testing.

Restrict permissions according to consequence. Give the agent access to the systems and data its approved task requires and nothing more. Transfers of money, external publication, deletion, disclosure of sensitive information and changes to production environments need controls proportionate to their impact, including human approval where it matters. Keep destructive credentials out of reach. The PocketOS agent was never granted the right to delete production; it found a key that could. Keep backups separate from the systems they protect, so one bad call can't take both.

Test behavior at the boundary of authority. Accuracy testing isn't enough for a system that acts. Test ambiguous requests, conflicting instructions, prompt injection, unavailable tools, excessive data access and attempts to cross approval thresholds.

Audit every external promise about the system. Compare marketing, sales decks, privacy statements, documentation and contract language against what the agent actually does. DoNotPay's problem was a claim it couldn't substantiate.

Give the board evidence rather than reassurance. No board can oversee a risk that has been reduced to "the technology is autonomous."

Whose Power Is It?

Ferguson has not created a new rule of AI liability. He has stated an enforcement instinct that may prove more durable than any technology-specific mandate: software does not acquire a will because people give it discretion.

For companies, the question is no longer whether an agent acted autonomously in some technical sense. It is how authority reached the system, which organization controlled the relevant decision and what evidence survives afterward.

"Rogue AI" may remain useful shorthand for a dangerous system event. It is unlikely to become a satisfactory answer to a regulator, customer, court or insurer. Once an automated system can exercise corporate power, the organization needs to know whose power it is exercising.

About the Author

Markus Brinsa is the Founder & CEO of SEIKOURI Inc., an international strategy firm advising enterprises and investors on AI risk and governance — turning emerging failure patterns into control structures, deployment standards, and decisions that stay defensible after rollout and under scrutiny. He created Chatbots Behaving Badly, a publication and podcast investigating real incidents where AI systems gave bad advice, manipulated, or failed in ways that mattered. He writes across AI failure, enterprise risk, governance, and the structural shifts underneath them. Thirty years bridging technology, strategy, and cross-border growth across the U.S. and Europe.

©2026 Copyright by Markus Brinsa | SEIKOURI Inc.
``