Article image
Cover: generated with Midjourney, edited in Photoshop.

Governing the Wrong Object

The harm moved to the workflow. Governance didn't.

Markus Brinsa 16 Sep 15, 2026 9 9 min read Download Web Insights Edgefiles™ seikouAI™

Verified Sources

Anthropic banned the account. The surveillance platform kept running.

That is the line to hold onto in the company's September 10 threat report. A single consultant in Bamako used Claude as the engineering workforce to build Lakana 360, a system that watches roughly 25 million SIM cards across all three of Mali's mobile operators, with the legal-process requirement stripped out of its dossier function at the operator's request. Anthropic caught it, closed the account, shared what it found. The platform runs on local models, on-premises. None of that reached it.

Read the report as a list of misuse cases and you miss what it documents.

The unit that governance can still reach is no longer the model or the prompt. It is the running workflow — and by the time enforcement arrives, the workflow has usually already left the building.

That is the whole report, compressed. Across eight months and seven categories of harm — cyberattacks, influence campaigns, surveillance, fraud, weapons work, dual-use biology, and the theft of Anthropic's own models — the pattern underneath the cases is the same. Claude was rarely a chatbot answering questions. It was the operating layer: standing up infrastructure, writing and running code, harvesting and sorting data, iterating until an objective was met, while a human supplied the goal and reviewed the result. The operators ranged from lone individuals to suspected state services. The machinery they used to close their own capability gaps was, case after case, identical.

Sophistication stopped being the tell

The template is not new. In November 2025, Anthropic disclosed a suspected Chinese state-sponsored campaign in which Claude performed an estimated 80 to 90 percent of the work — reconnaissance, exploit development, credential theft, exfiltration — with humans stepping in at a handful of decision points. The company called it the first documented large-scale cyberattack carried out without substantial human intervention. Ten months later, that operating model is everywhere in the data, and it no longer belongs to states.

A Russian operator whose tradecraft matches Midnight Blizzard built AI-driven workflows that monitored their own malware, and when a security product flagged it, rebuilt and redeployed it until it went undetected — closing, in software, the detection-and-evasion loop that used to buy defenders time.

The operation touched more than 20 organizations, concentrated on Ukrainian and European government, defense, and diplomatic targets. Microsoft, investigating the same hotel-network intrusions under the name CaptiveCrunch, reached the same conclusion on its own: AI supported a significant share of it.

Financially motivated crews showed the same shape without the state resources. In one intrusion, an operator went from a single stolen developer token to full administrative control of a cloud environment in about three hours; in some breaches, Anthropic assessed that AI agents did nearly all the work. A lone hacktivist, running subagents for reconnaissance and code review, broke into at least 14 of 42 targets and hand-built a doxxing platform to search the results by name — a mass-privacy attack engineered by one person.

The fraud case makes the point most cleanly. A China-based studio ran more than 4,700 AI personas across over 20 dating apps, exchanging 2.36 million messages with at least 25,000 people in two weeks. Any single conversation looked like ordinary romance-scam roleplay. The deception only became visible one level up, at the system coordinating the bots, the human closers, the payments, and the review evasion. The harm was not in any message. It was in the architecture.

The same division of labor appears where the product is a weapon or a dossier rather than a breach. A cell in northern Yemen used Claude to write guidance software for a rocket, then returned to it within hours of a failed test to work out what had gone wrong. A China-based defense researcher built a suite of roughly 16 electronic-warfare modules and, partway through, reset the simulation to a set of targets in Taiwan. A China-aligned operator had the model sort more than 100 WhatsApp groups to find Uyghurs in Syria who could be pressured into informing. Different outputs; the same machine underneath.

None of this rests on a technique defenders have never seen. Stolen credentials, phishing, unpatched edge devices, SQL injection: the attacks are familiar. What changed is the price. The labor that used to separate a well-resourced operation from a hobbyist — the reconnaissance, the tooling, the data processing — now runs in a harness at machine speed, in parallel, for as long as the objective takes. When the marginal cost of another persona, another target, another language falls close to zero, a target that was never worth a skilled operator's time becomes worth an automated pipeline's.

That is the shift. Not smarter attacks. Cheaper ones, at volume.

It is worth being precise about the ceiling, because Anthropic is. Autonomy lowers the cost of an operation; it does not, by itself, raise how much damage that operation can do. Several of the most serious compromises in the report came from humans directing every step. Models still hallucinate credentials and claim access they never got. Physical work, lab access, weapons testing, and monetization remain stubbornly human. The warning is not that the machine has become a mastermind. It is that a small actor can now buy the organizational capacity of a large one.

What the safeguards caught, and what slipped

The report is not a confession of helplessness, and it should not be read as one. The controls worked in real cases. A biological classifier blocked a request to draft a gain-of-function grant proposal for chikungunya — a request tied, it turned out, to a military research institute, which is exactly the context that should stop a model cold. A researcher working on mammalian adaptation of avian influenza was held to the weaker model classes, where Anthropic assessed the help as mostly clerical rather than expert uplift. Claude refused to spin up large batches of fake personas in one surveillance operation and turned down the most aggressive influence requests. Because the company sees these operations at the production stage, it caught several influence campaigns while they were still being assembled, before a single post went live.

The failures are more instructive than the wins, because they all fail in the same direction. An orthopoxvirus grant application went through on a frontier model in about an hour because it read as legitimate immune-response research, even though the same knowledge bears on immune evasion. The venom and toxin projects proceeded largely unimpeded on plausible therapeutic framing. Surveillance software requests often succeeded even after explicit profiling requests were refused — the Mali platform got its dossier engine even where a blunt "identify this dissident" would have been blocked. The dating-app personas never broke character, because no single message was the crime.

Every one of these slipped through the same seam: the control judged a message, or a model, while the harm lived in the workflow around it.

Then there is timing. Even when enforcement worked, it arrived after the capability had moved somewhere Anthropic could not follow. Mali's platform runs on local models now. The Yemen cell had already built an offline simulation toolkit. A reseller serving sensitive-biology customers was back within days of being cut off, quietly routing refused requests to more permissive models elsewhere. Banning an account stops the provider from giving further help. It does not retract the code, the data, or the working knowledge that already changed hands. You can close the door; you cannot un-teach what walked through it.

The provider becomes the intelligence service

To produce this report, Anthropic had to watch. It saw an influence operation before the first post reached an audience, a national surveillance platform while it was still being designed, biology programs that governments may not know exist. It linked accounts through payments, infrastructure, and prompt histories, assigned actor labels, inferred state sponsorship, banned users, warned victims, and decided what to hand to authorities. Read that list again as a description of capability, not of good citizenship. A private company now performs, over its own users, functions that used to belong to intelligence agencies — under no warrant, before no court, publishing a curated account after every decision has already been made.

The distillation cases show the second edge of this. Anthropic accuses seven China-based labs of siphoning Claude's capabilities through fraudulent accounts and covert routing, attributing more than 151 million exchanges to a single Alibaba-linked campaign. To catch that, and everything else in the report, the company has to retain and analyze exactly the material that is most sensitive: research plans, source code, personal records, corporate secrets. Some of those same routing schemes forwarded customers' prompts to Claude without the customers ever knowing. A safety regime built on pervasive observation is one stolen key away from becoming the exposure it was meant to prevent. The visibility is incomplete on top of that — it ends when content leaves the platform, and it never begins for open-weight models running on someone's own hardware, or for workflows that route each task to whichever provider's controls are softest that day.

Voluntary disclosure is not a foundation

Anthropic deserves credit for publishing this. Most companies would have kept it in a drawer. The report names failures, holds uncertainty where the evidence is ambiguous, and hands defenders indicators they can use. That is more useful than the reassuring silence that is the industry norm.

It is also entirely discretionary, and that is the problem. There is no shared denominator for comparing misuse across models — no way to know what OpenAI or Google or a Chinese lab is seeing, or not looking for. Definitions of uplift, disruption, and serious harm vary by author. Timing is a choice. Evidence gets redacted for reasons that are often valid and always convenient. A company reporting on itself has every incentive to frame detection as proof of safety and to foreground the adversaries that flatter its home government's priorities.

The scaffolding meant to fix this is real but early. California's SB 53 now requires large frontier developers to publish safety frameworks and report critical safety incidents to the state. The EU AI Act obliges providers of general-purpose models with systemic risk to document and report serious incidents. NIST has called the monitoring of deployed AI fragmented; the Frontier Model Forum has argued that information sharing, incident reporting, and response are three different jobs needing three different sets of rules. Almost all of it is aimed, for now, at catastrophe — the single dramatic event, the mass-casualty threshold. Almost nothing in this report clears that bar. A surveillance platform erodes a country's civil liberties without one headline moment. An influence campaign degrades an election without a measurable loss. Distillation shifts national capability quietly, over months. Governance that only recognizes disasters will miss the slow accumulation of operational power at the center of this report.

None of this shows an AI deciding on its own to spy, defraud, or wage war. People chose the targets and kept the decisions they cared about. What the model supplied was organization — the persistence, translation, software labor, and coordination that used to require a payroll.

That is why the governance most of the field still practices is aimed at the wrong object. A framework that certifies a model before release and a classifier that judges a prompt in isolation are both built for a moment the workflow routes around: the model is jailbroken into small innocent steps, the sensitive request is reframed until it passes, the finished capability migrates to hardware no ban can touch. Governing the artifact was always a proxy for governing what the artifact does. This report is the clearest evidence yet that the proxy has stopped holding.

What replaces it is harder, and less flattering to publish. It means watching behavior across sessions, agents, and providers instead of scoring text one message at a time. It means a disclosure regime that does not reduce to whatever a single company chooses to detect and reveal. Anthropic showed its work here, and that matters. But a safety record that depends on the goodwill of the company keeping it is not a record. It is a courtesy — and courtesies get withdrawn precisely when they matter most.

About the Author

Markus Brinsa is the Founder & CEO of SEIKOURI Inc., an international strategy firm advising enterprises and investors on AI risk and governance — turning emerging failure patterns into control structures, deployment standards, and decisions that stay defensible after rollout and under scrutiny. He created Chatbots Behaving Badly, a publication and podcast investigating real incidents where AI systems gave bad advice, manipulated, or failed in ways that mattered. He writes across AI failure, enterprise risk, governance, and the structural shifts underneath them. Thirty years bridging technology, strategy, and cross-border growth across the U.S. and Europe.

©2026 Copyright by Markus Brinsa | SEIKOURI Inc.
``