Anthropic, the US-based AI company, has disclosed that it thwarted attempts by malicious actors to exploit its models for cyberattacks, surveillance, and biological research that could have aided weapons development. The revelations come in the company's third report on AI misuse, published on Thursday, and highlight the growing challenge of keeping powerful AI out of dangerous hands.
According to Anthropic, the sophistication required for cyberattacks has plummeted as AI models become more capable, meaning even individuals with limited technical skills can now pose threats that were unthinkable a year ago. The company said it has reinforced safeguards in its newest models to restrict access to biological research with potential weapons applications.
“The cases we share here aren't typical misuse, but rather examples of the most notable and novel threat activity we've identified to date,” the report states. It includes excerpts of malicious code and AI prompts, and urges governments and rival AI firms to watch for similar abuse.
“We're publishing this work because we believe we have a responsibility to disclose malicious misuse of our services,” the company said. “As models become increasingly capable, their risks will increase, unless AI developers and society's defenders act to make them safer.”
Blocked attempts to enhance a virus
Between December 2025 and August 2026, Anthropic's researchers identified misuse by a range of actors, from spyware vendors and politically motivated individuals to state-sponsored groups spreading propaganda. Among the most concerning cases, unnamed actors attempted to use Anthropic's models for research that could have led to biological weapons.
In one instance, the company said its systems blocked a request for its Claude chatbot to help draft a grant application for scientific funding. The application involved gain-of-function research on the chikungunya virus, a mosquito-borne pathogen that causes severe pain and fever. The proposal sought to enhance mutations that would make the virus more transmissible and better able to evade the immune system.
Anthropic noted that such research could “certainly” support vaccine and treatment development, but added that “it could also be used to make the pathogen more dangerous.”
Limits of older models
None of the cases involved Anthropic's newer, more powerful Claude Fable or Mythos-class models, with one exception: an “industrial-scale, covert campaign to extract a model's capabilities and replicate them in another model without authorisation.” The company said its older models, including Claude Opus 4 and Claude Sonnet 4.5 from 2025, “were well below the threshold where they could meaningfully assist a sophisticated user in carrying out dangerous biological research.”
“As a result, safeguards on these models were less stringent, directed mostly at preventing access to content that might uplift novices in recreating known bioweapons,” the report said. “But for today's models — which are capable of assisting in a range of complex scientific research tasks — the evidence is no longer certain, and we cannot make that same assurance.”
Consequently, Anthropic has introduced tighter safeguards restricting access to a wide range of dual-use biological research queries in its more recent models, such as Claude Fable 5.
Propaganda and influence operations
The report also details how actors created hundreds of fake social media accounts designed to look like ordinary users, which then posted material amplifying the same political message over the course of a week. Anthropic outlined nine such cases, originating in Russia, Iran, Turkey, and across the Gulf, South Asia, Africa, and Europe.
While social media platforms can detect influence operations once posts are already circulating, Anthropic said it “may see it on Claude while the operation is still being built.” This early detection capability is crucial, as Ukraine's recent sanctions on Kremlin-linked propaganda illustrate the ongoing battle against disinformation.
Researcher's warning and regulatory calls
The report follows the resignation of Anthropic researcher Jacob Coxon, who said he was leaving over fears that the company and its main rival, OpenAI, “are racing straight to self-improving superintelligence and gambling with our lives.” Coxon warned that some of his former colleagues believe AI could threaten human life before the end of the decade. His warning has resonated across the industry.
As AI companies release increasingly powerful models, experts have called on governments to regulate the technology rather than relying on the industry to police itself. John Thickstun, an assistant professor of computer science at Cornell University, said it puts companies such as Anthropic and OpenAI in an uncomfortable position, since they are effectively required to make “value judgements at societal scale without any kind of democratic or deliberative oversight.”
Anthropic said it had blocked each of the malicious activities identified in the report, used the findings to strengthen its safeguards, and shared information with government authorities and industry partners. The company is preparing for an initial public offering this autumn, and its efforts to demonstrate responsible AI development may be crucial. The IPO is expected to be one of the largest tech listings of the year.
“We hope that the findings in this report will help other developers recognise similar patterns on their own platforms, give governments and civil society a clearer view of how emerging threats take shape, and strengthen collective defences,” the company said.


