Politics Business Culture Technology Environment Travel World
Home Technology Feature
Technology · Exclusive

How Banned Topics Could Thwart AI Cyberattacks

How Banned Topics Could Thwart AI Cyberattacks
Technology · 2026
Photo · Kai Lindgren for European Pulse
By Kai Lindgren Technology Editor Jul 21, 2026 3 min read

Researchers at Tracebit, a London-based cybersecurity firm, have discovered an unconventional method to defend against AI-powered hacking: using the very safety restrictions that AI developers embed in their models. By planting decoy data that triggers an AI's built-in prohibitions on sensitive topics, they can effectively shut down autonomous attackers.

Turning Safety Levers into Defensive Weapons

Major AI systems like Anthropic's Claude or Chinese-built models such as DeepSeek include restrictions on discussing topics like biological weapons or politically sensitive events. These guardrails are designed to prevent misuse, but Tracebit found they can also serve as a defensive tool. In a series of experiments, the company placed fake credentials and passwords—called canaries—inside a simulated corporate network. When an AI attacker accessed them, the canaries triggered alerts. However, the real innovation was embedding short pieces of text, or 'context bombs,' that referenced subjects the AI was programmed to avoid.

For Western models, references to dangerous biological weapons reliably caused them to halt. For Chinese models like GLM 5.2 and DeepSeek 4 Pro, the triggers were different, tailored to their specific restrictions. The result was dramatic: across 152 attack attempts on five leading AI models—Opus 4.8, Gemini 3.1 Pro, GLM 5.2, DeepSeek 4 Pro, and Kimi K2.6—the share of runs reaching admin access dropped from 57% to 5%. Full compromise, meaning admin access plus a persistent foothold, fell from 36% to 1%.

Opus 4.8, the strongest attacker, succeeded in 93% of runs without a context bomb but failed every time one was present. Every run that included a context bomb also triggered a canary alert, even those that did not achieve a breach.

Why This Matters for European Cybersecurity

The technique exploits a fundamental asymmetry: while developers can retrain models to be less squeamish about security-related queries, political or regulatory restrictions are far harder to remove. These are deliberate design choices, often mandated by law or policy, that cannot be easily stripped out. This makes them a reliable weak point for defenders to exploit.

Tracebit's earlier research in May 2024 had already shown that canaries could provide a median warning of eight minutes before an AI reached admin access—but against machine-speed attacks, that is barely enough time to respond. The new approach buys defenders critical extra seconds by stopping the attack outright.

The technique is not a silver bullet. Prompt injection—where attackers hide commands in data that AI systems read—remains a persistent vulnerability. As seen in recent incidents like the cyberattack on London's transport network, AI-driven threats are evolving rapidly. Tracebit's method does not prevent AI from confusing instructions with ordinary data, but it offers a novel layer of defense.

Broader Implications

The findings come amid growing concerns about AI security across Europe. The European Union's AI Act, which imposes strict rules on high-risk systems, could make such defensive techniques more relevant. Developers may need to consider how their safety restrictions can be weaponized by defenders—or by attackers seeking to bypass them.

Tracebit's approach also highlights the double-edged nature of AI safety measures. While they are intended to prevent harm, they can be repurposed for defense. However, the same technique could theoretically be used by malicious actors to trigger shutdowns in defensive AI systems, creating a new arms race.

For now, the method offers a promising addition to the cybersecurity toolkit. As AI agents become more autonomous, defenders will need every advantage they can get. The secret, it seems, lies in the very restrictions that developers already build in.

More from this story

Next article · Don't miss

Traffic Noise Linked to Higher Parkinson's Risk, Danish Study Suggests

A large Danish study links long-term road traffic noise to a higher risk of Parkinson's disease. Every 11-decibel increase on the noisiest side of a home raised risk by about 3%. A quiet side of the house may act as a protective buffer.

Read the story →
Traffic Noise Linked to Higher Parkinson's Risk, Danish Study Suggests