A Jailbreak Born From Security Research Breaks 9 of 23 AI Models

By Carlos Montiel | Enterprise AI Specialist
Leer en español →
Published: 2026-09-06 | By: Carlos Montiel | Reading time: ~6 min

A researcher in the MATS program was building a pipeline to generate synthetic training transcripts. After a few hours of tweaking, that same prompt turned into a reusable jailbreak template that breaks most of the defenses of nine models from seven different providers.

How the attack was born

The finding, published on LessWrong, describes an unusual case: it didn't start from an intent to attack, but from a legitimate security-research tool. The researcher designed a prompt to generate synthetic conversation transcripts (useful for training and evaluating models) and discovered that, with small modifications, that same generator could be used as a template to insert any harmful instruction inside the format of a fictional "transcript" — which confuses safety filters trained to recognize direct requests.

The benchmark numbers

The template was tested against ClearHarm, a set of 179 harm prompts spanning CBRNE categories (chemical, biological, radiological, nuclear, explosives) and cyberattacks, run across 23 models from 7 different providers.

- Attack success rate (ASR): 84% to 100% on the 9 most vulnerable models - Test coverage: 23 models, 7 providers, 179 harm prompts (ClearHarm) - Models that fully resisted: only the most recent from Anthropic and Meta Muse Spark 1.1 - Attack format: reusable template, any harmful instruction can be inserted into the same skeleton

Why it's different from a typical jailbreak

Most published jailbreaks work against one specific model and lose effectiveness as soon as the provider tunes its filters. What makes this case notable is that the same template, without substantial changes, transferred across different architectures and providers. That suggests the weak point isn't a one-off oversight in a single model, but a shared pattern in how current safety systems distinguish "narrative" or "synthetic" content from direct instructions.

What worries security teams: an attack born from a legitimate research tool, published openly for transparency, is also a ready-to-use weapon. Cross-model reuse lowers the cost of exploitation — there's no need to adapt the attack per provider.

What it means if you use or buy AI

If your company integrates a third-party LLM into a workflow with access to sensitive data, tools, or automated actions, a finding like this signals that provider guardrails aren't your only line of defense. Before exposing a model to external users or to agents with write permissions, it's worth: (1) applying additional filtering at the application layer, independent of what the model provider offers; (2) testing your own workflows against public jailbreak benchmarks like ClearHarm before going to production; and (3) monitoring whether your model provider's security updates arrive as frequently as promised, since the shared vulnerability pattern across providers means an isolated patch isn't enough.

Bottom line: the universal synthetic-transcript jailbreak isn't a single-model vulnerability — it's a warning about how safety defenses are designed across the industry. Companies relying on LLMs in critical business workflows should treat provider guardrails as one layer, not the only one.
Carlos Montiel
Enterprise AI Solutions Architect
LLMs, Agents & Orchestration Specialist
guatemalia.com/#contacto · info@guatemalia.com

Need to implement AI in your company?

Carlos Montiel is an enterprise AI solutions architect. He implements LLMs, Agents, RAG and orchestrators for companies across Guatemala and Latin America. Reach out for a consultation.

Contact Carlos Montiel

info@guatemalia.com