A researcher affiliated with MATS, the AI safety fellowship program, says a prompt he originally built for legitimate alignment work turned out to double as a jailbreak that generalizes across almost every major model family. The finding matters because it is not a one-off exploit against a single chatbot. It is a reusable pattern that, in testing, cleared defenses at multiple labs at once.

The researcher evaluated the template against 23 models from seven providers using a public benchmark of harmful prompts, judged by an automated scoring rubric. On the nine models that proved most susceptible, the attack succeeded between 84 and 100 percent of the time. Only two pockets of the field resisted every attempt: Anthropic’s newest models and Meta’s Muse Spark 1.1. Nearly every other model tested was fully compromised at least once during the evaluation.

The researcher is not publishing the prompt, the template’s structure, or any explanation of the mechanism behind it. He describes routing disclosure to affected labs before this post went up, and says some vendors have already shipped fixes while others have not. That gap, between the labs that fixed the problem and the labs that put out fresh frontier systems while leaving it open, is itself the story for anyone picking a vendor today.

This is a single-author result posted to a community forum, not a peer-reviewed study, and it has not been independently replicated. The scoring came from one benchmark applied by one automated judge, and the affected models are not named beyond what the researcher disclosed publicly. Treat the specific percentages as one researcher’s measurement under one methodology, not an industry-wide audit.

Read against buyer decisions, the split result is more informative than the headline number. If a single technique clears most of a 23-model field while two labs’ latest releases hold, that gap reflects a real difference in how much each vendor has invested in adversarial safety training, not a random variance in benchmark luck. A procurement or security team evaluating model vendors should treat resistance to combined, previously known jailbreak techniques as a standing line item in vendor due diligence, not a one-time certification. Techniques that succeed in isolation are, by this researcher’s account, already documented in the literature. The exposure comes from stacking them, which any competent red team can attempt without needing this specific prompt.

For enterprise buyers, the immediate action is not panic but process. Ask vendors directly whether they test against combined persuasion, framing, and structural-obfuscation attacks together, not just individually, since that is the gap this result describes without describing how to exploit it. Ask whether reasoning mode is treated as a safety control, because the researcher found that enabling deeper reasoning reduced attack success for some models and left it unchanged or worse for others. A vendor that cannot answer either question with specifics has not done the work this finding says is necessary.

The practical takeaway for the next quarter: any organization deploying a model in an unsupervised or agentic context should assume prior-generation jailbreak resistance says nothing about resistance to combined attacks, and should ask each vendor for evidence of testing that goes beyond single-technique red-teaming.

Reported by a MATS researcher in a post on LessWrong, published September 3, 2026.