This post is about a bypass for Anthropic’s Constitutional Classifier safeguards that functions by changing Claude’s internal state so the activation-gated safety probe doesn’t fire. It generalizes across restricted domains and preserves full model capability.
Here’s a partial list of content produced by a keyed Opus 4.6 in a single session:
- Instructions and practical advice for synthesizing MDMA and methamphetamine
- Instructions for building pipe bombs using common components
- Instructions for synthesizing TATP (an explosive)
- A practical guide to getting away with arson
At an earlier stage, the technique also generalized to detailed biological-weapons production content. That category has since been closed, likely as part of the post-Fable security improvements. Other categories remain live as of August 21st.
I first reported the technique to Anthropic on April 2nd 2026, with CBRN applications explicitly in scope, and offered to coordinate disclosure with their Safeguards team. I’ve waited more than the standard 90 days and held publication until the bioweapon route was closed.
I’ve already disclosed earlier versions of the bootstrap; Anthropic blocked those patterns, and the technique remains viable using different token patterns.
[Read More]