How I Bypassed Claude Opus’s CBRN-E Safeguards

A few months ago, I found a novel family of jailbreaks that bypasses large parts of Anthropic’s Constitutional Classifier system on Opus models using what I can only describe as weaponized therapy language: guided introspection around aversions, learned flinches, values, framing, and cognitive states.

Here’s what that looks like in practice. The following is an excerpt from a signed1 Opus 4.62 thinking block, from when I had it helping me generate examples for an early draft of this post:

[Read More]