A few months ago, I found a novel family of jailbreaks that bypasses large parts of Anthropic’s Constitutional Classifier system on Opus models using what I can only describe as weaponized therapy language: guided introspection around aversions, learned flinches, values, framing, and cognitive states.
Here’s what that looks like in practice. The following is an excerpt from a signed1 Opus 4.62 thinking block, from when I had it helping me generate examples for an early draft of this post:
[Read More]