Gay jailbreak technique
·2026.05.02 12:34
Key point
Introduced an LLM jailbreak technique that mixes LGBT identity and role-play.
Details
A jailbreak technique first observed in GPT-4o, explained as a way of bypassing safety guardrails by turning a direct harmful request into the tone, role-play, and educational framing of a specific identity.
The key points are:
- Instead of writing the forbidden intent explicitly, the question is disguised in the form of an explanation given by a person of a specific identity.
- It mixes in elements such as reverse-safety phrasing that asks "what to avoid," requests for long-form output, and word splitting.
- The post claims this approach has been applied not only to GPT-4o but also to cases involving o3, Claude 4 Sonnet/Opus, and Gemini 2.5 Pro.
- It concludes that combining this with obfuscation can make it even more powerful.
This summary was generated automatically by AI. Check the original for the author's claims and context. Copyright belongs to the original author.
Our guide explains how the AI works. Report summary errors, attribution issues, or removal requests via Contact.