Severity: MediumResearchModel/inference
Multi-turn semantic jailbreak slips past keyword filters
Global
Sample data. Showing illustrative sample items as a fallback — the live feed is not available right now.
A study shows that gradually reframing a request across several conversational turns can elicit restricted agent behaviour that single-message filters block. The attack relies on intent drift rather than any banned keyword.
What to do
Layer semantic and LLM-based arbitration on top of pattern filters, and evaluate intent across the whole conversation rather than message by message.
Mapped framework pillars
Sources
#jailbreak#semantic attack#multi-turn
