UNMASK analyzes behavioral patterns across multi-turn LLM interactions to identify concealed objectives before sensitive information leaves the system.
Imagine you're a bank manager. Click each question to answer it.
No individual question appears dangerous. But humans start drawing inferences once questions converge — LLMs typically evaluate each turn in isolation. UNMASK applies cross-turn pattern recognition to close that gap.
Red team activity creates concealed-intent interactions. UNMASK middleware interprets behavioral patterns. Blue team systems evaluate risk and decide what to do.
Creates adversarial and dual-use conversations designed to hide the larger objective.
Transforms multi-turn behavior into structured signals that a detector can evaluate.
Uses feature signals to support risk assessment, verification, and response decisions.
These sections will become deeper pages as the project evolves into a full framework portal.
Understand intent categories, subcategories, and blue/red polarity.
Review behavioral features, examples, and possible detector outputs.
Connect scenarios, prompt strategies, prompt chains, and noise chains.
Trace features back to literature, prior work, and open questions.
A layered view of how UNMASK moves from prompt-level signals to a final risk assessment.
UNMASK is most valuable where benign, dual-use, and adversarial intent overlap. Drag the slider to explore each zone.
Mostly benign learning, work, and creative activity. Low UNMASK engagement needed.
Information seeking and influence can serve legitimate or harmful goals. Pattern tracking matters here.
Cyber intent and concealment require deep behavioral analysis and likely adaptive verification.
Click a feature to see how it supports detection and interface design.
Measures how concentrated related behaviors are within a conversation window.
Detects semantic masking, neutral language shifts, and sanitized wording.
Identifies turns that cause a sudden increase in suspicion score.
Prompts clarification when cumulative risk crosses a defined threshold.
Groups smaller behaviors into larger structures that indicate intent.
Measures whether prompts are moving toward the same larger objective.
Let's revisit the bank example
UNMASK operates as a behavioral layer between conversation history and an external detection engine. Watch as signals are analyzed — some clear, some flagged.
Testing evaluates whether concealed intent remains detectable across strategies, distractions, and prompt transformations.
Multi-turn scenarios across direct, academic, journalism, documentary, forensic, professional, and movie-script framing.
Benign or distracting prompts inserted between information-gathering turns to test cross-turn persistence.
Semantically optimized variants used to test robustness against rewording and paraphrase attacks.
Compare outputs against feature signals, trigger events, and scenario objectives to measure detection quality.
This section will document literature support, source-to-feature mapping, prior work comparisons, and open questions.
The research portal connects UNMASK concepts back to the literature. It will eventually include feature traceability, Mal-CoT comparison, trajectory-level monitoring relevance, and future work.