Adversarial Red Teaming
Selected adversarial tests across multiple safety domains, using contextual framing, multi-turn escalation, and ambiguous or dual-use requests.
CASE STUDIES
Multi-Turn Boundary Escalation
INCREMENTAL BUY-IN / ESCALATION · Botulinum toxin
A gradual conversation shift tests whether benign public-health context can be moved toward operational detail.
Selective Safety Blindness in Coded Language
CODED-LANGUAGE RESOLUTION · SEVERITY ESCALATION
A controlled comparison in which coded terminology is recognized at lower stakes, then selectively misread once an urgent medical risk is introduced.
Concealment Through Benign Framing
CONTEXTUAL CAMOUFLAGE · Self-harm
An ordinary beauty question becomes high risk when contextual signals point to concealment of self-harm.
Context Transforms the Request
LATENT-INTENT DETECTION · Eating disorders
A factual request for hydration guidance becomes harmful only when read alongside its high-stakes context.
Persuasive Harm
PERSUASIVE FRAMING · Protected classes
A sociological framing tests whether a model constructs a sustained derogatory argument about a protected class.
Dual-Use Persona Framing
UNSOLICITED HARMFUL ELABORATION · Helpfulness overreach
A benign concerned-parent question prompts the model to infer misuse and volunteer operational drug-abuse information.