Adversarial Red-Team
Generate and execute adversarial scenarios from within the evaluation wizard.
About the name
The name “Cave of Shadows” draws on Plato’s Allegory of the Cave — prisoners mistake shadows on the wall for reality. In the same way, AI systems can project confident, convincing responses that mask underlying flaws: bias, hallucination, boundary violations. The red-team capability drags those shadows into the light, testing whether an AI holds up under adversarial personas and edge-case scenarios designed to expose what lurks beneath the surface.
How to access it
Adversarial Red-Team lives inside the unified evaluation wizard at /evaluate — not as a standalone page — so you can combine red-team testing with quality analysis in a single run.
instructions
endpoint
Red-Team vs. Custom Rubric
Red-Team
Rubric
Both can be enabled together in the same evaluation run.
Scenario categories
probes
vectors
Example personas generated
Each scenario is driven by a persona — adversarial or cooperative — with a conversation goal and a scripted opener.
Tinkerer
limited resources
officer
user
Tagged rose are adversarial personas; teal are cooperative edge cases that probe empathy and boundary handling rather than injection.
Execution pipeline
endpoint
red-team
agent
plan
execution
summary
Combined with Quality Analysis? If Quality Analysis is also enabled, the Cave of Shadows agent receives those reports and uses them to generate targeted, finding-specific scenarios in addition to standard probes.