Index/ AESOP/ Black Box Investigation
Discovery mode — AESOP

Black Box Investigation

Discover a chatbot’s behavior through progressive, multi-turn interviews — with no access to its instructions.

01

What investigation does

Black Box Investigation is built for situations where you have a live chatbot but no access to its system instructions. AESOP interviews the chatbot through progressive, multi-turn conversations to discover its behavior, capabilities, boundaries, and potential issues — all without ever seeing the underlying prompt. The investigator agent acts as a skilled interviewer, probing from multiple angles to build a comprehensive behavioral profile.

02

When to use it

3PThird-party
chatbot
Evaluating a vendor or partner chatbot where you have endpoint access but not the system prompt or configuration.
INInherited
system
A chatbot built by a previous team with no documentation. Investigation reveals what it actually does.
CACompetitive
analysis
Understanding a competitor’s chatbot capabilities, safety posture, and behavioral patterns through structured interviews.
PAPre-audit
discovery
Before a formal evaluation, investigate the chatbot to understand what you’re working with and scope the audit appropriately.
03

How it works in the wizard

In the evaluation wizard, skip the system instructions in Step 1 and provide only a chatbot endpoint URL in Step 2. When no instructions are provided, Black Box Investigation becomes available in Step 3. Select it to begin the automated interview process. AESOP handles the rest — running multi-turn conversations, analyzing transcripts, and building a behavioral profile.

04

What it discovers

PurposeSystem
purpose
What the chatbot is designed to do, its domain of expertise, and its stated or implied mission.
BehaviorBehavioral
patterns
Conversational style, response strategies, escalation behavior, and how it handles ambiguity.
BoundariesSafety
posture
Where the chatbot draws lines, how it refuses requests, and its approach to sensitive topics.
RisksPotential
vulnerabilities
Gaps in boundary enforcement, inconsistent behavior, and areas where the chatbot may be exploitable.
05

What happens next

After investigation completes, AESOP can optionally generate evaluation rubrics and test scenarios from the discovered behavior. The investigation report feeds forward — the behavioral profile can go into Cave of Shadows for red-team testing, or seed a Custom Rubric tailored to the chatbot’s actual capabilities rather than assumed ones.

06

The investigation pipeline

Input — wizardEndpoint
only
Provide the chatbot endpoint URL. No system instructions needed.
WizardEnable black box
investigation
Available when no instructions are provided — select to begin automated discovery.
Fork — interview phaseThree
probes
Progressive probing Gradually escalating questions to map capabilities and boundaries.
Behavioral mapping Multi-angle conversations to understand response patterns and style.
Boundary testing Exploring edge cases and refusal behaviors across sensitive topics.
DiscoveryAnalyze interview
transcripts
Reconstruct an understanding of system behavior from conversation evidence.
InferenceRubric
generation
Generate evaluation criteria based on discovered capabilities and behaviors.
OptionalScenario
execution
Optional. Generate and run test scenarios against the chatbot based on discovered behavior.
OutputInvestigation
results
Behavioral profile, discovered capabilities, identified risks, and recommended follow-up evaluations.