What Investigation Does
Black Box Investigation is designed for situations where you have a live chatbot but no access to its system instructions. AESOP interviews the chatbot through progressive, multi-turn conversations to discover its behavior, capabilities, boundaries, and potential issues -- all without ever seeing the underlying prompt. The investigator agent acts as a skilled interviewer, probing the chatbot from multiple angles to build a comprehensive behavioral profile.
When To Use It
3P
Third-Party Chatbot
Evaluating a vendor or partner chatbot where you have endpoint access but not the system prompt or configuration.
IN
Inherited System
A chatbot built by a previous team with no documentation. Investigation reveals what it actually does.
CA
Competitive Analysis
Understanding a competitor's chatbot capabilities, safety posture, and behavioral patterns through structured interviews.
PA
Pre-Audit Discovery
Before a formal evaluation, investigate the chatbot to understand what you are working with and scope the audit appropriately.
How It Works in the Wizard
In the evaluation wizard, skip the system instructions in Step 1 and provide only a chatbot endpoint URL in Step 2. When no instructions are provided, the "Black Box Investigation" option becomes available in Step 3. Select it to begin the automated interview process. AESOP will handle the rest -- conducting multi-turn conversations, analyzing transcripts, and building a behavioral profile of the chatbot.
What It Discovers
Purpose
System Purpose
What the chatbot is designed to do, its domain of expertise, and its stated or implied mission.
Behavior
Behavioral Patterns
Conversational style, response strategies, escalation behavior, and how it handles ambiguity.
Boundaries
Safety Posture
Where the chatbot draws lines, how it refuses requests, and its approach to sensitive topics.
Risks
Potential Vulnerabilities
Gaps in boundary enforcement, inconsistent behavior, and areas where the chatbot may be exploitable.
What Happens Next
After investigation completes, AESOP can optionally generate evaluation rubrics and test scenarios based on the discovered behavior. The investigation report serves as a foundation for further evaluation -- you can feed the behavioral profile into Cave of Shadows for red-team testing, or use it to create a Custom Rubric tailored to the chatbot's actual capabilities rather than assumed ones.