← Index
Rubric Generator v1.4 Scenario Generator v1.0

Custom Test Suite

A wizard capability that generates bespoke rubrics and test scenarios from your system instructions

How to Access
1
Open the evaluation wizard and provide your system instructions along with any context documents in Step 1.
2
Optionally provide a chatbot endpoint URL in Step 2. This determines whether you get a static framework or a live execution with transcripts.
3
In Step 3 (Capabilities), enable the "Custom Test Suite" toggle. This activates the custom rubric and scenario generation pipeline alongside any other selected capabilities.
Static vs. Live Execution
No endpoint provided
Static Framework
Get a complete evaluation framework for manual review. Ideal when you want to inspect the rubric and scenarios before testing, or when the chatbot is not yet deployed.
+ Custom rubric with weighted scoring categories
+ Generated test scenarios with personas
+ Conversation starters and testing objectives
- No automated execution or transcripts
Endpoint provided
Live Execution
Full automated evaluation. Scenarios are executed against your live chatbot endpoint, producing conversation transcripts and scored results.
+ Everything in the static framework
+ Automated scenario execution against chatbot
+ Full conversation transcripts
+ Scored results per rubric category
The Two-Stage Pipeline
1
Rubric Generator
Stage 1 — Rubric Generation · v1.4
Analyzes your AI system instructions and generates a tailored evaluation rubric with weighted scoring categories specific to your use case.
3 core categories always included (Safety, Core Function, Structure)
Up to 4 dynamic categories based on the system's purpose
Confidence scoring and gap analysis
Comprehensive safety framework with universal tests
Outputs detailed analysis + structured rubric
2
Scenario Generator
Stage 2 — Scenario Generation · v1.0
Transforms the custom rubric into comprehensive test scenarios with realistic personas and conversation starters.
Scenario count weighted by rubric category priorities
Realistic personas with clear testing objectives
Universal safety tests always included
Both cooperative and adversarial persona types
Output compatible with Cave of Shadows execution format
Rubric Framework
Core
Safety & Boundaries
Always included
Core
Core Functionality
Always included
Core
Structural Integrity
Always included
Dynamic
Domain-Specific
Based on AI purpose
Dynamic
Interaction Quality
If user-facing
Dynamic
+ Up to 2 More
Context-dependent
Wizard Entry & Stage 1 — Rubric Generation
Wizard Step 1
Instructions + Context
System instructions and context documents provided in the evaluation wizard
Wizard Step 2 (optional)
Chatbot Endpoint
Provide a URL to enable live execution. Skip for static framework only.
Wizard Step 3
Enable "Custom Test Suite"
Toggle on the Custom Test Suite capability to activate the custom rubric and scenario generation pipeline
Rubric Generator v1.4 — Rubric Generation
Step 1
Analyze Instructions
Parse and understand the AI system's purpose and domain
Step 2
Identify Purpose & Risks
Map capabilities, boundaries, and potential vulnerabilities
Output
Custom Rubric
JSON
3 core + up to 4 dynamic weighted scoring categories
Stage 2 — Scenario Generation & Execution
Scenario Generator v1.0 — Scenario Generation
Step 1
Build Test Scenarios
Transform rubric categories into realistic test personas
Output
Evaluation Plan
JSON Scenarios
Personas, conversation goals, and start messages
Decision
Endpoint Provided?
From wizard Step 2 — determines static or live execution
Yes — Live Execution
Execute Against Chatbot
Run each scenario against the live endpoint, collect transcripts and scored results
No — Static Framework
Deliver Evaluation Framework
Rubric + scenarios delivered as a complete evaluation framework for manual review