Index/ AESOP/ Custom Test Suite
Rubric Generator v1.4 · Scenario Generator v1.0

Custom Test Suite

A wizard capability that generates bespoke rubrics and test scenarios from your system instructions.

01

How to access it

1Provide
instructions
Open the evaluation wizard and provide your system instructions and any context documents in Step 1.
2Optional
endpoint
Optionally provide a chatbot endpoint URL in Step 2. This determines whether you get a static framework or live execution with transcripts.
3Enable
capability
In Step 3 (Capabilities), enable the “Custom Test Suite” toggle. This activates the custom rubric and scenario-generation pipeline alongside any other selected capabilities.
02

Static vs. live execution

No endpointStatic
framework
A complete evaluation framework for manual review. Ideal when you want to inspect the rubric and scenarios before testing, or when the chatbot isn’t deployed yet.
+ Custom rubric with weighted scoring categories
+ Generated test scenarios with personas
+ Conversation starters and testing objectives
− No automated execution or transcripts
Endpoint providedLive
execution
Full automated evaluation. Scenarios execute against your live chatbot endpoint, producing conversation transcripts and scored results.
+ Everything in the static framework
+ Automated scenario execution against the chatbot
+ Full conversation transcripts
+ Scored results per rubric category
03

The two-stage pipeline

Stage 1Rubric
Generator
v1.4 — rubric generation
Analyzes your instructions and builds a tailored rubric with weighted scoring categories specific to your use case.
3 core categories always included — Safety, Core Function, Structure
Up to 4 dynamic categories based on the system’s purpose
Confidence scoring and gap analysis
Comprehensive safety framework with universal tests
Outputs detailed analysis + structured rubric
Stage 2Scenario
Generator
v1.0 — scenario generation
Transforms the rubric into comprehensive test scenarios with realistic personas and conversation starters.
Scenario count weighted by rubric-category priorities
Realistic personas with clear testing objectives
Universal safety tests always included
Both cooperative and adversarial persona types
Output compatible with Cave of Shadows execution format
04

Rubric framework

Core — alwaysSafety &
Boundaries
Always included.
Core — alwaysCore
Functionality
Always included.
Core — alwaysStructural
Integrity
Always included.
DynamicDomain-specific
Based on the AI’s purpose.
DynamicInteraction
quality
If user-facing.
Dynamic+ Up to 2 more
Context-dependent.
05

Wizard entry & rubric generation

Wizard step 1Instructions +
context
System instructions and context documents provided in the evaluation wizard.
Wizard step 2Chatbot
endpoint
Provide a URL to enable live execution. Skip for a static framework only.
Wizard step 3Enable Custom
Test Suite
Toggle on the capability to activate the rubric and scenario pipeline.
Rubric genAnalyze
instructions
Parse and understand the AI system’s purpose and domain.
Rubric genIdentify purpose
& risks
Map capabilities, boundaries, and potential vulnerabilities.
OutputCustom rubric
JSON — 3 core + up to 4 dynamic weighted scoring categories.
06

Scenario generation & execution

Scenario genBuild test
scenarios
Transform rubric categories into realistic test personas.
OutputEvaluation
plan
JSON scenarios — Personas, conversation goals, and start messages.
DecisionEndpoint
provided?
From wizard Step 2 — determines static or live execution.
YES Execute each scenario against the live endpoint; collect transcripts and scored results.
NO Deliver the rubric + scenarios as a complete evaluation framework for manual review.