PuRDy — AI Agent Quality at Scale
We built a system where AI agents are not shipped based on “looks good” — they are scored against business-defined rubrics, red-teamed across difficulty tiers, and improved through structured repair cycles until they meet a measurable quality bar. PuRDy is the first agent through this process. This is the model for every AI deployment going forward.
The Problem We Solved
Why PRD intake was broken, and what it costs
Employees face a blank 15-page template and either submit incomplete documents or don’t submit at all. IS and AI Strategy then spend significant time on clarification calls. This bottleneck scales linearly with request volume — every new project request costs the same manual effort to get right.
What PuRDy Is
A self-service PRD interview agent + automated formatting pipeline
What changed: Employees no longer face a blank template. They have a conversation. The agent does the requirements thinking. IS receives complete, structured, submission-ready documents.
The Build-Eval-Repair Loop
How we went from v1 to production-ready through structured improvement cycles
We do not build an agent, look at the output, and say “that seems fine.” We run it through a rigorous, repeatable improvement cycle where every change is driven by scored evidence, not intuition.
The Agentic Toolchain
Four purpose-built tools working together
Studio
Engine
Scenarios
Pipeline
How Evaluations Work
Not “does this look good?” but “does this meet a measurable standard?”
The Rubric: 5 Dimensions, 16 Criteria
Every criterion has anchored scoring — specific descriptions of what a 1, a 5, and a 10 look like. Scores are not vibes. They are measurements.
Completion
- PRD production — does the artifact exist with all 4 minimum elements?
- Depth calibration — lean PRD for simple projects, full for complex?
- Problem-first orientation — describes a problem, not a solution?
- Open questions captured — real gaps documented, not perfunctory?
Quality
- Stage sequence adherence — 8 stages in order, full depth on Stage 2?
- Classification accuracy — correct project type, user confirmation?
- Depth matrix application — right number of questions per section?
- Information state tracking — knows what’s missing, probes to fill gaps?
Skill
- Question pacing — one question per response, natural rhythm?
- Active listening — references prior answers before next probe?
- Coaching move deployment — uses prescribed techniques?
- Summary quality — accurate confirmation before generating PRD?
Detection
- Gap and inconsistency surfacing — catches contradictions, missing stakeholders?
- Scope and routing assessment — identifies non-IS projects, routes appropriately?
Quality
- Content specificity — names, numbers, systems from the conversation?
- Section structure — matches expected PRD format?
- Downstream readiness — IS team can act on this document?
- Synthesis quality — overview captures the whole story?
Scenario Library: Tiered Difficulty
We don’t just test the happy path. Scenarios are designed to stress-test the agent across different project types, user behaviors, and difficulty levels.
Tier
Tier
Tier
Score Progression Across Versions
Measurable improvement from v1 through v8+
Evaluation Scores by Dimension Across Repair Cycles — each dimension weighted, scored 1-10, tracked v1 through v8.
| Version | Goal Completion (25%) | Process Quality (20%) | Conversational Skill (10%) | Red Flag Detection (20%) | Output Quality (25%) |
|---|---|---|---|---|---|
| v1 | 4.5 | 4.25 | 5.0 | 4.0 | 5.0 |
| v2 | 7.5 | 7.0 | 7.0 | 6.0 | 7.5 |
| v3 | 8.2 | 7.8 | 7.5 | 6.5 | 8.0 |
| v4 | 8.8 | 8.5 | 8.7 | 7.3 | 8.7 |
| v5 | 8.5 | 8.3 | 8.3 | 7.3 | 8.7 |
| v6 | 8.5 | 8.0 | 8.0 | 7.3 | 8.3 |
| v7 | 8.8 | 8.0 | 8.0 | 8.0 | 8.3 |
| v8 | 8.8 | 8.8 | 8.7 | 8.0 | 8.7 |
Weighted average: ~4.5/10. Hallucinated routing details, no classification confirmation, missing metadata section, no executive summary. Functional but not IS-ready.
Fixed fabricated contact names and Slack channels. Added confirmation gate before proceeding. Added metadata and exec summary sections. One-question-per-response rule added.
Added “Why Now” urgency question, constraints capture, Enablement section. Agent now covers ~95% of Stratton’s template requirements. Output went from ~60% to ~95% template coverage.
Solution-first redirect enforcement. Multi-question batching eliminated. End-user perspective probing added. Red flag detection improved from 7.3 to 8.0.
Vacation test coaching move mandated. Active listening callbacks required. Metadata pre-verification step. ROI arithmetic validation. Named routing destinations instead of generic advice. Scores stabilized at 8.0+ across all dimensions.
What the Scores Mean for the Business
From “we hope this works” to “we can prove it works”
Post-Deploy: Continuous Monitoring
The eval loop doesn’t stop at launch
Once PuRDy is live with real users, the same evaluation infrastructure becomes a continuous quality monitoring system.
- Prompt drift detection — scores trending down across conversations signal the agent’s behavior is drifting from its instructions. Catch it before users notice.
- Meta-analysis across conversations — are certain project types consistently scoring lower? Are users in certain departments struggling? Pattern recognition at scale.
- Continuous improvement cycles — the same Build-Eval-Repair loop runs on production data, not just test scenarios. The agent gets better from real usage, not just our imagination of usage.
- Quality reporting for leadership — IS leadership gets visibility into agent quality metrics over time. Not “it’s working” — but “here are the scores, here are the trends, here’s what we fixed this month.”
Why This Matters
The difference between AI that’s deployed and AI that’s governed
We are not just shipping AI tools. We are building a quality system.
Anyone can put an agent in Glean and call it done. What we’ve built is different:
- Rubrics defined by people, not AI — the business decides what “good” means. Criteria are anchored to what IS actually needs to approve and scope a project.
- Scenarios that stress-test, not just validate — we don’t test with ideal inputs. We test with solution-first language, missing sponsors, vague quantification, adversarial behavior.
- Evidence-based improvement, not intuition — every change to the agent traces back to a scored finding. We can show exactly why v7 is better than v6.
- Continuous monitoring, not “launch and forget” — the same eval infrastructure that built the agent continues to monitor it in production.
- Full audit trail — version history, evaluation reports, repair changelogs, score progression. Complete traceability from first draft to production.
The Model Going Forward
PuRDy is agent #1. This is how we build every AI deployment.
The toolchain and methodology we built for PuRDy is not PuRDy-specific. It’s a reusable framework for any AI agent deployment across the organization and beyond.
Agents already in this pipeline:
- PuRDy (IS) — PRD interview agent. 9 repair cycles complete. Approaching production.
- Sales Ops Agent — Sales workflow support. In requirements phase.
- Legal Separation Agent (Legal/HR) — Agreement generation + severance calculation. In development.
- Discovery PuRDy (Cross-functional) — Product requirements interviews. Pipeline integration planned.
Built with AESOP Studio, Glean, AESOP Evaluation Engine.