2.5 years. Eight MVPs. Over a million tests exercised across nine hundred builds. A method for building AI that doesn't lie to you about whether it's working.
How the work gets made. A stack of habits. Compound interest.
What the work taught. Some obvious in retrospect. None obvious at the time.
The fingerprints. The things that show up across every project, every repo, every artifact, whether anyone asked for them or not.
What all of this is in service of. The thesis. The operating model. The thing being assembled.
The work that distinguishes transformation from automation. Not mapping what exists and speeding it up. Questioning whether what exists is the right answer.
Stuart Winter-Tear's organizing insight: "Organizations are compensating mechanisms for the characteristics and limitations of human intelligence." Departments exist because people specialise. Hierarchies exist because managers cannot coordinate unlimited work. Meetings exist because context has to move between human beings. Approvals exist because judgement is scarce.
If AI changes the economics of those constraints — context becomes cheaper to retrieve, knowledge becomes persistent rather than personal, routine judgement becomes delegable, coordination costs fall — then the organizational structures built around those constraints deserve to be questioned. Not optimized. Questioned.
AESOP's discovery pipeline splits into two distinct phases, each with a different relationship to truth:
The output of authored discovery is a decision surface — a structured, actionable JSON artifact that captures:
Most AI transformation work automates the organization that exists. Authored change asks what the organization should become when the constraints that created it no longer apply. This is the difference between deploying chatbots and redesigning how legal work gets done. Between faster contract review and questioning whether contract review needs to be a separate function at all.
The decision surface connects directly to executive decision-making: it gives leadership a structured artifact they can act on, with named owners, measurable receipts, and explicit boundaries. It turns "we should think about how AI changes our operating model" into "here are the 7 decisions you need to make, here's what depends on each one, here's who owns them, and here's what happens after." Not advice. A surface.
Why it was built. How. How it was graded. How it was watched.
| Platform | What It Is & Why It Matters | Scale |
|---|---|---|
| MVPs | ||
| AESOP OS | Enterprises deploy AI agents with no independent evaluation. Vendors grade their own homework. AESOP is the external evaluation layer: a full pipeline from discovery through authored organizational change to deployment, with 6 evaluation modes and a scoring engine that blocks unsafe output before it ships. Certification tiers stakeholders can trust. Drift detection that catches degradation before it becomes liability. | 990 tests 6 eval modes |
| Legal AI OS | Harvey agents deployed across the enterprise need independent evaluation. Harvey can't grade its own homework. Legal AI OS provides the external governance layer: 9 built legal functions, independent 4-dimension agent evaluation, drift monitoring, and an agentic help system. A reusable 5-layer architecture for any enterprise legal AI program. | 62 tests 9 functions |
| SuperAssistant | AI assistant platform for a professional services firm. Multi-tenant architecture where client data can never leak between tenants — enforced at the foundation, not in application code. Chat interface with knowledge management. Verified against production with 428 automated tests. | 428 tests TBG-owned |
| TSE-Stripped | Voice interviews produce rich qualitative data. Processing it manually doesn't scale. A fully automated serverless pipeline: voice agents capture the interview, structured analysis is delivered automatically. The architecture for any voice-to-insight pipeline. | Production deployed |
| Walter | Persistent AI analysis that survives individual sessions. Knowledge that lives beyond the conversation. Session analysis, insight extraction, social-safe excerpts. A firm's collective intelligence, indexed and retrievable. | TBG-owned infra |
| POCs & Portfolio | ||
| Matter Intake Evaluator | Legal intake evaluation with configurable rubrics, not hardcoded rules. Staffing teams and evaluation criteria that adapt to the firm, not the other way around. Deployed and in use. | 24 tests |
| 132 Visual Explainers | A portfolio of thinking tools. Architecture overviews, process maps, decision flows, maturity models — each a single self-contained file that opens anywhere and survives the platform it was built on. Custom typography, intentional color, dual light and dark themes. The medium enforces portability. Nothing rots. | 132 pages |
The patterns that port. Built, tested, deployed across multiple projects. Ready to drop into anything.