← Index
Enterprise Legal AI

Legal AI Operating System

A roadmap for the three problems every large law firm is facing right now — and a platform for building, deploying, and governing AI where every decision is auditable, explainable, and traceable by design.

Three Things Keeping Your Firm Up at Night

If your firm is like every other AmLaw 100 firm right now, you are staring at these three numbers and trying to figure out what to do about them.

69%
Shadow AI Is Already Inside Your Firm
69% of legal professionals use generative AI for work — most without firm knowledge or approval. Every associate with a ChatGPT tab open is a confidentiality event waiting to happen. You cannot govern what you pretend is not happening. The question is not whether AI is being used. It is whether you have any visibility into how, or any control over what happens to the data.
43%
Most Firms Have No Governance — Or Fake Governance
43% of firms have no AI policy and no plans to create one. The remaining 57% mostly have a document nobody read after week one. ABA Formal Opinion 512 makes clear: the duty of competence now includes understanding the benefits and risks of relevant technology. A policy document is not governance. Governance is structural — it lives in the database, not in a PDF.
85%
The Billable Hour Model Is Under Structural Pressure
85% of clients say firms should disclose when AI is used on their matters. 64% of in-house teams expect to depend less on outside counsel. The billable hour is not dead — but the premium for routine work is evaporating. The firms that thrive will be the ones that use AI to deliver faster, better work — and can prove it with an audit trail.

These three problems share one root cause

Shadow AI exists because there is no safe, approved alternative. Governance is fake because nobody knows how to make it structural. The billable hour is under pressure because AI makes routine work cost zero — and clients know it. The solution to all three is the same: a governed platform that makes AI safe to use, provably compliant, and measurable in outcomes. That is what the Legal AI OS is.

How You Address All Three

Five layers. Each independent. Governance at the center — not as something you bolt on after the fact, but as the structural foundation every function is built atop. This is the map from "AI is happening to us" to "we control how AI happens."

Layer 1 — Non-Negotiable Foundation
Governance & Trust

Everything above this layer depends on it. Nothing below it can compromise it. Governance is not a compliance checklist. It is the architecture. Three pillars are non-negotiable — without them, the system is a liability, not an asset:

Pillar 1
Auditability
Every evaluation captures the full prompt, the full response, the model version, the rubric version, and the scoring formula. When a partner questions a classification six months later — or when a client's outside counsel guidelines audit asks for proof — the full decision context is retrievable. Not "the AI said so." Here is exactly why, with the raw data.
Full prompt capture. Response + reasoning chain. Deterministic score replay. Structured JSONL logging. Immutable.
Pillar 2
Explainability
Every AI output includes its chain of reasoning — visible to the reviewer, not hidden in a vector. Classification decisions cite the specific clause, signal, or pattern that drove the result. Risk scores decompose into their weighted dimensions. No black-box recommendations. The attorney understands why before they decide whether.
Chain of reasoning. Source attribution. Dimension-level score decomposition. Visible, not hidden.
Pillar 3
Traceability
Who ran the evaluation. When. Which model version. Which rubric version. What the score was. Whether a human overrode it, and why. Full chain of custody from upload through review. Compliance-ready artifacts on demand — SOC 2, ISO 42001, EU AI Act, ABA 512, client audit requests. The answer to "prove it" is always one query away.
Chain of custody. Who, when, what, why. Every override logged. Every artifact exportable.

These three pillars are enforced by four supporting layers: Confidentiality Architecture (Row-Level Security at the database — Client A's data cannot touch Client B's models), Human-in-the-Loop Gating (AI recommends, human decides — confidence below threshold triggers mandatory escalation), Model Governance (pre-deployment multi-agent evaluation with hard veto rules), and Compliance Readiness (ABA 512, SOC 2, ISO 42001, EU AI Act, Outside Counsel Guidelines — all covered by the architecture, not a policy document).

Confidentiality Architecture Audit Trail Explainability Traceability HITL Gating Model Governance Data Privacy
Layer 0 — Foundation
Knowledge Foundation

Unified knowledge base, precedent libraries, clause libraries, search infrastructure. Everything the AI functions need to reason over — inside the trust boundary. Embedding and analysis run within the firm's enterprise infrastructure. Enterprise LLM agreements cover all API calls: no training on client data, no retention, no leakage between clients. Without this layer, every function starts from zero. With it, two centuries of institutional memory becomes searchable in seconds.

Layer 2 — Execution
Legal AI Functions

Standalone applications. Each owns its UI, workflow, and data. Each exposes a governance contract — health, metrics, and evaluation targets — so the governance layer can verify compliance without coupling to implementation. Functions are independent: replace one without touching the others. Every function follows the same pattern: Router classifies, Evaluator reasons, programmatic layer applies weights and thresholds, audit trail captures everything. The LLM provides reasoning. The system provides judgment.

Layer 3 — Operations
Program Operations

The bridge between AI capability and organizational adoption. Portfolio dashboard tracks every investment — adoption rates, cost impact, quality metrics, ROI. AI literacy and enablement turns capability into daily practice. Stakeholder discovery ensures every project starts with how the work actually gets done — go in to learn, not tell. Client engagement and change management turn AI from a cost center into a retention asset: "Here is what AI saved this quarter, with the audit trail to prove it."

Layer 4 — Structure
Organizational Model

Maps AI functions to practice groups, divisions, and strategic priorities. Not every function applies to every group. The operations layer tracks adoption by division. The governance layer enforces boundaries between them — employment legal walled from commercial legal at the database layer. Corporate gets contract review and due diligence. Litigation gets matter intake and regulatory monitoring. IP gets KM intelligence with the highest confidentiality walls. Every group gets what it needs, and nothing it should not see.

What's Built

Nine functions built and deployed. Each follows the same governed pipeline — Input → Router → Evaluator → Scoring → Audit Trail. Plus Harvey Agent Monitoring and an agentic help system.

Built / Deployed
Built
Matter Intake & Triage
Two-stage pipeline: Router classifies practice area and urgency, Evaluator scores across 5 weighted dimensions. Programmatic scoring never trusts the LLM directly. Full audit trail — every prompt, response, and rubric version captured.
Under 10 seconds. Every decision auditable. The highest-leverage entry point for firm-wide AI adoption.
Built
Contract Review & Analysis
Five specialized agents — vendor, customer, employment, DPA, general. Classification to extraction to risk scoring to flagging. RAG-powered with 30+ legal standards. Human-in-the-loop checkpoints. Full audit trail.
360 contracts per hour. Production, not demo.
Built
Employment Legal Agents
Separation agreement generator covering US, EMEA, and Australia jurisdictions. Legal metrics analysis agent for outside-counsel spend by firm and practice area. State annual report filing automation.
Built inside a real legal department with real confidentiality constraints. RLS walls between employment and commercial legal.
Configured
Cowork Legal Plugin
Nine skills for in-house legal teams: contract review, NDA triage, compliance checks, risk assessment, meeting briefings, legal responses, vendor checks, signature requests, daily briefings. Playbook-driven.
Playbook-driven. legal.local.md is the control document. Attorney reviews every redline before sending.
Built / Deployed
Built
Due Diligence Accelerator
Bulk document review with batch ingest, parallel classification, clause-level comparison against target standards. Identical clauses grouped; only deviations shown to reviewer. 87% time reduction.
3-5 deals/year. ~60 attorney hours saved/year. Deployed on Fly.io + Vercel.
Built
Regulatory Change Monitor
Monitors 15-20 regulatory sources. Extracts changes, maps to active matters by jurisdiction and practice area. Classifies severity. Notifies within 24 hours via Slack/email/webhook.
2 critical changes caught before client impact. Celery-backed scheduled polling.
Built
KM & Precedent Intelligence
Natural language query → LLM refinement → semantic vector search → enriched results. "Have we handled a separation in Germany before?" — answered in seconds, not partner memory.
~120 hours saved/year across 12 lawyers. 100% of client precedent searchable.
Built
Client Value Reporting
Quarterly reports for all clients proving AI value. Every number backed by the audit trail. Aggregate → Calculate → Assemble (governance docs, ABA 512, audit samples) → Generate executive summary.
Minutes per report — was days of manual assembly. 100% audit-trail-backed.
Built
AI Maturity Assessment
Organizational AI readiness evaluation. Scored across governance, technology, data, culture, and process dimensions. Gap analysis with prioritized recommendations.
Baseline measurement for the firm's AI journey. Deployed with Celery backend.
Harvey Agent Monitoring & Help System

Two platform-level services that operate across all layers. Independent evaluation of Harvey AI outputs, and an agentic help system for the entire platform.

Built
Harvey Agent Monitoring
4-dimension independent evaluation of Harvey AI outputs: Accuracy, Safety, Bias, Compliance. Weighted scoring (Tier 1 at 1.5x). Hard veto at 75 on any Tier 1 dimension. Certification: Platinum → Bronze. Drift detection across 6 types (tone, scope, refusal, instruction erosion, hallucination, safety). Agent registry, evaluation history, drift alerts.
Harvey can't grade its own homework. Ported from AESOP evaluation engine. 21 tests.
Built
Agentic Help System
RAG-powered chat: Voyage AI embeddings → vector similarity search (pgvector) → LLM-generated answer with source references. 11 indexed docs, 89 chunks. Slide-out panel with suggested questions. Local FAQ fallback when API unreachable. Thumbs feedback on answers.
Same architecture as AESOP's help system. Deployed on Fly.io + Vercel.

Deterministic over LLM

The platform never trusts an LLM score directly. Every classification, risk score, and recommendation goes through a programmatic scoring layer. The LLM provides reasoning. The system applies weights, thresholds, and rules. This is the architectural expression of the three pillars: auditability means the scoring is replayable. Explainability means every dimension is decomposed and visible. Traceability means the full chain — prompt to model to rubric to score to human override — is captured and immutable. Nothing is a black box. The audit trail is the product.