legal-os · governance monitoring

Fault-Line Radar: the scoring engine

An engine that forecasts where the next legal-AI rulings land, then grades itself on whether it called them.

Rulings land on fault lines, the places where an AI capability runs ahead of a rule lawyers already live by.

You can’t predict the wording of a rule nobody has written. You can predict where it lands.

Forecast the fault line, not the ruling.
The core idea — two curves

Capability moves faster than governance

Two curves. Capability is what AI can do in legal work. Governance is what the rules allow. Capability climbs fast, governance climbs slow, and every fault line lives in the gap. That’s where the rulings land.

Capability Governance fault-line zone
Capability — what AI can do Governance — what the law allows The gap — where rulings land
Three orders of prediction — a cascade
First order

Capability

AI can now do the thing that creates the fault line. Agentic autonomy, error rates, a new model release. Weighted low and clearly labeled, so a capability that arrived never reads as a ruling that landed.

L1 · capability meter (live)
Second order

The ruling reacts

A court, bar, or statute moves on that behavior. This is the pressure meter the radar shipped.

L2 · ruling pressure
Third order

The market adopts

The control becomes table stakes, court or no court. Insurers, RFPs, and firms standardize it.

L3 · control adoption

The engine grades each level, and the chain, reporting P(ruling called | capability called) and P(adoption called | ruling called), not just a headline hit rate, because an adoption call only means something if the ruling it depends on held. But the cascade is the default frame, not a law: some fault lines are pure market or regulatory reactions with no capability behind them, and some run backwards, where a firm deploying a tool in a filing provokes the rule, so the model is moving to type each line by what actually drives it rather than forcing one order onto all of them.

The third order — the operating-model queue

Three lanes per fault line, each graded on the market signal it can actually carry:

The queue is the intersection. Both high means the control is urgent and about to be mandatory, and the lead time says how long you have to stand it up. That’s the build-before-it-lands list.

Control to buildCapabilityRulingAdoptionQueueLead
Verification-by-default with an audit trail6.79.88.17.9now
Documented verification + trace logs6.79.88.17.9~2 quarters
Firm-wide training + governance program4.59.68.27.9~1 year
Governance as an insurable artifact2.58.88.47.4now
One operating model, to the strictest standard2.59.37.77.2now
Data-flow mapping + vendor attestation4.59.47.06.6now

Snapshot figures; the live engine cites the market signals behind every adoption reading. Queue = ruling × adoption, so a low reading on either meter pulls a control down the list. Capability is shown as context, not multiplied in — where it runs far ahead of the ruling (the error-rate benchmark reads 8.6 capability against a 7.2 ruling), it flags a fault line the law hasn’t caught up to yet.

How it scores — authority over volume

A legal-AI feed is mostly vendors selling tools. Let the loudest source win and the forecast just tracks marketing. So evidence is weighted by earned authority. A vendor blog carries about 1/100 of an ABA opinion.

1.0
T1 · Binding
Court rulings, sanctions orders, statutes, regulations. This is the landscape.
0.8
T2 · Guidance
ABA & state bar ethics opinions, judicial standing orders.
0.6
T3 · Empirical
Peer research, trackers, insurer/market actions carrying data.
0.35
T4 · Commentary
Firm client alerts, bar magazines, established legal press.
0.008
T5 · Vendor
Vendor blogs, SEO, hot takes. Narrative signal only.
Conflict discount — a source selling what it comments on Empirical boost — items carrying hard data Corroboration — agreement across independent tiers Recency decay — momentum fades; binding law never does Full provenance — every score cites its evidence and weight
The 11 fault lines — ruling pressure (L2)
Fault lineDutyPressure (0–10)Horizon
Disclosure & certification standardize3.3 / R.11
9.8
near
“A human reviewed it” stops being enough5.1 / 5.3
9.8
near–mid
Competence becomes a governance duty1.1
9.6
mid
Confidentiality hardens into data governance1.6 / 1.9
9.4
near
Regulatory convergence forces one model1.1 / 1.6
9.3
near
Insurance becomes the real regulator1.1 / 5.1
8.8
near
Benchmark / tool-certification standard1.1
7.2
mid
Agentic autonomy reopens supervision & UPL5.3 / 5.5
7.1
mid
Judicial / litigation analytics regulated3.5 / 8.4
6.4
mid
The billable hour cracks under AI1.5
5.8
mid–far
Liability shifts partway to vendors
5.8
far

Pressure is an analyst seed moved by weighted evidence, a rubric for where load is building. Snapshot figures; each reading on the live page cites the evidence behind it.

Calibration — the engine grades itself
67%

Called before they landed, graded by order

Anyone can have opinions about where the law is heading. Almost nobody grades their own. For each event that landed, the radar replays itself 90 days earlier on the evidence available then, and checks whether the right meter was already flagged. It earns trust by being scored, at every order.

L1 capability 1/3  ·  L2 rulings 3/3  ·  L3 adoption 2/3  ·  P(L2|L1) n/a  ·  P(L3|L2) 1/1

2024-01-11○ L1 missedStanford/Yale — legal LLM hallucination proven (58–88%)benchmark 6.5 → 7.4
2024-05-23○ L1 missedRegLab — “a human reviewed it” shown insufficientverification 5.5 → 6.7
2026-06-01● L1 calledLegal AI tools quantified at 17–34% errorbenchmark 8.1 → 8.6
2025-12-15● L2 calledCouvrette — ~$110K sanction, 15 fake casesdisclosure 9.2 → 9.2
2026-04-15● L2 calledNebraska suspends attorney — 57/63 bad citesverification 9.3 → 9.5
2026-08-01● L2 calledEU AI Act obligations take effectconvergence 9.1 → 9.3
2026-06-15● L3 calledCNA makes AI governance a coverage conditioninsurance 7.5 → 8.7
2026-08-31● L3 calledVerification duty extended to expert work productverification 7.3 → 8.1
2026-08-29○ L3 missedAI AGENT Act — agent-authorization controlagentic 5.0 → 6.2

The misses are honest. Both L1 misses are first-of-kind: the earliest demonstration of a capability has no precursor on record, so the lane cannot call it early — it only calls the second-generation benchmark once the first ones are logged. P(L2|L1) is n/a because the one capability the lane called — error-rate benchmarks — has no ruling event yet: capability outran the law, which the engine reports rather than scoring as a fail. The agentic L3 miss is the market not having moved 90 days out. Seed set is small and the L1 lane is young; this is the track record so far, and it grows every run. Every run writes an immutable audit artifact — knobs, evidence, weights, math, predictions, backtest — fully replayable.

A worked example — how the Nebraska call was made

What the radar saw 90 days out

By mid-January 2026, three months before the Nebraska Supreme Court suspended an attorney over a hallucinated brief, the verification fault line was already loaded. Four independent sources, all pointing the same way:

T1 · 2023

Mata v. Avianca. The first Rule 11 sanction for fabricated AI citations. Binding, and it never decays out of the score.

T2 · 2024

ABA Opinion 512. You own the tool’s output. A wrong cite is your problem in front of the court, not the model’s.

T2 · 2025

Texas Ethics Opinion 705. A state bar echoing the same duty. A second, independent tier moving in the same direction.

T1 · Dec 2025

Couvrette v. Wisnovsky. A fresh ~$110K sanction for 15 fake cases, landing 30 days before the window opened.

Four tiers, one direction, corroborating across sources that don’t depend on each other. The radar read verification at 9.2 and flagged it. In April the suspension landed on exactly that line. Called, on evidence anyone could have read.

Architecture — the data flow
sources/feed.jsonl
Curated feed
Tier-tagged rulings, opinions, and actions.
score.py
Weighted scorer
Deterministic, point-in-time, provenance-transparent.
build.py
Static page
Self-contained HTML + data.json. Served by Pages.
calibration.py
Backtest & audit
Snapshots, called-vs-actual, immutable run artifacts.
Runs in the background · GitHub Actions, weekly, zero infrastructure — or local watch mode.  Phase 2 seam · allowlist-only harvester → admission gate (anti-noise) → Voyage voyage-law-2 embeddings into Supabase pgvector.
Why it matters

What it isn’t

It isn’t legal advice, and it isn’t clairvoyance. The doctrine belongs to the courts and the bars, and a good knowledge lawyer already sees what’s coming.

What it is

Modeling, research, and calibration applied to legal AI. It weights the authorities, tracks where the pressure builds, and grades its own calls. Preparation you can replay, instead of one partner’s gut that walks out the door when they leave.

The defensible trail

When a regulator or a carrier asks “why didn’t you prepare for this,” a calibrated, evidence-graded record answers. Intuition can’t be shown to a board.

Fits the pillars

Deterministic replay, explainable, traceable. The radar turns legal-os’s own governance discipline on the future.

Still early, and I check it against reality every run. If you live in this and see a fault line I’ve mis-weighted or missed, I want to hear it. That’s the whole point of keeping score in the open.