Research

How we evaluate legal AI models

The LawBlogger Research Lab publishes the benchmarks and methodology behind which models make it onto the panel.

Citation accuracy benchmark

We test every candidate model against a held-out set of questions with known, verified citations, and measure how often its cited sources actually say what it claims.

Jurisdictional consistency

Models are checked for whether they correctly distinguish between state and federal law, and between jurisdictions with conflicting rules.

Disagreement calibration

We measure how often model disagreement actually predicts a wrong answer, so our confidence notes mean something.

Continuous re-evaluation

Every model on the panel is re-benchmarked on a rolling basis; models that fall behind are retired.