The LawBlogger Research Lab publishes the benchmarks and methodology behind which models make it onto the panel.
We test every candidate model against a held-out set of questions with known, verified citations, and measure how often its cited sources actually say what it claims.
Models are checked for whether they correctly distinguish between state and federal law, and between jurisdictions with conflicting rules.
We measure how often model disagreement actually predicts a wrong answer, so our confidence notes mean something.
Every model on the panel is re-benchmarked on a rolling basis; models that fall behind are retired.