↓ Skip to main content

Posts

88% Recall, One Attorney, 18 Hours

A new working paper ran generative AI document review and a managed active-learning workflow head to head on the same 45,004-document corpus. Same review protocol, same reference labels, same scorecard. The GenAI system won on recall, and the paired test backs it up. The rest of the scorecard – precision, the 62-to-1 effort gap, the population extrapolation – needs more qualification than the headline suggests.

Building a Medicare Fraud Backtest in One Claude Code Session

A walkthrough of building a Medicare fraud backtest overnight in Claude Code – from a plain-English spec to 289 matched providers across 41 states, a fraud-similarity model with AUC 0.79, and a manual public-record check of high-scoring peers. Including the three times the pipeline failed, the data duplication bug, and the engineering decisions that shaped the final design.