Every major legal AI vendor shipped autonomous agents in Q1 2026. Here's what they actually do, what can go wrong, and why your ethical walls weren't built for this.
The PPP fraud pipeline worked because the SBA released unusually inspectable data. Medicare's public data is fragmented, de-identified, and missing the features detection needs. Here's what exists on GitHub, where it falls short, and what CMS would need to release to make outside healthcare-fraud analysis more practical.
Public PPP data can produce enforcement-relevant anomaly maps. An open-source fraud-scoring system, run against the SBA PPP dataset, surfaced lender and geographic concentrations that overlap with known enforcement patterns — while also showing why public data cannot prove fraud by itself.
Chinese labs aren't just catching up — they're pioneering the techniques Western models adopt, sharing them under open licenses, and training them on chips that weren't supposed to exist.
LLMs learned language without ever encountering what language refers to. That gap between syntax and semantics explains why they fabricate citations, can't count letters, yet write flawless code.
A proof of concept for using AI and synthetic SAR data to estimate how many FCPA enforcement actions describe transaction patterns that would plausibly generate suspicious activity reports — and what that tells us about the hidden plumbing of anti-corruption enforcement.
Medicare claims, tax returns, PPP and EIDL applications — the government increasingly holds the structured transaction data that lets fraud enforcement start with a query, not a tip.
LLMs ace simple retrieval benchmarks but collapse on the tasks that matter in fraud investigations — finding semantically disguised evidence buried in millions of documents