
From TAR to Transformers#
TL;DR
- E-discovery is the only completed AI adoption cycle in law. From the Grossman & Cormack study (2011) through Da Silva Moore and Rio Tinto to routine use today, it took a decade of empirical validation, judicial buy-in, and practitioner resistance before technology-assisted review became standard. Every other legal AI use case is in the first two years of that same curve.
- TAR’s validation model doesn’t transfer cleanly to generative AI. Recall and precision work because relevance is binary. Drafting quality, research accuracy, and contract analysis don’t reduce to a single metric — and the profession hasn’t built equivalent validation protocols for any of them.
- The 26(f) conference is where AI adoption actually happens. TAR didn’t become standard because judges mandated it. It became standard because producing parties proposed it at the meet-and-confer and courts declined to second-guess the methodology. The same mechanism is available now for GenAI-assisted review — and the early stipulated ESI protocols are already appearing.
- Twelve years from proof to routine use. Adjust expectations accordingly. The firms that adopted TAR early didn’t just save on review costs — they changed their litigation strategy. The same strategic advantage is available now, but only for firms that measure first and stop waiting for perfection.
In 2011, Maura Grossman and Gordon Cormack published a study demonstrating that technology-assisted review outperformed exhaustive manual review on both recall and precision. Not by a slim margin — TAR was consistently better at finding relevant documents and better at excluding irrelevant ones. The finding should have ended the debate. Instead, it started one that took a decade to resolve.
Fifteen years later, most firms have absorbed the e-discovery lesson. TAR is standard practice for document review. But the profession hasn’t applied the lesson anywhere else. Contract review, legal research, brief drafting — areas where generative AI is producing results comparable to TAR’s early demonstrations — remain in the same “promising but unproven” limbo that e-discovery occupied circa 2012. The objections sound identical: accuracy concerns, “black box” distrust, malpractice anxiety, the insistence on human review of every output.
E-discovery is the legal profession’s only completed AI adoption cycle. The pattern it traced — empirical proof, judicial endorsement, practitioner resistance, gradual acceptance, routine use — is the closest thing lawyers have to a roadmap for what comes next.
The TAR Decade#
The timeline matters because it reveals how slowly the profession absorbs even well-proven technology.
2011: The empirical proof. Grossman and Cormack’s JOLT study tested TAR against manual review across multiple document sets. TAR achieved higher recall (finding more relevant documents) and higher precision (excluding more irrelevant ones) than teams of human reviewers. The study gave the profession something it had never had for any technology: controlled, reproducible evidence that the machine outperformed the human baseline.
2012: The first judicial endorsement. In Da Silva Moore v. Publicis Groupe, Magistrate Judge Andrew Peck of the Southern District of New York approved the use of predictive coding for document review — the first federal court to do so explicitly. Peck’s opinion didn’t mandate TAR. It said something more important: a producing party’s reasonable choice to use predictive coding would not be second-guessed by the court, provided the party was transparent about its methodology. The standard was reasonableness, not perfection.
2015: Continuous active learning reaches the docket. In Rio Tinto PLC v. Vale S.A., 306 F.R.D. 125 (S.D.N.Y. 2015), Magistrate Judge Andrew Peck — the same judge who decided Da Silva Moore — approved the parties’ stipulated TAR protocol and discussed the practical implications of continuous active learning, where the model updates as reviewers code documents rather than training on a fixed seed set. Peck wrote that the case law had developed to the point where TAR was black-letter law, while stressing that what a producing party must disclose remains protocol- and context-specific.
2015–2020: The resistance phase. Despite Grossman & Cormack’s data and two significant judicial endorsements, adoption was slow. The Lighthouse AI in eDiscovery Report documents the pattern that persisted for years: budget constraints, lack of training, and trust issues. Senior partners who had built careers on manual review resisted a technology that implied their previous approach was suboptimal. Review vendors whose revenue depended on billing thousands of contract reviewer hours had no incentive to promote efficiency. And many litigators simply didn’t trust a process they couldn’t see — even after the data showed it produced better results than the process they could.
2020–present: Routine use. TAR became standard not because the holdouts were convinced but because the economics became undeniable. A matter requiring review of 500,000 documents costs roughly $1.5–2 million with manual review teams. TAR cuts that by 60–80%. As EDRM’s 2024 review transition report documented, the remaining resistance collapsed once clients started asking why they were paying for manual review when a proven, judicially endorsed alternative existed.
The entire arc — from controlled study to standard practice — took roughly twelve years. Not because the technology was unproven, but because the profession needed time to build the institutional infrastructure around it: validation protocols, judicial precedent, practitioner training, vendor tooling, and the proportionality framework under amended Rule 26(b)(1) that gave lawyers a structured way to argue that AI-assisted methods satisfy discovery obligations.
What Transfers and What Breaks#
The TAR decade produced a repeatable adoption pattern. Some of it applies directly to generative AI. Some of it doesn’t — and the differences matter more than the similarities.
What transfers#
Measure against the human baseline, not perfection. Grossman and Cormack didn’t prove TAR was perfect. They proved it was better than the alternative. Manual review by contract attorneys achieves recall rates around 60% — meaning 40% of relevant documents get missed. TAR consistently beat that number. The profession accepted TAR not because it was flawless but because the existing method was worse and more expensive. The same framing applies to generative AI: the question isn’t whether an LLM's contract summary is perfect, but whether it’s better than a first-year associate’s at 2 AM on a Friday before a Monday closing.
The judicial acceptance pattern. One magistrate judge in a permissive opinion → silence from appellate courts → gradual adoption by other trial courts → routine use. Da Silva Moore and Rio Tinto were decided by magistrate judges, not circuit courts. No appellate court weighed in to endorse or reject TAR. The adoption spread through district courts citing each other, building a body of practice that never needed appellate validation because no one appealed. Generative AI is tracing the same path. United States v. Heppner (S.D.N.Y. 2026) applied existing privilege doctrine to AI tools without inventing new rules — exactly the pattern Da Silva Moore set for TAR.
The proportionality framework. Rule 26(b)(1)’s requirement that discovery be “proportional to the needs of the case” gave TAR its legal foundation. A producing party could argue that technology-assisted review was a reasonable, proportional method — and courts agreed. That framework is a rule of federal discovery, not a general licence for AI-assisted work. Outside litigation it survives only as an economic analogy: a firm using an LLM to extract key terms from 500 contracts in a due diligence review is making a reasonableness argument to its client, not to a court. Rule 26 supplies the vocabulary; the engagement letter and the standard of care supply the obligation.
What breaks#
Recall and precision don’t generalize. TAR’s validation framework works because relevance is treated as binary. A document is coded responsive to a discovery request or it isn’t. Recall (what share of relevant documents did the system find?) and precision (what share of those it flagged were actually relevant?) can then be estimated — though the estimates rest on human coding judgments and sampling, so they carry error bars rather than certainty. A contract summary, a legal research memo, or a draft brief has no equivalent single-axis metric. A brief can cite every case correctly and still be unpersuasive. A contract review can capture every defined term and miss the commercial intent. LegalBench attempts to fill this gap with 162 tasks, but it measures pattern-matching on short legal problems — not the multi-dimensional quality judgment that practitioners actually apply.
Nondeterminism. TAR produces the same results on the same inputs every time. LLMs don’t. Run the same document through the same prompt twice and you’ll get slightly different output. For classification — is this document responsive? — the variance is usually harmless. For extracting a specific dollar figure or a deadline date that a deal team will rely on, it’s a problem TAR never had. The nondeterminism issue is solvable (lower the model’s temperature, use structured extraction, validate outputs against source), but it requires engineering work that the profession hasn’t standardized.
The measurement gap. Before Grossman & Cormack, the profession had no agreed-upon way to measure whether TAR worked. Their study created the measurement infrastructure. For generative AI across contract review, legal research, and drafting, that infrastructure doesn’t exist yet. legalbenchmarks.ai is the most serious attempt — their Phase 2 research found that specialized tools didn’t always produce better drafts than general-purpose LLMs, but offered better workflow integration. The profession needs its own Grossman & Cormack moment for each major use case: a controlled, reproducible study that establishes what “good enough” looks like. Until then, every adoption argument is qualitative.
The 26(f) Conference: Where Adoption Actually Happens#
TAR didn’t become standard because judges mandated it. It became standard because producing parties started proposing it at the Rule 26(f) conference — the meet-and-confer where parties agree on the scope, format, and methodology of discovery — and courts declined to second-guess reasonable methodological choices.
This is the mechanism that matters. The 26(f) conference is where abstract questions about AI adoption become concrete: What review methodology will the producing party use? What validation will it apply? What will it disclose to opposing counsel about the process? Read together, Da Silva Moore and Rio Tinto gave producing parties a workable template: propose a method, be transparent about the protocol, and expect the court to engage with the protocol rather than impose one. Both were context- and protocol-dependent decisions rather than general rules, and neither forecloses a challenge to how a method was actually executed.
The same mechanism is available now for generative AI-assisted review. The conversation at the 26(f) conference hasn’t changed structurally — only the technology under discussion.
What a GenAI disclosure looks like. The producing party identifies the technology (e.g., Relativity aiR, Everlaw AI, or a custom LLM pipeline), describes the workflow (initial classification via TAR, augmented by LLM-generated relevance rationales, with human review on a statistical sample), and commits to a validation protocol (precision and recall measurements on the sample, with results shared with opposing counsel). The disclosure tracks what TAR adopters proposed a decade ago — the methodology is different, the transparency framework is identical.
What opposing counsel’s objections sound like. They sound like 2012. “We can’t trust a black box.” “The AI might miss critical documents.” “There’s no precedent for this method.” TAR litigators heard every one of these objections and overcame them with the same response: the producing party’s methodology is reasonable, the validation is transparent, and manual review is neither more accurate nor more proportional. The Grossman & Cormack data was the trump card then. For GenAI-assisted review, the early platform data — Relativity reporting 85% faster reviews and 10–20% more relevant documents surfaced, Epiq claiming 90% faster than traditional TAR — serves a similar function. It’s vendor-reported, not independently validated, but it’s enough to support a proportionality argument at the meet-and-confer.
What stipulated ESI protocols are starting to address. Early protocols covering GenAI-assisted review address concerns TAR protocols didn’t need to: which foundation model the review platform uses, whether documents are sent to a third-party API or processed within the platform’s own infrastructure, data retention terms, and whether the model’s training data could include documents from the same litigation. These are new questions, but they fit within the existing ESI protocol structure — additional line items, not a new framework.
The 26(f) conference is where adoption gets decided. Not in law review articles, not at conferences, not in vendor demos. A litigator who understands the TAR precedent and can articulate a defensible GenAI review methodology at the meet-and-confer has the same advantage early TAR adopters had: lower costs, faster review, and a proportionality argument that opposing counsel will struggle to counter without proposing something more expensive and less accurate.
What the Timeline Teaches#
Twelve years from Grossman & Cormack to routine use. That timeline included controlled empirical evidence (2011), judicial endorsement from a respected magistrate judge (2012), validation of the improved methodology (2014), a supporting amendment to the Federal Rules (2015), and still the profession took another five years to make it standard.
Generative AI for contract review, legal research, and drafting is roughly where TAR was in 2013: strong early results, a handful of firms using it seriously, a much larger group watching from the sidelines, and no equivalent of the Grossman & Cormack study establishing what “good enough” looks like for any use case beyond e-discovery. The Lighthouse 2025 survey of 225 e-discovery professionals found the same adoption barriers slowing generative AI everywhere else — budget constraints, lack of training, trust issues. The FTI/Relativity General Counsel Report showed AI adoption doubling among corporate legal departments, but from a small base.
The firms that adopted TAR early didn’t just save money on document review. They changed their litigation strategy. Faster review meant earlier case assessment. Earlier case assessment meant better deposition questions, more targeted motions, and earlier settle-or-fight decisions. The cost savings were real — $1 million or more on large matters — but the strategic advantage of faster information processing mattered more.
The same shift is available now. A corporate team that can review a 500-document due diligence set in days instead of weeks makes better deal decisions. A litigation team that can summarize six depositions before the key witness’s deposition asks better questions. A regulatory practice that scans a year of enforcement actions every morning spots patterns that monthly manual reviews miss.
But only if the profession applies the e-discovery lesson: measure against the human baseline, not against perfection. Build validation protocols before you need them. Use the existing procedural frameworks — proportionality, the 26(f) conference, the producing party’s methodological discretion — rather than waiting for new ones. And adjust expectations for a timeline measured in years, not quarters.
The TAR decade proved that the legal profession can adopt AI tools successfully. It also proved that “successfully” means slowly, skeptically, and only after the economics become impossible to ignore. The evidence for generative AI across legal practice is accumulating faster than TAR’s evidence did — more firms, more use cases, more data. The profession’s absorption rate is the bottleneck, not the technology. It was last time, too.
Further Reading#
- Grossman & Cormack, “Technology-Assisted Review in E-Discovery Can Be More Effective and More Efficient Than Exhaustive Manual Review” (2011). The foundational study that started the TAR era.
- Da Silva Moore v. Publicis Groupe (S.D.N.Y. 2012). First federal judicial endorsement of predictive coding.
- Rio Tinto PLC v. Vale S.A., 306 F.R.D. 125 (S.D.N.Y. 2015). Judge Peck on continuous active learning and the state of TAR case law.
- EDRM, “eDiscovery Review in Transition” (2024). Comparison of manual review, TAR 1.0/2.0, and GenAI-assisted review methods.
- Lighthouse AI in eDiscovery Report 2025. Survey of 225 legal professionals on adoption barriers.
- FTI/Relativity General Counsel Report 2026. AI adoption data from 200+ general counsel.
- legalbenchmarks.ai Evaluation Framework. Open-access legal AI evaluation toolkit.
- Winter 2026 eDiscovery Pricing Survey. ComplexDiscovery and EDRM pricing benchmark.
This is a standalone post on LegalRealist AI. It is intended for informational and educational purposes only and does not constitute legal advice. Court rulings, vendor capabilities, and adoption data described here reflect publicly available information as of the publication date and are subject to change. E-discovery obligations and technology-assisted review standards vary by jurisdiction.



