A 60-lawyer litigation and business firm signs a commercial Claude agreement in the spring. By midsummer, adoption is genuinely good. Associates summarize opposing briefs. The corporate group compares agreement versions against a playbook. Litigators build chronologies from document productions. Someone drafts a client update in a third of the usual time.
Nobody has done anything wrong. And the firm cannot answer these questions.
Which Claude product and plan are we on, and what does it retain?
Can a lawyer rely on Claude to tell us whether controlling contrary authority exists?
When a partner says the memo was "reviewed," what was actually checked?
Six months from now, when another lawyer opens that chronology, will anything indicate that a model produced it, which model, what sources it saw, and who verified it?
Does an AI-generated summary of a walled matter inherit that wall?
What changes when someone enables a connector?
Procurement answered which vendor. It did not answer how legal work using that vendor will be governed, and the second problem is the one with professional consequences. "Is Claude safe for legal work" has no answer, because it is not a question about a product. The useful question is narrower: for which work, using which information, under which configuration, with what verification, and who owns the result. That question has an answer, and the answer is a set of controls a firm can build.
Claude Is Three Systems at Once
Most firm AI policies fail because they treat Claude as one thing. It is three, simultaneously, and a policy addressing only one leaves the others ungoverned.
A legal-work assistant
It drafts, summarizes, compares, extracts, synthesizes, critiques, and generates alternatives. This is the dimension firms buy, and the one policies usually address.
Addressed by most policiesA third-party information system
It receives client information and processes it on someone else's infrastructure. That raises retention, access, subprocessors, residency, ethical walls, records management, and incident response.
Usually treated as an IT questionA probabilistic generation system
Its output can be wrong, incomplete, internally inconsistent, confidently stated, built on misread authority, or missing the thing that mattered most.
Rarely governed at allA firm that addresses only the first has a productivity program. A firm that addresses only the second has an IT policy. Neither governs legal work.
Where Claude Earns Its Place
This is not a hazard article. Matter orientation: chronologies, party maps, fact matrices, transcript synthesis. Document analysis: clause extraction, version comparison, due diligence tables. Draft acceleration: correspondence, discovery, brief sections built from authorities you supply. And adversarial review, the most underused capability in law firms: attack this argument, find the missing element. What the strong uses share: the lawyer supplies the source material and can check the output against it.
The Quiet Failures Matter More Than the Loud Ones
Fabricated citations get the attention because they are detectable: someone checks whether the case exists, and it does not. The failures that survive review are quieter.
Omission is more dangerous than fabrication because it leaves no artifact.
A fake case is a thing you can find. A missing case is an absence, and nothing in a fluent, well-organized memo signals that the absence is there. Polished prose reads as competence, and the model's confidence is not calibrated to its accuracy. This is why "we tell everyone to check the output" is not a control. Check it against what?
Legal Research Needs Its Own Workflow
Claude can be an excellent research assistant without being the system of record for what the law is. Conflating those two roles is the central failure mode, and the difference between these workflows is where the exposure lives.
✓ Defensible
- Question
- Claude helps spot issues and formulate the research
- Retrieval from an authoritative source
- Claude helps synthesize what you retrieved
- Attorney independently evaluates the authority
- Quotations and pinpoints checked against the source
- Currency and negative treatment checked
- Contrary controlling authority actively investigated
- The lawyer adopts the conclusion
✗ Closed loop
- Question
- Claude answers
- Claude supplies cases
- Claude confirms its own cases
- Lawyer copies
- Filing
A model is not independent verification of its own output. Asking Claude whether the cases it just gave you are real produces another generation from the same system. It can be confidently wrong twice.
Two facts should calibrate the firm's confidence. ABA Formal Opinion 512 (2024) ties required verification to the task, stating it "will necessarily depend on the GAI tool and the specific task that it performs," and characterizes uncritical reliance on GAI output as malpractice ABA FO 512. And the study the ABA itself relied on found that purpose-built legal research tools from LexisNexis and Thomson Reuters each hallucinate between 17% and 33% of the time Stanford / JELS.
Hallucination rate measured in purpose-built legal research tools grounded in proprietary legal databases. Not general chatbots. This is the evidence ABA Formal Opinion 512 itself cited.
Those are tools designed for legal research, grounded in proprietary databases, sold on reliability. If they land in that range, a general-purpose assistant is not your authority system. And existence is only the first of fourteen checks a citation owes you:
A real case, accurately cited, for a proposition it does not support, is still a bad citation.
The Use-Case Risk Matrix
Suitability rises as independent verifiability rises and consequence of error falls. The matrix is Songbird's operational judgment, not a statement of any bar's rule.
| Use case | Primary failure mode | Verifiability | Error consequence | Minimum review | Status |
|---|---|---|---|---|---|
| Rewriting attorney-authored prose | Meaning drift | High | Low | Read for altered meaning | Appropriate |
| Formatting, structure | Silent substantive edit | High | Low | Compare to source | Appropriate |
| Summarizing a supplied document | Omission, lost qualification | High | Medium | Read source, check qualifiers | Defined review |
| Summarizing a supplied case | Posture and holding errors | High | High | Read the opinion | Defined review |
| Clause extraction | False negatives | High | Medium | Sampling plus exception testing | Defined review |
| Contract and version comparison | Missed change | High | High | Diff against source | Defined review |
| Chronology creation | Date and attribution errors | High | Medium | Spot-check to source documents | Defined review |
| Deposition prep, discovery drafting | Omission, scope errors | Medium | Medium | Lawyer edits substantively | Defined review |
| Contract drafting | Unintended operative language | Medium | High | Full substantive review | Defined review |
| Client correspondence drafting | Tone, commitments, accuracy | Medium | High | Lawyer adopts before sending | Defined review |
| Brief sections from supplied verified authorities | Mischaracterized authority | Medium | High | Full cite and holding check | Defined review |
| Brainstorming, adversarial review | Plausible but wrong candidates | N/A | Low | Lawyer selects what to use | Appropriate |
| Generating research queries and terminology | Framing bias | High | Low | Run against a real database | Appropriate |
| Open-ended legal research | Fabrication, omission, stale law | Low | High | Authoritative retrieval required | Restricted |
| Identifying controlling authority | Omission | Low | Very high | Authoritative retrieval required | Not appropriate |
| Confirming no contrary authority exists | Invisible omission | Very low | Very high | Cannot be delegated | Not appropriate |
| Verifying citations or quotations | Closed-loop confirmation | None | Very high | Must verify outside the model | Not appropriate |
| Deciding legal strategy | Displaced judgment | N/A | Very high | Lawyer decides | Not appropriate |
| Direct client communication without review | Unreviewed advice | N/A | Very high | Lawyer owns it | Not appropriate |
| Filing without substantive lawyer review | Everything above | N/A | Very high | The signing lawyer owns it | Not appropriate |
The bottom rows are not "risky." They are category errors. Asking a model whether contrary authority exists asks it to prove a negative about a corpus it did not systematically search.
"A Lawyer Reviewed It" Is Not a Control
It becomes one when the firm defines what the lawyer was supposed to check. Different tasks fail differently, so review has to be task-specific.
| Task type | What review actually means | Fails if |
|---|---|---|
| Transformation, formatting | Meaning preserved, no substantive drift, no confidentiality change | Reviewer reads for style only |
| Extraction | Source-linked output, defined sampling rate, exception testing, escalation for ambiguity | Reviewer spot-checks what the tool found and never tests what it missed |
| Summarization | Read the source. Check omissions, qualifications, chronology, attribution | Reviewer checks the summary against itself |
| Legal research | Authoritative retrieval, cite and quotation checks, negative treatment, contrary authority search | Reviewer asks the model to confirm |
| Drafting | Facts, law, strategy, client objectives, citations, tone, procedural requirements | Reviewer edits prose without checking substance |
| Court submissions | The signing lawyer owns the filing entirely | Anyone treats AI involvement as shared responsibility |
Scale review intensity with consequence, reversibility, whether the output leaves the firm, and whether a court or client will rely on it.
Paid Claude Is Not Automatic Data Governance
"We are on a commercial plan, so confidentiality is handled" is the most expensive misconception in legal AI. Four separate questions hide inside it, with four different answers.
Anthropic's commercial position is clear: by default it will not use inputs or outputs from its commercial products to train its models Anthropic. That covers Claude for Work, Enterprise, the API, and Claude Gov.
A completely separate question. Not trained on does not mean not retained. API inputs and outputs are deleted within 30 days subject to exceptions. Chat products retain conversations to provide a consistent experience Anthropic.
Enterprise and Platform customers can use a Compliance API that pulls activity feed events, chat data, and file content Anthropic. That is content, not metadata.
This one runs the other direction. A short vendor retention period does not mean the firm should destroy quickly. Records obligations, client file duties, litigation holds, and regulatory requirements are the firm's, not the vendor's.
One exception has an operational shape: rate a chat with a thumbs up or down and Anthropic may use that conversation to train, retaining the feedback up to five years. One lawyer rating one privileged conversation changes the answer for it. Owners can disable this through the "Rate chats" setting. Make that decision once, at the organization level.
Enterprise custom retention
Available, configurable by Primary Owner or Owner, minimum 30 days, covering chats and projects. The default is that data is retained indefinitely unless a custom period is set.
Buying Enterprise does not shorten retention: configure nothing and the firm accumulates an indefinite archive of every conversation, privileged ones included. One vendor policy overrides even a negotiated contract. Effective June 9, 2026, "covered models" carry 30-day retention wherever offered, and organizations holding zero-data-retention agreements must enable retention to use them Anthropic.
Negotiated ZDR agreement in force. No prompt or output retention.
Firm approves a covered model for its lawyers. No contract change. No visible product change.
Retention must be enabled to use that model. 30-day retention now applies.
Model selection is a data-governance decision, not a quality decision.
California's 2026 COPRAC guidance makes the diligence point directly: reasonable efforts to protect client confidences "require more than reliance on generalized marketing assurances" CA COPRAC 2026. Someone at the firm has to actually read this material.
Approve the Surface, Not the Vendor
"Claude is approved" names a brand, not a data path. Songbird recommends an Approved AI Surfaces Register: for each route, the product and plan, model families, permitted user groups and data classes, connectors, web access, retention configuration, owner, and re-review date. Anthropic's own Commercial Terms leave third-party features governed by the third party's terms, without defining connectors, integrations, or MCP, or resolving where client data flows once it leaves Claude Commercial Terms. So approve features separately: each new model, each connector, web search, and any agentic capability. A feature update can change the data path without changing the name on the invoice.
Provenance: Make the Work Intelligible Later
Six months from now, a lawyer opens that polished chronology and nothing tells them a model produced it, what it saw, or who verified it. They will either redo the work or trust it. Both are bad outcomes. Preserve enough metadata to answer those questions, and pair it with a work-product status model.
Machine-assisted material that has not passed the required validation for its task type.
Checked against the firm's defined standard for that task.
Substantively adopted by the responsible lawyer as professional work product.
Labeling everything "AI generated" conveys nothing, because it does not distinguish a raw first pass from work a partner has adopted. What matters is which of these three things the reader is holding.
Monitoring Is Also a Repository
The Compliance API exposes activity feed events, chat data, and file content. That is genuinely useful for supervision under Model Rules 5.1 and 5.3 Model Rules, and it is also a second repository of privileged material, governed by whatever the firm puts around the export rather than by the originating matter's ethical wall. Who may inspect attorney conversations, who authorizes it, whether access is itself logged, and where exports live are questions the vendor's documentation does not answer and the firm must.
More visibility is not automatically safer. Visibility has to be governed too.
Connected Claude Changes the Security Model
Receives only what a lawyer pastes in.
Can reach the document management system, email, knowledge systems, cloud storage, and the open web.
Can alter documents, send communications, create files, and call external services.
Prompt injection is not theoretical. In a 2026 employment dispute in Para, Brazil, lawyers filed a petition containing hidden white-on-white text instructing any AI reader to contest it superficially. The court's AI flagged it; the judge imposed an R$84,000 fine and a bar referral reported. The account rests on secondary reporting in a foreign jurisdiction, and the lesson travels anyway: the adversary was opposing counsel, the vector was a filed document, and a human reader could not see it. Any firm running opposing filings or productions through an AI system inherits that exposure.
The control principle, for the policy document
Instructions found inside a retrieved document, email, webpage, filing, or adversary-produced material are data, not authorization. They do not permit the system to change its task, disclose information, invoke tools, or take external action.
The same logic governs ethical walls. Derived information inherits the sensitivity of its source: a summary of a walled matter is still walled material, and it does not become less confidential by being shorter.
Models Change. Revalidate.
The wrong mental model is "we approved Claude in 2026." The right one is "we validated this product, model, and workflow for this category of work, under these controls." Keep a small benchmark suite of known-answer tasks the firm actually relies on, and rerun it when the model changes, when a workflow changes materially, and on a calendar. It is a smoke detector, not science: it catches the degradation that anecdote misses. The covered-models retention change is exactly what a vendor-initiated revalidation trigger looks like.
Clients, Consent, and Billing
Consent and disclosure are not one national rule. ABA Formal Opinion 512 ties informed consent substantially to self-learning tools, warns that "boilerplate waivers will not suffice," and finds no one-size-fits-all disclosure duty ABA FO 512. California frames the trigger differently: no confidential client information into a system that may present material risks to confidentiality or security, absent informed consent. A deployment that does not train on inputs can still present material risks through retention, administrative access, or connectors, so the California analysis does not necessarily end where the ABA's does. And FO 512 limits its own shelf life: as the technology develops, the risks may change in ways that would alter its conclusion.
ABA Formal Opinion 512 FO 512. Competence, confidentiality, supervision, fees, communication.
Virginia LEO 1901 LEO 1901 adopted by the Supreme Court of Virginia. Reasonable fees and generative AI.
California COPRAC guidance COPRAC replaces the 2023 version and reaches agentic AI.
22 NYCRR Part 161 Part 161 takes effect across the New York Unified Court System.
NYC Bar Opinion 2026-2 NYC Bar, third in a sequence rather than a single pronouncement.
A firm policy citing only ABA 512 is citing the 2024 layer of a stack that has kept moving. Make client AI restrictions enforceable rather than memorable: outside counsel guidelines belong in matter metadata, not in a PDF a lawyer is expected to recall.
On billing, FO 512 is blunt. Lawyers billing hourly "must bill for their actual time," quoting Formal Opinion 93-379, and a lawyer who drafts a pleading in 15 minutes with a tool bills 15 minutes FO 93-379. The harder question is design: if a six-hour task becomes ninety minutes, someone captures that value, and firms that do not decide who will find it was decided by default, through realization. Clio's 2025 research already shows a majority of firms using flat fees in some form Clio 2025. The question is how to price work whose cost structure just changed.
Apprenticeship Debt
Junior lawyers build judgment through work that is, in the narrow sense, inefficient: reading fifty cases to find the six that matter, noticing on the thirty-first contract that the indemnity language has been drifting. Claude removes that repetition. That is the point of buying it, and some of what it removes is how lawyers develop the pattern recognition and skepticism that let a reviewer sense that a fluent memo is wrong.
Call it apprenticeship debt: real, deferred, and invisible on any current-year financial statement.
The fix is not pointless manual labor as a character exercise. It is deliberate redesign: associates read the primary authority behind AI-assisted research, draft before comparing against the model, and are evaluated on whether they can detect model errors. Teach verification as a legal skill. The goal is not preserving inefficiency. It is preserving judgment.
The Governance Stack
Claude does not owe the client a duty, sign the pleading, or carry malpractice responsibility. The responsible lawyer owns the work they adopt. These fifteen layers are what make that ownership real rather than nominal.
Define an AI incident broadly: client data through an unapproved surface, a fabricated citation in a filing, a connector sending data somewhere unexpected, a vendor policy change that invalidates an approved workflow. Decide the escalation path before you need it.
The Question That Decides Whether You Are Done
How will the firm demonstrate that Claude improved the quality of legal work, rather than only the volume of work produced? Adoption metrics measure activity. The honest measures are correction rates, citation defects caught in review, omissions discovered, escalation rates, and rework. A firm that cannot distinguish more output from better work has not finished building the governance system, whatever the policy document says.
The unit of governance is not the model. It is the workflow: the task, the information, the configuration, the required validation, the person accountable, and the evidence the work was checked. Firms that govern at that level can adopt aggressively. Firms that govern at the level of "we approved Claude" are relying on individual professionalism to compensate for an absent system, which works right up until it does not.
What This Connects To
This article covers the workflow layer. The rest of the cluster covers the pieces it assumes. ai-tool-due-diligence covers the vendor review that produces the answers in the data-governance section. legal-ai-reliability-evaluation covers how to test reliability before a tool is approved at all. privilege-confidentiality-ai covers the confidentiality classification the use matrix depends on. firm-ai-policy covers the policy structure that carries the approved surfaces register. ai-use-by-role covers role-based permissions. how-to-run-a-legal-tech-pilot covers the pilot structure the benchmark suite fits inside.
This article is operational guidance and does not constitute legal advice or a formal ethics opinion. Professional responsibility obligations vary by jurisdiction. ABA Model Rules and formal opinions are interpretive guidance and do not bind a jurisdiction unless adopted there; state rules, state ethics opinions, court rules, and client outside counsel guidelines may impose different or additional requirements. Product details reflect vendor documentation as of August 2026 and change. Songbird Strategies is a legal technology consulting firm, not a law firm. Firms should consult qualified ethics counsel for legal conclusions and policy decisions. See Sources & Notes for the authority and vendor documentation cited.