AI & Confidentiality

Your Firm Approved Claude. What Has to Exist Before It Touches a Matter.

A workflow-level governance model for firms that have already bought commercial Claude: approved surfaces, task-specific review standards, research verification, provenance, and the questions a firm should be able to answer before the work reaches a client.

A 60-lawyer litigation and business firm signs a commercial Claude agreement in the spring. By midsummer, adoption is genuinely good. Associates summarize opposing briefs. The corporate group compares agreement versions against a playbook. Litigators build chronologies from document productions. Someone drafts a client update in a third of the usual time.

Nobody has done anything wrong. And the firm cannot answer these questions.

Which Claude product and plan are we on, and what does it retain?

Can a lawyer rely on Claude to tell us whether controlling contrary authority exists?

When a partner says the memo was "reviewed," what was actually checked?

Six months from now, when another lawyer opens that chronology, will anything indicate that a model produced it, which model, what sources it saw, and who verified it?

Does an AI-generated summary of a walled matter inherit that wall?

What changes when someone enables a connector?

Procurement answered which vendor. It did not answer how legal work using that vendor will be governed, and the second problem is the one with professional consequences. "Is Claude safe for legal work" has no answer, because it is not a question about a product. The useful question is narrower: for which work, using which information, under which configuration, with what verification, and who owns the result. That question has an answer, and the answer is a set of controls a firm can build.

Claude Is Three Systems at Once

Most firm AI policies fail because they treat Claude as one thing. It is three, simultaneously, and a policy addressing only one leaves the others ungoverned.

01

A legal-work assistant

It drafts, summarizes, compares, extracts, synthesizes, critiques, and generates alternatives. This is the dimension firms buy, and the one policies usually address.

Addressed by most policies
02

A third-party information system

It receives client information and processes it on someone else's infrastructure. That raises retention, access, subprocessors, residency, ethical walls, records management, and incident response.

Usually treated as an IT question
03

A probabilistic generation system

Its output can be wrong, incomplete, internally inconsistent, confidently stated, built on misread authority, or missing the thing that mattered most.

Rarely governed at all

A firm that addresses only the first has a productivity program. A firm that addresses only the second has an IT policy. Neither governs legal work.

Where Claude Earns Its Place

This is not a hazard article. Matter orientation: chronologies, party maps, fact matrices, transcript synthesis. Document analysis: clause extraction, version comparison, due diligence tables. Draft acceleration: correspondence, discovery, brief sections built from authorities you supply. And adversarial review, the most underused capability in law firms: attack this argument, find the missing element. What the strong uses share: the lawyer supplies the source material and can check the output against it.

The Quiet Failures Matter More Than the Loud Ones

Fabricated citations get the attention because they are detectable: someone checks whether the case exists, and it does not. The failures that survive review are quieter.

Controlling contrary authority that never appears Law that was correct three years ago A holding from the wrong jurisdiction Dicta presented as holding A party's losing argument summarized as the court's reasoning A real case, correctly cited, for a proposition it does not support An accurate quotation attached to the wrong pinpoint A summary that drops the qualification that made the paragraph mean something Inconsistent treatment across a large document set

Omission is more dangerous than fabrication because it leaves no artifact.

A fake case is a thing you can find. A missing case is an absence, and nothing in a fluent, well-organized memo signals that the absence is there. Polished prose reads as competence, and the model's confidence is not calibrated to its accuracy. This is why "we tell everyone to check the output" is not a control. Check it against what?

Legal Research Needs Its Own Workflow

Claude can be an excellent research assistant without being the system of record for what the law is. Conflating those two roles is the central failure mode, and the difference between these workflows is where the exposure lives.

Defensible

  1. Question
  2. Claude helps spot issues and formulate the research
  3. Retrieval from an authoritative source
  4. Claude helps synthesize what you retrieved
  5. Attorney independently evaluates the authority
  6. Quotations and pinpoints checked against the source
  7. Currency and negative treatment checked
  8. Contrary controlling authority actively investigated
  9. The lawyer adopts the conclusion

Closed loop

  1. Question
  2. Claude answers
  3. Claude supplies cases
  4. Claude confirms its own cases
  5. Lawyer copies
  6. Filing

A model is not independent verification of its own output. Asking Claude whether the cases it just gave you are real produces another generation from the same system. It can be confidently wrong twice.

Two facts should calibrate the firm's confidence. ABA Formal Opinion 512 (2024) ties required verification to the task, stating it "will necessarily depend on the GAI tool and the specific task that it performs," and characterizes uncritical reliance on GAI output as malpractice ABA FO 512. And the study the ABA itself relied on found that purpose-built legal research tools from LexisNexis and Thomson Reuters each hallucinate between 17% and 33% of the time Stanford / JELS.

17% to 33%

Hallucination rate measured in purpose-built legal research tools grounded in proprietary legal databases. Not general chatbots. This is the evidence ABA Formal Opinion 512 itself cited.

Those are tools designed for legal research, grounded in proprietary databases, sold on reliability. If they land in that range, a general-purpose assistant is not your authority system. And existence is only the first of fourteen checks a citation owes you:

ExistenceCourtJurisdictionDateCitation format Quotation accuracyPinpoint accuracyThe actual holding Procedural postureRelevance to your propositionPrecedential status Current validityNegative treatmentContrary authority

A real case, accurately cited, for a proposition it does not support, is still a bad citation.

The Use-Case Risk Matrix

Suitability rises as independent verifiability rises and consequence of error falls. The matrix is Songbird's operational judgment, not a statement of any bar's rule.

Use case Primary failure mode Verifiability Error consequence Minimum review Status
Rewriting attorney-authored prose Meaning drift High Low Read for altered meaning Appropriate
Formatting, structure Silent substantive edit High Low Compare to source Appropriate
Summarizing a supplied document Omission, lost qualification High Medium Read source, check qualifiers Defined review
Summarizing a supplied case Posture and holding errors High High Read the opinion Defined review
Clause extraction False negatives High Medium Sampling plus exception testing Defined review
Contract and version comparison Missed change High High Diff against source Defined review
Chronology creation Date and attribution errors High Medium Spot-check to source documents Defined review
Deposition prep, discovery drafting Omission, scope errors Medium Medium Lawyer edits substantively Defined review
Contract drafting Unintended operative language Medium High Full substantive review Defined review
Client correspondence drafting Tone, commitments, accuracy Medium High Lawyer adopts before sending Defined review
Brief sections from supplied verified authorities Mischaracterized authority Medium High Full cite and holding check Defined review
Brainstorming, adversarial review Plausible but wrong candidates N/A Low Lawyer selects what to use Appropriate
Generating research queries and terminology Framing bias High Low Run against a real database Appropriate
Open-ended legal research Fabrication, omission, stale law Low High Authoritative retrieval required Restricted
Identifying controlling authority Omission Low Very high Authoritative retrieval required Not appropriate
Confirming no contrary authority exists Invisible omission Very low Very high Cannot be delegated Not appropriate
Verifying citations or quotations Closed-loop confirmation None Very high Must verify outside the model Not appropriate
Deciding legal strategy Displaced judgment N/A Very high Lawyer decides Not appropriate
Direct client communication without review Unreviewed advice N/A Very high Lawyer owns it Not appropriate
Filing without substantive lawyer review Everything above N/A Very high The signing lawyer owns it Not appropriate

The bottom rows are not "risky." They are category errors. Asking a model whether contrary authority exists asks it to prove a negative about a corpus it did not systematically search.

"A Lawyer Reviewed It" Is Not a Control

It becomes one when the firm defines what the lawyer was supposed to check. Different tasks fail differently, so review has to be task-specific.

Task type What review actually means Fails if
Transformation, formatting Meaning preserved, no substantive drift, no confidentiality change Reviewer reads for style only
Extraction Source-linked output, defined sampling rate, exception testing, escalation for ambiguity Reviewer spot-checks what the tool found and never tests what it missed
Summarization Read the source. Check omissions, qualifications, chronology, attribution Reviewer checks the summary against itself
Legal research Authoritative retrieval, cite and quotation checks, negative treatment, contrary authority search Reviewer asks the model to confirm
Drafting Facts, law, strategy, client objectives, citations, tone, procedural requirements Reviewer edits prose without checking substance
Court submissions The signing lawyer owns the filing entirely Anyone treats AI involvement as shared responsibility

Scale review intensity with consequence, reversibility, whether the output leaves the firm, and whether a court or client will rely on it.

Paid Claude Is Not Automatic Data Governance

"We are on a commercial plan, so confidentiality is handled" is the most expensive misconception in legal AI. Four separate questions hide inside it, with four different answers.

Training

Anthropic's commercial position is clear: by default it will not use inputs or outputs from its commercial products to train its models Anthropic. That covers Claude for Work, Enterprise, the API, and Claude Gov.

Retention

A completely separate question. Not trained on does not mean not retained. API inputs and outputs are deleted within 30 days subject to exceptions. Chat products retain conversations to provide a consistent experience Anthropic.

Access

Enterprise and Platform customers can use a Compliance API that pulls activity feed events, chat data, and file content Anthropic. That is content, not metadata.

Preservation

This one runs the other direction. A short vendor retention period does not mean the firm should destroy quickly. Records obligations, client file duties, litigation holds, and regulatory requirements are the firm's, not the vendor's.

One exception has an operational shape: rate a chat with a thumbs up or down and Anthropic may use that conversation to train, retaining the feedback up to five years. One lawyer rating one privileged conversation changes the answer for it. Owners can disable this through the "Rate chats" setting. Make that decision once, at the organization level.

Enterprise custom retention

Available, configurable by Primary Owner or Owner, minimum 30 days, covering chats and projects. The default is that data is retained indefinitely unless a custom period is set.

Buying Enterprise does not shorten retention: configure nothing and the firm accumulates an indefinite archive of every conversation, privileged ones included. One vendor policy overrides even a negotiated contract. Effective June 9, 2026, "covered models" carry 30-day retention wherever offered, and organizations holding zero-data-retention agreements must enable retention to use them Anthropic.

Before

Negotiated ZDR agreement in force. No prompt or output retention.

Event

Firm approves a covered model for its lawyers. No contract change. No visible product change.

After

Retention must be enabled to use that model. 30-day retention now applies.

Model selection is a data-governance decision, not a quality decision.

California's 2026 COPRAC guidance makes the diligence point directly: reasonable efforts to protect client confidences "require more than reliance on generalized marketing assurances" CA COPRAC 2026. Someone at the firm has to actually read this material.

Approve the Surface, Not the Vendor

"Claude is approved" names a brand, not a data path. Songbird recommends an Approved AI Surfaces Register: for each route, the product and plan, model families, permitted user groups and data classes, connectors, web access, retention configuration, owner, and re-review date. Anthropic's own Commercial Terms leave third-party features governed by the third party's terms, without defining connectors, integrations, or MCP, or resolving where client data flows once it leaves Claude Commercial Terms. So approve features separately: each new model, each connector, web search, and any agentic capability. A feature update can change the data path without changing the name on the invoice.

Provenance: Make the Work Intelligible Later

Six months from now, a lawyer opens that polished chronology and nothing tells them a model produced it, what it saw, or who verified it. They will either redo the work or trust it. Both are bad outcomes. Preserve enough metadata to answer those questions, and pair it with a work-product status model.

AI-DRAFT

Machine-assisted material that has not passed the required validation for its task type.

VALIDATED

Checked against the firm's defined standard for that task.

LAWYER-FINAL

Substantively adopted by the responsible lawyer as professional work product.

Labeling everything "AI generated" conveys nothing, because it does not distinguish a raw first pass from work a partner has adopted. What matters is which of these three things the reader is holding.

Monitoring Is Also a Repository

The Compliance API exposes activity feed events, chat data, and file content. That is genuinely useful for supervision under Model Rules 5.1 and 5.3 Model Rules, and it is also a second repository of privileged material, governed by whatever the firm puts around the export rather than by the originating matter's ethical wall. Who may inspect attorney conversations, who authorizes it, whether access is itself logged, and where exports live are questions the vendor's documentation does not answer and the firm must.

More visibility is not automatically safer. Visibility has to be governed too.

Connected Claude Changes the Security Model

Standalone

Receives only what a lawyer pastes in.

Retrieval-connected

Can reach the document management system, email, knowledge systems, cloud storage, and the open web.

Agentic

Can alter documents, send communications, create files, and call external services.

Prompt injection is not theoretical. In a 2026 employment dispute in Para, Brazil, lawyers filed a petition containing hidden white-on-white text instructing any AI reader to contest it superficially. The court's AI flagged it; the judge imposed an R$84,000 fine and a bar referral reported. The account rests on secondary reporting in a foreign jurisdiction, and the lesson travels anyway: the adversary was opposing counsel, the vector was a filed document, and a human reader could not see it. Any firm running opposing filings or productions through an AI system inherits that exposure.

The control principle, for the policy document

Instructions found inside a retrieved document, email, webpage, filing, or adversary-produced material are data, not authorization. They do not permit the system to change its task, disclose information, invoke tools, or take external action.

The same logic governs ethical walls. Derived information inherits the sensitivity of its source: a summary of a walled matter is still walled material, and it does not become less confidential by being shorter.

Models Change. Revalidate.

The wrong mental model is "we approved Claude in 2026." The right one is "we validated this product, model, and workflow for this category of work, under these controls." Keep a small benchmark suite of known-answer tasks the firm actually relies on, and rerun it when the model changes, when a workflow changes materially, and on a calendar. It is a smoke detector, not science: it catches the degradation that anecdote misses. The covered-models retention change is exactly what a vendor-initiated revalidation trigger looks like.

Clients, Consent, and Billing

Consent and disclosure are not one national rule. ABA Formal Opinion 512 ties informed consent substantially to self-learning tools, warns that "boilerplate waivers will not suffice," and finds no one-size-fits-all disclosure duty ABA FO 512. California frames the trigger differently: no confidential client information into a system that may present material risks to confidentiality or security, absent informed consent. A deployment that does not train on inputs can still present material risks through retention, administrative access, or connectors, so the California analysis does not necessarily end where the ABA's does. And FO 512 limits its own shelf life: as the technology develops, the risks may change in ways that would alter its conclusion.

Jul 2024

ABA Formal Opinion 512 FO 512. Competence, confidentiality, supervision, fees, communication.

Nov 2025

Virginia LEO 1901 LEO 1901 adopted by the Supreme Court of Virginia. Reasonable fees and generative AI.

2026

California COPRAC guidance COPRAC replaces the 2023 version and reaches agentic AI.

Jun 2026

22 NYCRR Part 161 Part 161 takes effect across the New York Unified Court System.

Aug 2026

NYC Bar Opinion 2026-2 NYC Bar, third in a sequence rather than a single pronouncement.

A firm policy citing only ABA 512 is citing the 2024 layer of a stack that has kept moving. Make client AI restrictions enforceable rather than memorable: outside counsel guidelines belong in matter metadata, not in a PDF a lawyer is expected to recall.

On billing, FO 512 is blunt. Lawyers billing hourly "must bill for their actual time," quoting Formal Opinion 93-379, and a lawyer who drafts a pleading in 15 minutes with a tool bills 15 minutes FO 93-379. The harder question is design: if a six-hour task becomes ninety minutes, someone captures that value, and firms that do not decide who will find it was decided by default, through realization. Clio's 2025 research already shows a majority of firms using flat fees in some form Clio 2025. The question is how to price work whose cost structure just changed.

Apprenticeship Debt

Junior lawyers build judgment through work that is, in the narrow sense, inefficient: reading fifty cases to find the six that matter, noticing on the thirty-first contract that the indemnity language has been drifting. Claude removes that repetition. That is the point of buying it, and some of what it removes is how lawyers develop the pattern recognition and skepticism that let a reviewer sense that a fluent memo is wrong.

Call it apprenticeship debt: real, deferred, and invisible on any current-year financial statement.

The fix is not pointless manual labor as a character exercise. It is deliberate redesign: associates read the primary authority behind AI-assisted research, draft before comparing against the model, and are evaluated on whether they can detect model errors. Teach verification as a legal skill. The goal is not preserving inefficiency. It is preserving judgment.

The Governance Stack

Claude does not owe the client a duty, sign the pleading, or carry malpractice responsibility. The responsible lawyer owns the work they adopt. These fifteen layers are what make that ownership real rather than nominal.

01 Vendor and architecture Someone has read the actual terms, retention documentation, and DPA, and can state what applies to your plan
02 Approved surfaces and features A register naming each approved route, model family, and connector, with an owner and a re-review date
03 Identity and access SSO, provisioning and deprovisioning tied to HR, roles scoped to groups, admin privileges deliberately limited
04 Data and matter classification Categories defined and mapped to permissions, so "what may I put in" has an answer
05 Use-case permissions The matrix above, adopted and communicated by practice group
06 Task-specific review standards Each task type has a defined validation, not "an attorney reviewed it"
07 Research and citation controls Authoritative retrieval mandated, verification defined, closed-loop confirmation prohibited
08 Provenance and lifecycle status AI-DRAFT, VALIDATED, and LAWYER-FINAL applied, with metadata retained
09 Monitoring and audit access Who may see content, who authorizes it, whether access is logged, where exports live
10 Testing and regression A benchmark suite with an owner and defined triggers
11 Incident response AI incidents defined beyond breach, with an escalation path and preservation duties
12 Training and development Verification taught as a skill, apprenticeship redesigned rather than eliminated
13 Client and OCG restrictions Client AI limits enforced in matter metadata, not memory
14 Billing and economics A decision about who captures efficiency, and billing practices consistent with it
15 Periodic re-review A calendar date and a named owner, plus event triggers

Define an AI incident broadly: client data through an unapproved surface, a fabricated citation in a filing, a connector sending data somewhere unexpected, a vendor policy change that invalidates an approved workflow. Decide the escalation path before you need it.

The Question That Decides Whether You Are Done

How will the firm demonstrate that Claude improved the quality of legal work, rather than only the volume of work produced? Adoption metrics measure activity. The honest measures are correction rates, citation defects caught in review, omissions discovered, escalation rates, and rework. A firm that cannot distinguish more output from better work has not finished building the governance system, whatever the policy document says.

The unit of governance is not the model. It is the workflow: the task, the information, the configuration, the required validation, the person accountable, and the evidence the work was checked. Firms that govern at that level can adopt aggressively. Firms that govern at the level of "we approved Claude" are relying on individual professionalism to compensate for an absent system, which works right up until it does not.

What This Connects To

This article covers the workflow layer. The rest of the cluster covers the pieces it assumes. ai-tool-due-diligence covers the vendor review that produces the answers in the data-governance section. legal-ai-reliability-evaluation covers how to test reliability before a tool is approved at all. privilege-confidentiality-ai covers the confidentiality classification the use matrix depends on. firm-ai-policy covers the policy structure that carries the approved surfaces register. ai-use-by-role covers role-based permissions. how-to-run-a-legal-tech-pilot covers the pilot structure the benchmark suite fits inside.

This article is operational guidance and does not constitute legal advice or a formal ethics opinion. Professional responsibility obligations vary by jurisdiction. ABA Model Rules and formal opinions are interpretive guidance and do not bind a jurisdiction unless adopted there; state rules, state ethics opinions, court rules, and client outside counsel guidelines may impose different or additional requirements. Product details reflect vendor documentation as of August 2026 and change. Songbird Strategies is a legal technology consulting firm, not a law firm. Firms should consult qualified ethics counsel for legal conclusions and policy decisions. See Sources & Notes for the authority and vendor documentation cited.

Can Your Firm State Which Workflows Are Approved, and What Each One Requires?

If your firm has approved Claude but cannot yet say what validation each workflow requires, what client data may enter the system, and who owns re-evaluation when the product changes, the adoption decision is ahead of the governance system. That is a common place to be and a fixable one. We help firms inventory current AI use, define approved surfaces, and build the review standards that make the work defensible.

Book a Free Strategy Call

30 minutes. No sales pitch.

See the Legal AI Matrix →