All articles

Why Law Firms Are Moving To On-Prem AI

Cloud AI is structurally incompatible with attorney-client privilege. Here's why law firms are moving to on-prem private AI and how the engagement works.

A managing partner at a mid-size litigation firm in Carmel Valley told us this last month: three of his associates had been quietly pasting deposition transcripts into ChatGPT to generate summaries. Not because they were reckless. Because it was fast, it was good, and nobody had told them not to. When he found out, he spent a weekend trying to figure out whether he had to disclose it to clients, to opposing counsel, or to his malpractice carrier. The answer, depending on which state bar opinion you read, was somewhere between "probably" and "definitely."

That firm is not unusual. It's the median. Every law firm in North County San Diego right now has attorneys using consumer AI tools on privileged material, and most managing partners either don't know or are actively trying not to know. The productivity gains are real. The exposure is worse.

The structural problem with cloud AI in legal practice

Attorney-client privilege depends on a specific thing: the communication has to be kept confidential. Not "kept confidential except for the vendor's terms of service." Not "kept confidential unless the model provider decides to use it for training." Confidential.

When an associate pastes a client email into a hosted large language model, that content leaves the firm's control. It transits a third-party network, sits on a third-party server, gets processed by a third-party inference engine, and — depending on the plan tier and the vendor's current policy — may be retained, logged, reviewed by human trainers, or used to improve the model. Some vendors offer enterprise agreements that promise no training use and short retention windows. Those agreements are contractual, not architectural. The data still leaves the building.

Work-product doctrine has the same problem, in a slightly different shape. Work product protects an attorney's mental impressions, legal theories, and case strategy from discovery. The moment those impressions get typed into a prompt on a hosted service, a reasonable opposing counsel is going to ask whether privilege was waived by voluntary disclosure to a third party. The case law is thin because the practice is new. That's not comforting. That's the risk.

And then there's conflicts. A firm representing Party A in a dispute against Party B cannot have its systems ingesting Party B's confidential material, even accidentally. If both parties' work has been fed into the same hosted model — even under an enterprise contract with segmentation promises — the firm is now relying on a vendor's technical controls to enforce an ethical wall. State bar opinions have not been kind to that arrangement.

What the state bars are actually saying

The pattern across jurisdictions is consistent even when the specifics vary. Florida Bar Ethics Opinion 24-1, California's State Bar practical guidance on generative AI, and the New York State Bar Association's April 2024 report all converge on the same set of duties: competence, confidentiality, supervision, and communication with the client.

The through-line: attorneys must understand how the AI tool handles data before using it on client matters, must not input confidential information into systems that could disclose or retain it improperly, and in many cases must inform clients when AI is being used on their work. Several opinions go further and suggest that using a public AI tool on privileged material could itself be a competence violation.

None of these opinions ban AI. All of them require the attorney to know where the data goes. That's the operative question. Most attorneys using ChatGPT, Claude, or Gemini on client work today cannot answer it with any specificity — and "the vendor promised" is not a defensible answer in a bar complaint.

The workflows lawyers actually want

The reason this problem is urgent, and not something firms can just wait out, is that the workflows AI is genuinely good at are exactly the workflows that generate the most billable hour compression and the most associate frustration.

Document review. Reviewing 40,000 pages of production for a mid-sized commercial dispute is the kind of work that used to take an associate three weeks. A well-tuned model can surface responsive documents, flag privilege candidates, and cluster themes in a fraction of the time. This is the single highest-value use case in litigation practice, and it involves the most sensitive material in the case.

Deposition preparation. Ingesting prior testimony, correlating it against document exhibits, and generating a first-pass outline of contradictions and areas to probe. Extraordinarily useful. Extraordinarily privileged.

Contract analysis. Comparing a redlined draft against a firm's standard playbook, flagging deviations, summarizing risk. Transactional attorneys have wanted this tool for a decade. It exists now. The contracts contain terms the client considers deeply confidential.

Discovery summarization. Turning ten thousand pages of ESI into a working chronology. Pure gold for a case team. Pure poison if it leaks.

First-pass drafting. Motions, briefs, correspondence, memoranda. Every first draft an attorney writes is work product from the moment the cursor blinks. Sending it to a hosted API to "help clean it up" is the exact scenario the ethics opinions describe.

The problem is not that lawyers shouldn't use AI for this work. They should. The problem is that the current default — hosted consumer or enterprise cloud models — is structurally wrong for it.

What on-prem private AI actually is

A private AI deployment for a law firm looks like this: an appliance — a purpose-built server with modern GPU hardware — sits in the firm's server room or IT closet. Open-weight models (Llama, Mistral, Qwen, and their fine-tuned legal variants) run on that hardware. Attorneys interact with the system through a familiar chat interface, or through integrations with document management systems like NetDocuments, iManage, or Clio.

Every prompt, every document uploaded, every response generated stays on that appliance. Nothing transits the public internet. Nothing goes to a vendor's cloud. Nothing is logged by anyone outside the firm. The model doesn't learn from the firm's queries in a way that could leak across matters, because there is no shared tenancy. It's the firm's hardware, the firm's model weights, the firm's data, the firm's network.

This is what private AI infrastructure means in a legal context. It's not a policy. It's not a contract with a vendor. It's a physical and architectural separation between the tool and the outside world.

The modern open-weight models are good enough to make this practical. A 70-billion-parameter model running locally on the right hardware handles document review, summarization, deposition prep, and drafting at a quality level that firms describe as indistinguishable from GPT-4-class hosted services for most legal tasks. For specialized work — contract analysis against firm-specific playbooks, for instance — fine-tuning the model on the firm's own document corpus produces results that hosted general-purpose models can't match, because the hosted models have never seen the firm's playbook.

Why this changes the malpractice calculus

The exposure with hosted AI isn't just theoretical bar complaints. It's the concrete question a malpractice carrier will ask after a data incident: what were your controls? "We had a written policy telling associates not to use ChatGPT on client work" is a weak answer, because associates use it anyway. "The AI tool physically cannot send data outside the building" is a strong answer.

The same logic applies to client communications. Increasingly, sophisticated clients — general counsel at public companies, in-house teams at private equity firms, family offices — are asking outside counsel directly: do you use AI on our work, and if so, where does our data go? A firm that can answer "we use AI extensively, and your data never leaves our office" is in a fundamentally different competitive position than a firm that has to explain a hosted vendor's data handling policy.

The engagement letter can now say something specific. AI tooling is used. It runs on firm-controlled infrastructure. No client material is transmitted to third-party AI services. That's a statement a client can rely on and a bar counsel can verify.

The engagement model

For firms in the ten-to-fifty-attorney range — which is most of the firms we work with across San Diego — the practical deployment looks like this:

  1. Site assessment and infrastructure review. We evaluate the firm's existing server room, power, cooling, and network. Most firms need modest upgrades. Some need a proper IT closet built out. This is the same infrastructure work SentriCraft has always done for business network deployments, extended to accommodate GPU hardware.

  2. Appliance sizing and installation. Model choice, hardware sizing, and integration architecture are matched to the firm's caseload, document volume, and existing systems. The appliance is delivered, racked, networked, and configured on-site.

  3. Model deployment and tuning. Open-weight base models are installed and, where valuable, fine-tuned on the firm's own document corpus — brief banks, contract templates, prior work product. The firm owns the resulting model weights entirely.

  4. Attorney interface. A chat interface that looks and feels like ChatGPT, plus integrations with the firm's document management system. Attorneys don't need to learn new workflows. They point their queries at the internal system instead of the public one.

  5. Ongoing management. Model updates, security patching, monitoring, and capacity planning are handled as a managed service. The firm doesn't need to hire an AI engineer. The system is maintained the way the phone system is maintained — quietly, reliably, in the background.

The appliance lives in the firm's office. Attorneys query it like ChatGPT. The data never leaves the building. That's the whole architecture, and it's the only architecture that satisfies the duties a modern state bar opinion is going to hold attorneys to five years from now.

The firms that get this infrastructure in place now will practice differently than the ones that don't. Private AI on privileged work isn't a productivity upgrade. It's the baseline competent practice of law in the next decade.

Ready for a network that just works?

Book a Free Consult