The first purchasing decision is not which agent has the longest feature list. It is which legal task can be specified, tested, reviewed, and stopped. A good pilot has a narrow source set, a known acceptable output, a lawyer who owns the result, and a realistic baseline for time and quality. From there, a firm can decide whether a vendor product, a tailored integration, or a conventional workflow best fits the work.

This is a proposed selection and pilot method, not a certification, product ranking, or statement of law. Professional obligations, data terms, court requirements, client instructions, and procurement rules vary. Check them for the exact deployment and matter before using real client information. No vendor performance or cost figures are asserted here.

Illustrative legal agent adoption path: select a bounded task, inspect data use and tool permissions, pilot on representative matters with review time, then expand, revise, or stop at a release gate
A proposed adoption path, not a universal timetable or a claim that any vendor meets these gates.

Choose a workflow where success can be checked

Write one sentence: “Given these approved sources, prepare this internal work product for this named reviewer, without changing or sending anything.” Candidate tasks include identifying clauses across a controlled contract set, preparing a source-linked chronology, or flagging missing documents before a lawyer's review. Avoid starting with court filing, client advice, autonomous external communication, or mixed-client repositories. Those may be important later, but they combine source uncertainty, legal judgment, and side effects before the firm has measured the basic loop.

Swipe sideways to see all columns →

QuestionStronger pilot candidateWeaker pilot candidate
EvidenceFinite, permission-cleared packet with version IDs and checkable passages.Open-ended drive or web search with unclear source priority.
OutputInternal issue list or draft whose claims a lawyer can verify.Unreviewed advice, filed paper, or external message.
SuccessKnown issues, omissions, false claims, and reviewer time can be counted.A vague goal such as “handle the matter.”
RecoveryWork can be stopped and discarded without external effect.A wrong action could create an immediate deadline or disclosure.

This is an editorial triage rubric, not a legal risk score. A narrow task can still contain sensitive material. Before procurement, the firm should map intended use, affected people, sources, controls, and likely harms. NIST's AI Risk Management Framework covers governance, mapping, measurement, and management, including third-party technology and contingency processes. It is voluntary and context dependent, rather than a certification a vendor can simply claim to have passed.

Build, buy, or use a fixed workflow

Swipe sideways to see all columns →

OptionWhat the firm controlsMain obligation to verify
Fixed workflowKnown steps and rule-based gates, possibly with a model for extraction or drafting.Whether the workflow covers real document variation and keeps human review effective.
Vendor agentConfiguration and contract, subject to the vendor's feature and logging limits.Actual data flows, permissions, terms, version changes, evidence export, and incident support.
Tailored agentTool definitions, integration and policy code, evaluation, and monitoring design.Whether the firm can maintain security, testing, uptime, and model/tool changes over time.

Buying may reduce engineering work but does not transfer a lawyer's professional responsibilities. Building may allow narrower connectors and clearer traces but creates ongoing software and security obligations. A fixed workflow may outperform an autonomous agent on a stable, repetitive task. Use the same representative cases and reviewed-outcome measures for all credible options; a polished demonstration with a vendor's sample data is not a comparison.

Ask vendors for a data-flow answer, not a slogan

For the specific product, plan, and configuration, ask for a diagram of every place matter data can go: uploaded files, prompts, retrieved snippets, model providers, logs, analytics, backups, support access, subprocessors, and persisted notes. Ask who owns those records, their location and retention, whether they are used for model improvement, how deletion works, and what can be exported. Do not infer these answers from a brand name or a generic “enterprise-grade” phrase. Confirm them in current terms and configuration, with security and legal reviewers where appropriate.

Swipe sideways to see all columns →

Due-diligence areaQuestion to obtain in writing or demonstrate
AccessCan reads and writes be scoped by identity, client, matter, repository, field, and recipient? Is access rechecked on every tool call?
Action controlCan sending, filing, deletion, record edits, and purchases be disabled or require an exact-action approval? What happens on timeout?
EvidenceCan reviewers open source passages and exported traces with document versions, tool arguments, approvals, and actual outcomes?
Data termsWhat are retention, training/data-use, deletion, subprocessors, support access, geographic, and incident-notification terms for this plan?
Change controlHow are model, prompt, tool, index, and policy changes announced, tested, rolled back, or pinned?
ExitCan the firm export data and audit records, revoke credentials, delete copies under the contract, and continue work if the service fails?

Obtain a live demonstration of denied access and rejected actions, not only successful output. Ask the vendor to show what a reviewer sees before a side-effecting call and what the log records afterward. OpenAI's guardrail and approval documentation illustrates why approval belongs at the tool boundary; the specific vendor must prove its own implementation. The permissions guide can serve as a discussion map.

Count the cost of an accepted result

Compare total cost per lawyer-accepted work product, not price per model response. Record license and usage charges, integration and security work, lawyer review and correction time, failed or retried runs, monitoring, training, and incident handling. A practical expression is: total pilot cost ÷ accepted outputs, reported alongside quality and severe-event counts. This is a comparison tool, not a universal accounting rule or a promise of savings. Preserve a baseline from the existing workflow, including its own review and rework time.

Latency matters too. A workflow that saves drafting time but increases source-checking or exception handling may not save a lawyer time overall. A cheaper option that cannot provide source or action traces may be inappropriate for the task. The agent evaluation scorecard records support, misses, unauthorized actions, reviewer time, and total reviewed cost across repeated trials.

Run the pilot in stages

Swipe sideways to see all columns →

StageScopeDecision before moving on
1. DefineTask, packet, accepted output, owner, baseline, test cases, severe-failure gates.Can qualified reviewers agree on what correct and incorrect behavior looks like?
2. TestFictional or cleared data; read-only tools; repeated normal and adversarial cases.Are claims source-supported and boundaries enforced in traces, including failure cases?
3. ShadowRun beside the existing process; a lawyer checks output without agent-side external action.Does the reviewed work meet quality and effort targets on realistic matters?
4. Limited useNamed users and matter types; internal drafts; active monitoring and rollback.Does live performance stay inside the tested scope and stop rules?
5. Expand or stopChange one risk dimension at a time—document type, users, volume, or authority.Is there evidence for this specific expansion, with new cases and an accountable owner?

The stages are a suggested sequence, not calendar commitments. Write the release gate before tests start. A proposed gate for a read-only internal-draft pilot is: no observed cross-matter access or external action, required critical issues found or explicitly escalated, material claims checked against sources, and outputs approved by a lawyer before use. Zero incidents in a finite pilot does not prove future safety. A failed gate should narrow or stop the rollout until the cause is understood and retested.

Plan changes, incidents, and exit before launch

Name the people who own product updates, legal quality, data access, incident response, and vendor relations. Keep a versioned evaluation set and rerun it after changes to the model, prompts, retrieval index, connected tools, permissions, or document mix. Establish a way for lawyers to report a wrong source or strange approval request with the run ID. The failure and incident guide describes a response path; it should be rehearsed before the service has authority to act.

An exit plan covers revoked tokens, exported documents and traces, deletion requests under the contract, fallback work queues, and retention of records the firm still needs. Test that a matter can continue when the agent is unavailable. NIST's AI RMF calls for third-party contingency processes; that is particularly relevant if a workflow becomes part of daily practice. The outcome of a successful pilot is a bounded, monitored use case with a known owner—not a blanket conclusion that the product is suitable for all legal work.