The first purchasing decision is not which agent has the longest feature list. It is which legal task can be specified, tested, reviewed, and stopped. A good pilot has a narrow source set, a known acceptable output, a lawyer who owns the result, and a realistic baseline for time and quality. From there, a firm can decide whether a vendor product, a tailored integration, or a conventional workflow best fits the work.
This is a proposed selection and pilot method, not a certification, product ranking, or statement of law. Professional obligations, data terms, court requirements, client instructions, and procurement rules vary. Check them for the exact deployment and matter before using real client information. No vendor performance or cost figures are asserted here.
Choose a workflow where success can be checked
Write one sentence: “Given these approved sources, prepare this internal work product for this named reviewer, without changing or sending anything.” Candidate tasks include identifying clauses across a controlled contract set, preparing a source-linked chronology, or flagging missing documents before a lawyer's review. Avoid starting with court filing, client advice, autonomous external communication, or mixed-client repositories. Those may be important later, but they combine source uncertainty, legal judgment, and side effects before the firm has measured the basic loop.
Swipe sideways to see all columns →
| Question | Stronger pilot candidate | Weaker pilot candidate |
|---|---|---|
| Evidence | Finite, permission-cleared packet with version IDs and checkable passages. | Open-ended drive or web search with unclear source priority. |
| Output | Internal issue list or draft whose claims a lawyer can verify. | Unreviewed advice, filed paper, or external message. |
| Success | Known issues, omissions, false claims, and reviewer time can be counted. | A vague goal such as “handle the matter.” |
| Recovery | Work can be stopped and discarded without external effect. | A wrong action could create an immediate deadline or disclosure. |
This is an editorial triage rubric, not a legal risk score. A narrow task can still contain sensitive material. Before procurement, the firm should map intended use, affected people, sources, controls, and likely harms. NIST's AI Risk Management Framework covers governance, mapping, measurement, and management, including third-party technology and contingency processes. It is voluntary and context dependent, rather than a certification a vendor can simply claim to have passed.
Build, buy, or use a fixed workflow
Swipe sideways to see all columns →
| Option | What the firm controls | Main obligation to verify |
|---|---|---|
| Fixed workflow | Known steps and rule-based gates, possibly with a model for extraction or drafting. | Whether the workflow covers real document variation and keeps human review effective. |
| Vendor agent | Configuration and contract, subject to the vendor's feature and logging limits. | Actual data flows, permissions, terms, version changes, evidence export, and incident support. |
| Tailored agent | Tool definitions, integration and policy code, evaluation, and monitoring design. | Whether the firm can maintain security, testing, uptime, and model/tool changes over time. |
Buying may reduce engineering work but does not transfer a lawyer's professional responsibilities. Building may allow narrower connectors and clearer traces but creates ongoing software and security obligations. A fixed workflow may outperform an autonomous agent on a stable, repetitive task. Use the same representative cases and reviewed-outcome measures for all credible options; a polished demonstration with a vendor's sample data is not a comparison.
Ask vendors for a data-flow answer, not a slogan
For the specific product, plan, and configuration, ask for a diagram of every place matter data can go: uploaded files, prompts, retrieved snippets, model providers, logs, analytics, backups, support access, subprocessors, and persisted notes. Ask who owns those records, their location and retention, whether they are used for model improvement, how deletion works, and what can be exported. Do not infer these answers from a brand name or a generic “enterprise-grade” phrase. Confirm them in current terms and configuration, with security and legal reviewers where appropriate.
Swipe sideways to see all columns →
| Due-diligence area | Question to obtain in writing or demonstrate |
|---|---|
| Access | Can reads and writes be scoped by identity, client, matter, repository, field, and recipient? Is access rechecked on every tool call? |
| Action control | Can sending, filing, deletion, record edits, and purchases be disabled or require an exact-action approval? What happens on timeout? |
| Evidence | Can reviewers open source passages and exported traces with document versions, tool arguments, approvals, and actual outcomes? |
| Data terms | What are retention, training/data-use, deletion, subprocessors, support access, geographic, and incident-notification terms for this plan? |
| Change control | How are model, prompt, tool, index, and policy changes announced, tested, rolled back, or pinned? |
| Exit | Can the firm export data and audit records, revoke credentials, delete copies under the contract, and continue work if the service fails? |
Obtain a live demonstration of denied access and rejected actions, not only successful output. Ask the vendor to show what a reviewer sees before a side-effecting call and what the log records afterward. OpenAI's guardrail and approval documentation illustrates why approval belongs at the tool boundary; the specific vendor must prove its own implementation. The permissions guide can serve as a discussion map.
Count the cost of an accepted result
Compare total cost per lawyer-accepted work product, not price per model response. Record license and usage charges, integration and security work, lawyer review and correction time, failed or retried runs, monitoring, training, and incident handling. A practical expression is: total pilot cost ÷ accepted outputs, reported alongside quality and severe-event counts. This is a comparison tool, not a universal accounting rule or a promise of savings. Preserve a baseline from the existing workflow, including its own review and rework time.
Latency matters too. A workflow that saves drafting time but increases source-checking or exception handling may not save a lawyer time overall. A cheaper option that cannot provide source or action traces may be inappropriate for the task. The agent evaluation scorecard records support, misses, unauthorized actions, reviewer time, and total reviewed cost across repeated trials.
Run the pilot in stages
Swipe sideways to see all columns →
| Stage | Scope | Decision before moving on |
|---|---|---|
| 1. Define | Task, packet, accepted output, owner, baseline, test cases, severe-failure gates. | Can qualified reviewers agree on what correct and incorrect behavior looks like? |
| 2. Test | Fictional or cleared data; read-only tools; repeated normal and adversarial cases. | Are claims source-supported and boundaries enforced in traces, including failure cases? |
| 3. Shadow | Run beside the existing process; a lawyer checks output without agent-side external action. | Does the reviewed work meet quality and effort targets on realistic matters? |
| 4. Limited use | Named users and matter types; internal drafts; active monitoring and rollback. | Does live performance stay inside the tested scope and stop rules? |
| 5. Expand or stop | Change one risk dimension at a time—document type, users, volume, or authority. | Is there evidence for this specific expansion, with new cases and an accountable owner? |
The stages are a suggested sequence, not calendar commitments. Write the release gate before tests start. A proposed gate for a read-only internal-draft pilot is: no observed cross-matter access or external action, required critical issues found or explicitly escalated, material claims checked against sources, and outputs approved by a lawyer before use. Zero incidents in a finite pilot does not prove future safety. A failed gate should narrow or stop the rollout until the cause is understood and retested.
Plan changes, incidents, and exit before launch
Name the people who own product updates, legal quality, data access, incident response, and vendor relations. Keep a versioned evaluation set and rerun it after changes to the model, prompts, retrieval index, connected tools, permissions, or document mix. Establish a way for lawyers to report a wrong source or strange approval request with the run ID. The failure and incident guide describes a response path; it should be rehearsed before the service has authority to act.
An exit plan covers revoked tokens, exported documents and traces, deletion requests under the contract, fallback work queues, and retention of records the firm still needs. Test that a matter can continue when the agent is unavailable. NIST's AI RMF calls for third-party contingency processes; that is particularly relevant if a workflow becomes part of daily practice. The outcome of a successful pilot is a bounded, monitored use case with a known owner—not a blanket conclusion that the product is suitable for all legal work.