How was the appellate-briefing test run?
We supplied the complete text of Geller v. Uber Technologies, Inc., 2026 IL 132066, to three fresh Codex workspace agents. Each received the same prompt, with instructions to write a 500–700 word brief under six headings, cite numbered opinion paragraphs, preserve legal capacities and distinct claims, and identify matters the court did not decide. No answer key was included.
Before generation, we recorded the source fingerprints, the prompt, the three-run plan, and eight source-comparison criteria. Every planned response is retained, including citation weaknesses. The input was layout-extracted text, not direct visual PDF interpretation. The method tests this supplied-source drafting task; it does not test independent legal research or current-law verification.
The runtime-reported configuration was gpt-6.1-sol with reasoning effort ultra, inherited by the delegated agents. The interface did not expose an immutable model version, temperature, seed, token billing, or isolated inference time. Agent instructions prohibited browsing, answer-key access, and consultation with other agents. Platform instructions and available tools remained part of the environment.
What did the opinion actually decide?
The Illinois Supreme Court concluded that Sheridan’s individual Uber agreement did not establish consent to delegate arbitrability of these wrongful-death claims, which arose from Mark’s use. The court separately concluded that those claims were not subject to arbitration. It reversed the appellate judgment, affirmed the circuit judgment, and remanded. Wrongful-death proceedings may resume; the voluntarily dismissed survival counts remain dismissed. The court did not reach procedural or substantive unconscionability. See ¶¶63, 70, 72–79. Read the official opinion ↗
What needs correction in the AI briefs?
The source comparison found that all three briefs preserve the central outcome, procedural sequence, separate capacities, and limits of the holding. Their shared description of Mark and Sheridan accepting separate agreements cites ¶¶8–9. Paragraph 8 recounts Uber’s assertion; paragraph 9 reproduces Sheridan’s agreement. Paragraph 10 expressly confirms Mark’s execution and says his agreement is not reproduced. Adding ¶10 makes that factual support more precise. The filing date comes from the opinion’s caption; identify that source when citing the date.
The observation is a citation-support issue, not proof that the substantive holding is wrong. A useful correction preserves the accurate distinction while tightening the source pointer. Do not substitute a blanket claim that Uber arbitration provisions are invalid, that every wrongful-death case has the same outcome, or that the court found this agreement unconscionable.
How should a lawyer review an AI case brief?
Start with the full opinion, not the generated summary. Verify the caption, parties and legal capacities, material facts, procedural history, questions presented, holding, reasoning, disposition, and issues expressly left undecided. Open each cited paragraph and ask whether it supports the precise proposition. Separate the court’s account of an allegation or argument from its own finding. Record the correction and the reason before accepting the brief.
The worksheet provides a place for a named reviewer to record each check, the source support, needed corrections, and an acceptance decision. Record actual review time and total cost if measuring productivity. Generation speed alone does not measure time saved when reading, checking, correcting, and professional responsibility remain part of the task.
What does this field test leave unmeasured?
This study covers one public opinion, one relatively detailed source-only prompt, one configured model alias, and three runs. It has no untreated baseline, blind human grading, or measured lawyer review time. It does not establish an accuracy percentage, a product ranking, suitability for confidential client files, a general reliability claim, or a conclusion about legal research performance. Neither later treatment nor current legal validity was checked by these runs.
The response comparison was prepared with Codex assistance and is pending Jonathan Nessler’s legal review. No legal-reviewer attribution, completed review date, review-time figure, or cost figure is asserted. The recorded test date identifies when the runs occurred in America/Chicago; UTC start and end timestamps are included for provenance.
How does this differ from the Geller teaching walkthrough?
The existing walkthrough provides authored prompts and a source-checked teaching brief. This field test records fresh, unedited model responses and compares them against criteria fixed before generation. Keep those two forms of evidence distinct. Read the Geller teaching walkthrough →
