An AI model capability is an ability to perform a task under specified conditions, including the inputs and tools available and the time allowed. What matters in practice is whether the complete system can produce an accurate, reviewable result for your task. Stronger models can increasingly:

  • Reason across evidence: Compare documents, trace claims to sources, and surface conflicts.
  • Interpret varied inputs: Work with text, images, charts, and long files together.
  • Use tools: Search, calculate, write code, and prepare artifacts a person can inspect.
  • Sustain difficult work: Plan, test, revise, and continue through multiple steps.
  • Search for leads: Explore large data sets and suggest patterns experts can test.

The release of a new AI model usually comes with a familiar set of numbers: benchmark scores, context-window size, speed, and price. Those figures are worth reading. They do not, by themselves, answer the question a lawyer, business owner, researcher, or manager needs answered: What work can this system now do well enough to be useful, and what must I still check?

That question has become more urgent. A capable model can move beyond producing a polished paragraph. Given the right tools and instructions, it can work through a collection of documents, compare competing explanations, write and test code, use software, and continue a difficult task through multiple rounds of work. Some systems can coordinate many model sessions at once. Lower prices make it practical to apply those abilities repeatedly, across a whole workflow.

Understanding these capabilities changes what we ask AI to do. It also changes how we supervise it. A reliable summary of one document is different from a research plan, and a research plan is different from an agent authorized to change files or send information outside an organization. The same model may be suitable for the first task and require strict limits on the third.

Illustration of documents, a photograph, charts, and data flowing toward a report
AI-generated illustration. A stronger system can bring different materials into one reviewable result. The result still needs evidence and human judgment.

Capability is more than a model name

It is tempting to describe one model as “smarter” than another. For actual work, that label is too thin. A useful assessment asks whether a system can complete a particular task with the information, tools, time, and safeguards available to it.

The model interprets inputs and generates responses. The system around it may supply documents, search, code execution, a browser, memory, and rules for when to pause. A long-running AI agent is the combination. Giving a model access to a filing system does not prove it can find the right filing. Giving it a browser does not prove it will recognize an outdated page. A large context window does not prove it attended to every paragraph placed inside it.

The meaningful unit is the accepted result: work that satisfies the task, survives review, and is worth the time and money spent producing it. That is why model capabilities matter more than a single leaderboard position. They tell us where to use a model, how to structure the work, and where human attention creates the most value.

What stronger models make possible

Reason across documents and conflicting evidence

A model can now do more than summarize documents one at a time. It can compare accounts, trace a claim back to its source, find apparent conflicts, propose explanations, and organize the result for review. This is especially useful when the information is scattered across many files.

Consider a litigation team examining an opinion, several deposition excerpts, and a chronology. A capable system might prepare an issue map that separates what the court held from what a witness said, identifies dates that appear to conflict, and points the lawyer to each supporting passage. The value is not that it “knows the case.” The value is that it can create a structured first pass that makes the lawyer’s own reading more focused.

The same pattern applies outside law. A business team could compare supplier proposals against an approved requirements list. A researcher could ask for the assumptions behind several competing papers. In each case, the model’s strongest contribution is often to reveal the shape of the problem and make the evidence easier to inspect.

This ability still has a condition: the system must show its work in a checkable form. A citation is useful only if the cited passage exists and supports the point. Fluency can make an unsupported conclusion sound settled.

Analyze text, images, charts, and long documents

Recent models can take images as well as text. GPT-6 Sol and GPT-6 Luna, for example, accept both. Their published specifications also allow very large inputs. That opens workflows involving diagrams, screenshots, charts, scanned pages, and lengthy collections of text.

Picture a contract review in which key terms appear in a table, an exhibit contains an annotated diagram, and a handwritten note changes the meaning of a date. A system that can examine the pages together may catch a relationship that disappears when the file is reduced to plain text. A model that can review a large set of materials at once may also preserve context that would otherwise be lost between separate conversations.

Capacity is not the same as comprehension. Scans can be unreadable, tables can be misinterpreted, and a detail near the middle of a long input can still be missed. For important work, ask the model to identify which files and pages it used. Check those pages, especially where the answer depends on a small qualification.

Use tools to produce reviewable work

A model connected to appropriate tools can search for information, run calculations, write code, edit a document, or operate a computer interface. The output can be a draft brief, a tested software change, a spreadsheet, or a collection of research leads. That is a different level of usefulness from a chat response that merely describes what someone else should do.

For example, a software agent can inspect a repository, make a limited change, run the relevant tests, and present the diff. A business workflow might extract data from incoming forms, compare it with an approved policy, and prepare exceptions for a person to decide. These are illustrative workflows; their reliability must be measured on the actual files and tools in use.

Tool access also creates real consequences. If an agent can edit a database, publish a page, or email a client, a mistaken interpretation can become an external action. Define what the system may read, what it may change, which actions require approval, and how its work will be reviewed. This is part of a responsible AI workflow. The OpenAI model documentation lists tool support for Sol, but the surrounding application determines which tools are available and what the agent is allowed to do.

Sustain multistep work on difficult problems

Some work takes a sequence of attempts. The system must make a plan, test an approach, recognize a dead end, preserve useful results, and continue. Stronger models, combined with software that manages the process, are becoming capable of carrying out more of that sequence.

A September 2026 account hosted by Anthropic offers a striking example. Physicists at Anthropic used Claude Fable 5.1 in Claude Science to calculate a nine-loop scattering amplitude in a simplified theory used to test particle-physics methods. The model pursued two established approaches, using code and sustained computation. Physicist Lance Dixon checked the result. Another research group had independently made substantial progress on the same problem with some GPT-6 assistance.

The accomplishment is substantial because the calculation was difficult and fragile, and the model kept the work moving through many steps. The account does not claim that Claude discovered a new physical principle. It applied methods developed by researchers, within a tool-supported system, to a problem those researchers knew how to validate. This is a good picture of what greater capability can mean: expert methods carried through to a checkable result with less continuous human labor.

Search large data sets for research leads

In another September 2026 report, Anthropic described Claude agents searching DNA data for unusual reverse transcriptase systems. Roughly 950 agents worked for 21 hours. One flagged a known enzyme beside an accessory gene and a repeating DNA array. Anthropic’s scientists investigated the combination and named the previously uncharacterized system array-associated reverse transcriptases, or ART.

The distinction matters. The enzyme itself had been identified before. The model’s contribution was to notice a potentially meaningful arrangement in a large search space and produce a lead for scientists to examine. Humans performed the laboratory experiments. The biological function of ART remains unknown, and any use as a gene-editing tool is unproven. The report is an early research result, not a finished biotechnology application.

For a professional outside biology, the lesson is still relevant. More capable AI can generate and filter hypotheses across data too large for one person to survey closely. The expert’s role shifts toward framing the search, deciding which leads deserve attention, and designing the tests that separate a promising pattern from a real finding.

The economics are part of the capability

A task may be technically possible and still too expensive to repeat. That is why the September 22 model releases matter beyond their benchmark charts. OpenAI introduced GPT-6 Sol and Luna as different balances of capability and cost. Sol is aimed at complex coding and agent work; Luna is aimed at focused, high-volume tasks. As checked on September 25, 2026, their listed standard API prices for shorter inputs are $2 and $0.10 per million input tokens, respectively, with higher prices for generated output. Pricing changes with input length and processing tier. These are developer usage prices, not chat subscription prices.

Anthropic introduced Claude Opus 5.5 the same day, reporting stronger performance than Opus 5 and an estimated 40 percent lower cost on typical workloads. Its published input and output token prices are each 20 percent lower than Opus 5. Anthropic attributes the larger typical-workload saving partly to lower cache-read prices and fewer tokens used per task.

These are provider claims and price schedules, not a guarantee that one model will be cheapest for your work. A more capable model can cost more per request while saving time by making fewer attempts or producing a draft that needs fewer corrections. A cheaper model can be ideal for a well-defined, repetitive task and poor value when every error creates expensive cleanup.

Measure cost per accepted result. Include model charges, tool charges, retries, reviewer time, and the cost of misses. The most useful model is the one that clears the quality bar for the specific job at an acceptable total cost.

Better scores do not settle the question

Benchmarks are useful signals. They are also snapshots of a particular task set, tool setup, reasoning setting, and scoring method. A model can improve sharply on one evaluation and barely change on a task that matters to you.

OpenAI reports that Sol made about half as many factual errors as GPT-5.6 Sol on an internal evaluation built from conversations in which users had flagged errors. OpenAI explicitly says those conversations are not representative of typical use. Anthropic likewise cautions that small benchmark margins at this capability level can be a poor guide to real-world differences. Both qualifications belong beside the impressive scores.

The right question is whether the result transfers to your workflow. Can the model handle the documents your organization actually receives? Does it recover after a tool fails? Does it admit when a source is missing? Does it preserve the distinction between evidence and inference? Does a reviewer understand what happened well enough to approve the work?

Scientific demonstrations deserve the same discipline. The nine-loop result shows a successful, specialist-checked calculation using established methods. The ART report shows a promising lead from a large search, with function still unknown. Each expands the picture of what AI can contribute. Neither establishes a general ability to solve every hard science problem.

A practical way to test a model for your work

Start with one bounded task whose result you can judge. “Help with legal research” is too broad. “Identify the controlling holding and every limitation on relief in these ten opinions, with paragraph citations” can be evaluated.

  1. Define the deliverable. Specify the source materials, output format, required citations, and what counts as a correct answer. Decide which mistakes would be inconvenient and which would be unacceptable.
  2. Use representative examples. Begin with public, fictional, or properly approved material. Include routine items, messy documents, missing information, and the edge cases your team has learned to watch for. A polished demo document tells you little about the difficult files in your own practice.
  3. Review the evidence, not the style. Check citations, calculations, omitted qualifications, and any external facts against primary sources. Record whether the model noticed uncertainty or filled gaps with plausible text.
  4. Test the complete workflow. If the proposed use involves search, file access, code execution, or other software, test the whole system. The model’s score without those tools may not predict its behavior with them.
  5. Count the whole cost. Track retries, review time, delays, tool charges, and corrections. Compare the cost of an accepted result with your current process.
  6. Set the action boundary. A system may be allowed to suggest a change, prepare it for approval, or carry it out within limits. Choose the boundary based on the consequences of an error, then make it visible to everyone using the workflow.
  7. Recheck after change. Model versions, prompts, data sources, and connected tools change. Revisit the examples that originally justified the workflow whenever one of those parts changes materially.

This process often reveals a useful division of labor. A fast, inexpensive model may sort clear cases. A stronger model may handle ambiguous ones. A person reviews the matters where context, accountability, or professional judgment is decisive. The division should come from observed performance, not a product name.

The question worth asking next

More powerful models expand the range of work that can be started, advanced, and sometimes finished with AI. They can reason across evidence, work with images and long files, use tools, keep a difficult project moving, and search enormous spaces for leads. Falling costs bring those abilities within reach of more ordinary work.

Knowing the capabilities is how we use that progress responsibly. It lets us recognize an opportunity that would have been impractical a year ago, choose a system suited to the task, and decide what proof we need before accepting its result. With evidence, permissions, and review in place, capable tools can make human judgment better informed and more effective.