An AI agent can occupy a virtual computer while repeatedly asking a model what to do next. Those activities have different costs. The computer supplies a working environment: a browser, files, memory, and a place to run code. Model inference supplies the computation that interprets information, reasons through a problem, and produces the next response or tool call.

That distinction changes how we should think about demand. The same person can move from a short conversation to a long investigation, then ask several agents to explore different parts at once. AI usage can grow much faster than user counts when each user delegates more work and each task consumes more computation. Whether that becomes higher spending depends on prices, efficiency, and how much useful work gets done.

Conceptual miniature computer workspace connected to a separate cluster of processing tiles that branch into parallel paths
Conceptual illustration of the working environment and model computation behind an agent. The objects do not represent a particular product or measured capacity.

What are the two main costs of an AI agent?

The virtual computer is the environment in which an agent's tools operate. It might be a remote desktop, browser session, container, or isolated virtual machine, often called a sandbox. The model processes inputs and generates outputs. In a common hosted setup, the sandbox calls a model service elsewhere; renting the sandbox does not itself pay for those model calls.

Swipe sideways to see all columns →

Cost layerWhat it suppliesWhat can increase usage
Virtual computerCPU, allocated memory, a browser or code runtime, files, and network accessMore active environments, longer runtime, larger allocations, and more storage or traffic
Model inferenceProcessing context, reasoning, generating answers, and proposing tool callsMore calls, longer inputs, deeper reasoning, more output, and additional agent branches

An agent that uses only APIs may not need a full virtual desktop. A locally hosted model may share infrastructure with its tools. The accounting distinction still matters, even when a vendor combines the charges into one subscription or usage allowance. Our agent architecture guide explains how the model and surrounding application exchange work.

Why a cheap virtual computer does not mean a cheap agent

Consider a sandbox that spends several minutes waiting for model responses. It may use little CPU during those waits while keeping memory allocated. Vercel's Sandbox billing documentation illustrates the distinction: Active CPU excludes waiting for model calls, while provisioned memory is metered by allocated gigabytes and running time. E2B's pricing documentation uses per-second running-sandbox charges with CPU and memory components. Providers do not all meter the same way.

Neither sandbox meter tells you how much the model service consumed. One run might make two brief calls; another might repeatedly process documents, generate alternatives, and check its work. Similar environments can therefore have very different total costs. A computationally heavy tool can also make the environment the larger expense. Measure both layers before assuming which dominates.

One agent task uses a virtual computer and model inference, connected by tool calls and returned results, with different usage meters for each
Two resource meters within one agent task. Waiting for a model can keep a workspace allocated; the provider's billing rules determine the charge.

For budgeting, use total operating cost = environment + inference + other tools and data + operations and human review. Search services, document processing, licensed information, and the time to verify a result can all matter. If a vendor bundles components, map what its allowance covers before adding separate estimates; otherwise, the same expense can be counted twice.

Longer reasoning can increase inference without a longer answer

A short answer can follow substantial model computation. In reasoning-capable systems, internal reasoning may consume billable tokens even when the user sees only a concise result. Claude's thinking documentation states that thinking tokens are billed as output tokens and that visible thinking can differ from the amount billed. Reasoning carried into later context can also incur input charges under its documented rules.

The practical consequence is that answer length and prompt count are incomplete measures of usage. Giving a model more room to reason, asking it to compare alternatives, or extending a task through more tool calls can increase the work behind the final response. A higher reasoning limit allows more work; it does not prove that the model used the entire allowance or produced a better result.

For a token-priced service, calculate inference cost by summing each billed token category multiplied by its applicable rate across every call. Separate ordinary input, cached input, and output according to the provider's rules. If output already includes reasoning tokens, do not add reasoning a second time. Track models separately: equal token counts across different models do not imply equal cost or equal physical computation.

Parallel agents change speed, capacity, and total work differently

Suppose four agents divide a fixed document-review task into four non-overlapping parts. If they perform the same aggregate work as one agent, parallel execution may reduce elapsed time without quadrupling consumption. It can still raise peak demand because several model calls or environments are active together.

Now suppose the agents pursue four independent research approaches, compare competing explanations, or review one another's findings. The task has expanded. There is more investigation to perform, plus the work of assigning tasks and combining results. Repeated context, overlapping searches, retries, and discarded drafts can add further consumption. Those are possible costs of the chosen workflow, rather than an automatic multiplier attached to the word “parallel.”

There is a concrete historical example. In its June 2025 account of a multi-agent research system, Anthropic reported that agents used roughly four times the tokens of chat interactions in its data, while multi-agent systems used roughly fifteen times. The figures describe that company's observed workloads. They are neither universal ratios nor a forecast for every agent product.

The distinction is between elapsed time, peak concurrent capacity, and aggregate consumption. A system can finish faster while using more resources overall. It can also finish faster with similar total work. Measure all three.

How demand can grow with no new users

For a defined period, a useful planning relationship is:

Total inference usage = active users × tasks per user × average inference usage per task.

The final term includes every participating agent, the coordinating model, and any retries. Do not multiply it by agent count again. Within a comparable model and workload, token volume can help track this relationship; it is not a universal measure of accelerator-hours or electricity.

Here is a deliberately hypothetical example. Assume an unchanged model and billing-category mix, and count all input and output tokens across the entire task.

Swipe sideways to see all columns →

MeasureInitial workflowExpanded workflow
Active users100100
Tasks per user per day48
Average tokens per task, all agents included20,000100,000
Total tokens per day8 million80 million

User count stays flat while token volume rises tenfold: twice as many tasks, with five times the aggregate token usage in each task. The larger task budget could support longer reasoning, additional steps, or independent agent investigations. These numbers illustrate multiplication; they report no benchmark or expected adoption rate. They do not establish a tenfold increase in hardware requirements or spending.

Environment usage needs its own calculation. For example, ten equally sized sandboxes running for one hour total ten sandbox-hours, whether one person or ten people initiated them. CPU activity, allocated memory, persistence, and network use then determine the relevant meters. Nor does every subagent necessarily need its own sandbox.

Cheaper inference can coexist with greater total demand

Lower unit prices and more efficient models can reduce the cost of a fixed workload. Caching, smaller models, shorter context, and better task design can also reduce expense. Aggregate spending can rise if additional or more intensive workflows more than offset lower unit costs.

That is a plausible mechanism, not an inevitable outcome. For illustration, a tenfold increase in comparable token usage combined with a halving of every applicable token rate would produce five times the inference bill. Different changes in usage or rates produce different results. Provider revenue, token volume, and physical infrastructure demand should not be treated as interchangeable measurements.

The opportunity is that useful agent work need not stop when a person stops typing. A person could initiate a document comparison, a separate investigation, and a verification pass from one request. The limit becomes the value of the results, the available budget, system capacity, and the ability to review them.

What should firms measure before expanding agent use?

For a law firm, a practical unit is a completed, reviewed task that meets a defined quality standard. A first draft, an impressive token count, or a fast response does not establish that standard. Compare the same document set and required output, and include omitted issues and correction time in the assessment.

Record environment usage and inference charges separately. Track model and token categories, calls and retries, elapsed time, peak concurrency, and reviewer time. Put limits on total spending and delegation depth, and stop work when further investigation no longer improves the result enough to justify its cost. Our guides to evaluating agents and choosing a bounded pilot provide a way to connect those measurements to real workflows.

User growth remains a useful measure of adoption. Understanding agent demand requires another question: How much work does each user now set in motion? Longer reasoning and wider delegation can make that answer change far faster than the number of people using the system.