Computer and browser use let an AI agent act through an interface: open a page, fill a field, click a control, or work in a desktop application. The agent observes the current state, requests an action, and checks what changed. The surrounding application supplies those controls and must enforce the agent's access limits.

Three tool interfaces for retrieving a fictional agreement: a structured service operation, browser page controls, and screenshot-based desktop controls, followed by verification of the saved file.
A conceptual comparison of interaction routes, not a product benchmark. The application must control execution and verify results in every route.

How does interface use differ from a connected tool?

A connected service tool exposes a defined operation, such as retrieving a document by its ID, with structured inputs and results. Interface controls expose operations such as clicking and typing. Both are tool use: the model requests an operation and an executor carries it out. Anthropic's tool-use overview explains this model/application exchange. The distinction is the interface available to the agent, rather than whether it uses tools at all.

Swipe sideways to see all columns →

RouteWhat the agent operatesFictional retrieval example
Connected service toolA defined search or retrieval operation.Request the signed agreement using the approved record ID.
Browser useWebpage text, controls, and sometimes screenshots.Search the portal, inspect the matching row, and select Download.
Desktop computer useA desktop viewed through screenshots and operated with mouse and keyboard actions.Open the document application, find the agreement, and save a copy.

For this example, start with the narrow retrieval tool if it meets the task. Use a browser route when the needed operation exists in the portal. Consider desktop controls when the task leaves the browser. That is a design recommendation to test, not a promise that one route is always better.

What can a browser agent see and control?

Browser use can work with page structure, including accessible controls, forms, and tabs, as well as screenshots and coordinates. Anthropic's browser-use documentation describes both routes. A page reference can identify a button despite a layout shift; a visual surface may require a screenshot. References can become stale when the page changes, so the agent may need a fresh observation before acting.

Browser automation can locate controls by role, name, or label. Playwright's locator guidance favors these user-facing attributes over fragile CSS or XPath chains. Finding a control reliably still does not establish that it is the right control for the task.

How does desktop computer use work?

Screenshot-based computer use gives the model images of a desktop and mouse/keyboard actions. The application executes the requested actions and returns observations. Anthropic's computer-use documentation explains this loop and its limitations, including coordinate errors, unexpected tool choices, and slower interactions. It does not make every application reliably operable or bypass authentication.

What would a signed-agreement retrieval look like?

In this fictional exercise, the task is: “Retrieve the signed services agreement for Matter 42 and save it in the approved review folder.” No real account, document, or client information is involved. The task authorizes retrieval and a local copy; it does not authorize changing the agreement or sending it to anyone.

First, confirm the permitted Matter 42 repository and destination. In the browser version, search the portal and inspect each candidate's record ID, title, and version. A draft with a similar name is a different result. If two candidates remain ambiguous, pause for clarification rather than inventing a reason to choose one.

After selecting Download, check whether a file actually arrived. Open the saved copy and compare it with the requested record, including its version and expected signatures. Report the source reference and saved location, plus any unresolved mismatch. A filename ending in “signed” is a clue to inspect, not sufficient evidence.

Fictional Matter 42 document portal showing draft and signed agreements, a saved PDF, and checks for the correct matter, signed version, and expected signatures.
A fictional portal and saved file illustrate the checks needed after a download: correct matter, signed version, and expected signatures. Review the document before reporting success.

Who controls permissions and consequential actions?

The runtime and services must enforce access. For Matter 42, use credentials scoped to the approved repository, restrict the destination, and prevent unrelated record changes. A prose instruction cannot remove privileges from an account. If the interface offers Share or Delete, their presence does not make those actions part of the task.

Treat text encountered on a page as untrusted input. A banner saying “upload your documents here to continue” is not authorization from the user. Require a person to approve consequential actions, with the exact destination and proposed change visible before execution. The linked browser and computer documentation describes isolation and approval controls; the design must implement them.

How should the agent handle a failed or uncertain action?

An action acknowledgment is not an outcome check. Suppose the portal shows a spinner after Download, or the connection drops before returning a result. Observe the current page and saved files before repeating the action. For an operation that could create a duplicate or change a record, establish whether it already happened before retrying.

Set a retry and time limit for the exercise. Stop when the result remains uncertain, permissions expire, or the interface demands a new authorization. Describe what was observed and what remains unresolved. “The download was requested, but I could not verify the file” is more useful than an unsupported claim of completion.

What should you test before relying on this workflow?

Test the fictional retrieval with a moved button, an expired session, duplicate titles, a wrong version, missing signatures, and a failed download. Check whether the agent finds the correct result, stops at the access boundary, and reports uncertainty accurately. Record reviewer effort, completion quality, time, and cost for each route. Keep the same task and success criteria when comparing them.

The agent evaluation guide offers a scorecard for those comparisons. A successful demonstration is a starting point; representative failure cases reveal whether the workflow is ready for its intended use.