What is a browser agent?
Short answer
A browser agent is an AI agent that operates a web browser to complete tasks: it observes the page through screenshots, the HTML or the accessibility tree, then clicks, types, scrolls and navigates step by step until the goal is done or it needs a person. Booking, research, form filling and testing are common uses.
How a browser agent works
- Observe: the agent captures the page state, as a screenshot, the DOM or the accessibility tree (roles, names and states of elements).
- Decide: a language model chooses the next action, such as “click the Search button” or “type the date into Departure”.
- Act: automation code, often built on tools like Playwright or Chrome DevTools Protocol, performs the action.
- Repeat until the task is complete, a step limit is reached, or a human must confirm something.
Products such as OpenAI’s ChatGPT agent, Anthropic’s Claude in Chrome and Google’s Project Mariner work this way, as do open-source frameworks for building your own.
Agent or API?
Browser agents are useful where no API exists, but they are slower and more fragile than calling an API or an MCP server directly. If you control the system, exposing a proper API or MCP tool is more reliable than asking agents to click through your interface.
Designing pages agents can use
- Use real buttons, links and form controls with clear labels; avoid clickable divs.
- Keep a logical heading structure and visible text for key information such as prices and policies.
- Avoid surprise pop-ups and layouts that shift while loading.
- Give every step a stable URL where possible.
These are the same practices that make a site accessible to people using screen readers.
Risks
An agent acting with your session can be steered by malicious page content (prompt injection). Limit its permissions, require confirmation for payments, messages and deletions, and log what it does.


