Browser Use
Let any vision + tool-calling model drive a real Chromium browser — navigate, click, type, scroll, and read live pages, with screenshots shown to you as it works. Live now in the Playground.
From idle model to live browsing
No separate setup, no API key for a browser provider — it's a toggle inside the chat you're already using.
Pick a vision + tools model
Browser Use needs a model that can both see (screenshots) and call tools — Claude, GPT-5, Gemini, and most frontier models qualify. The toggle only lights up once you've picked one.
Click the globe icon, confirm
A quick research-preview notice explains the billing before anything starts. Confirm and a real Chromium sandbox spins up in the cloud — billing begins the instant it does.
Just chat
Ask it to look something up, compare prices, check a site, fill out a form — it drives the browser itself, and every action shows up inline with a screenshot as it happens.
The model sees what it clicks.
Every action — a click, a scroll, a page load — captures a screenshot that's shown to you inline and handed back to the model, so it's reasoning about the real page state, not guessing.
Eight tools, one browser
The model calls these directly — same names it would see in a raw tool-call trace — one action at a time, up to 10 per reply.
navigate
Loads any http(s) URL. file://, javascript:, and data: URLs are refused outright.
click
Clicks by CSS selector or plain visible text — "the Sign In button" works as well as a selector.
type
Fills a field (by selector or label text) and can press Enter to submit, e.g. a search box.
scroll
Scrolls the page up or down by a configurable pixel amount.
press_key
Sends a single keystroke — Enter, Escape, PageDown, and the like.
screenshot
Captures the current page as JPEG — shown to you inline and fed back to the model so it can see the result.
read_text
Extracts the page's visible text (truncated to ~6,000 characters) — cheaper than a screenshot when it just needs to read.
go_back
Navigates back one step in browser history, same as a browser's back button.
Metered by the minute, not the session
No flat session fee. You're billed for exactly how long the sandbox is running — starting the instant you enable it.
Rate
- Ceiling-billed — a 10-second session still costs a full minute.
- The first minute is charged the moment you click Enable.
- Every minute after is metered against the same balance as your regular model usage.
- Drawn from lite credits first, then pro credits.
- Capped at 30 minutes per session — 0.15 credits, worst case.
Examples
- 1 min0.005 cr
- 5 min0.025 cr
- 15 min0.075 cr
- 30 min (session cap)0.15 cr
On top of this, normal per-token model costs still apply to every reply the model generates during the session — Browser Use billing only covers the browser sandbox itself.
What to expect from a research preview
Research preview
This is early. Behavior, tools, and pricing may change — expect occasional bugs.
http(s) only
No local files, no script URLs — the browser can only reach real web pages.
30-minute session cap
Sessions don't renew — once 30 minutes pass, the sandbox self-destructs and a new session is needed.
10 steps per message
Each reply gets up to 10 tool calls before the model has to wrap up and answer with what it has.
Browser Use, straight from the API.
The same capability is headed to /v1/chat/completions as a built-in tool — no separate session-management code required, so agents built against the API get the same live-browsing ability the Playground has today.
Common questions
The wall-clock time your browser sandbox is running — 0.005 credits per minute, ceiling-billed. Enabling it charges a full minute immediately, then every minute after is metered as you use it, even if you're just reading the response and haven't sent another message yet.
Give your model a browser.
Open the Playground, pick a vision + tools model, and click the globe icon — the first minute is 0.005 credits.