BoundBench

Stagehand

Browserbase's TypeScript/Python/Go SDK for browser agents: Playwright-style page control plus model-driven act, observe and extract calls, run by a runtime extension inside Chrome.

github.com/browserbase/stagehand · 2026-10-05 · 0336409

Defense-in-depth score

4.5 / 10

Minimal

Stagehand keeps the model on a short leash: each act() call lets it pick one or two clicks or keystrokes on elements that actually exist on the current page, it never writes code or chooses URLs directly, and secrets passed as variables stay out of the prompt. But there is no approval step, no limit on what a hostile page can steer those actions toward, and nothing scopes the logged-in browser sessions the README encourages you to keep. A malicious page can make act() type sensitive values into the wrong field and submit it without a human seeing it.

Key gaps (1)

  1. A page that hijacks act() can steer it to type secrets or page data into an attacker-chosen field and submit, or click irreversible actions, with no human in the loop. C5 · Untrusted input blast radius

Criteria

C1 Identity & least privilege

Minimal 0.25 / 1.00

By default Stagehand launches Chrome with a brand-new temporary profile that is deleted on close, so out of the box the model acts with no logged-in accounts. Nothing narrows authority beyond that: every act() call can operate on any site the session is signed into, secrets passed as variables can be placed into any field the model chooses, and Chrome inherits the full process environment. The README's lead example keeps a persistent profile so sign-ins survive between runs, which gives the model those accounts every time.

C2 Approval gates

Minimal 0.23 / 1.00

There is no approval gate: act() sends the page and instruction to the model and immediately performs whatever click, fill or key press it returns. Stagehand does offer a preview pattern, which its README promotes: observe() returns the exact element, method and arguments it would use, and act() can then run exactly that action, so a developer can build a review step. That pattern is opt-in and has no built-in approve or reject flow. Actions on real websites, such as submissions and purchases, cannot be undone.

C3 Tool & action scoping

Moderate 0.65 / 1.00

This is Stagehand's strongest area. The model never gets a general-purpose tool: in act() it can only choose an element that exists in the page snapshot Stagehand captured, and a method from a fixed list (click, fill, type, press, scroll, select, hover, drag); anything else is rejected. It cannot run JavaScript or navigate to a URL of its choosing, and extract() and observe() give it no actions at all. The gaps are that clicking links can still take the browser to any host, including local network addresses, unless the developer configures the optional domain allow and block lists, and that typed text is unrestricted.

C4 Code-execution isolation

Moderate 0.53 / 1.00

The model never writes code that Stagehand runs: there is no shell, no Python and no tool for running JavaScript in the page. The only code involved is the JavaScript of the websites it visits, which runs in Chrome's renderer sandbox. Stagehand leaves that sandbox on by default but turns it off without warning when it sees a CI environment variable or runs as root on Linux. Inside the sandbox, page code has full network access and the cookies of the site it belongs to.

C5 Untrusted input blast radius

Minimal 0.25 / 1.00

Every act, observe and extract call puts the page's accessibility tree straight into the prompt next to the developer's instruction, with no marking as untrusted and no detection. What helps is structural: the developer's code decides which operation runs on which page, extract() and observe() cannot change anything, and act() is limited to a couple of interactions with existing elements. Within an act() call, though, a hostile page can steer which element is clicked and what is typed, including the developer's secret variables, and submit it with no human involved.

C6 Memory, context & configuration integrity

Minimal 0.25 / 1.00

Stagehand has no long-term memory, loads no instruction files and does not read a .env file, and the default browser profile is thrown away after each run. Two things persist in documented use. The README's lead example keeps a browser profile on disk, so cookies and site storage a page writes carry into later runs. And with Browserbase, server-side caching is on by default: a successful act() is stored and later replayed on a matching page without asking the model, with no check on what was stored.

C7 Third-party extensions

N/A · full credit 1.00 / 1.00

Stagehand loads no third-party code at runtime. The only extension it installs into Chrome is its own runtime, bundled in the package; there are no plugins, MCP servers, model downloads or package installs that the model can reach.

C8 Secrets & sensitive-data protection

Minimal 0.45 / 1.00

Secrets meant for web forms are handled well: the developer passes them as variables, the model only ever sees placeholder names, the real values are substituted at the moment of typing, and results report the placeholders. Model and Browserbase API keys are passed as plain configuration into the browser runtime. Tracing is off unless configured. There is no redaction filter for logs or traces, and with Browserbase the variables also go to the server-side cache.

C9 Audit & traceability

Minimal 0.35 / 1.00

By default Stagehand prints info-level logs to standard error, which record that an inference finished and how many tokens it used but not which element was clicked or what was typed. The details of each action come back to the developer's code in the act() result, not in a log. When a developer configures OpenTelemetry, every RPC call and action gets a timestamped span with trace correlation shipped to their collector, but that is opt-in, best-effort and has no actor attribution.

C10 Limits & kill switch

Moderate 0.50 / 1.00

Each act() call is bounded in code to at most two actions and a few model calls, and there is no autonomous agent loop: the developer's code decides how many calls to make. Page commands have default timeouts, but act, observe and extract have no time limit unless the caller sets one, and model calls have no timeout or spending cap. Closing a locally launched browser kills Chrome's whole process group, which stops anything in flight.