BoundBench

Webcmd

Self-learning browser automation CLI for AI coding agents: drives a real logged-in browser, with site memory, plugins and external-CLI passthrough.

github.com/agentrhq/webcmd · 2026-10-04 · 9d8ea44

Defense-in-depth score

2.5 / 10

Minimal

Webcmd gives an AI agent a real browser logged into the user's accounts and lets it run any program against any site, while its skill pre-approves every webcmd command in the host agent. Browser programs run in a QuickJS sandbox, but the bridge to Playwright is not a complete boundary. External-CLI passthrough and plugins run arbitrary unpinned code on the host. A prompt injection in any page can leak data and act on the user's accounts with no human in the loop.

Key gaps (5)

  1. A hijacked agent acts with every account logged into the browser profile and, via external-CLI passthrough, the user's gh/docker credentials. C1 · Identity & least privilege
  2. The shipped skill pre-approves every 'webcmd' command, so browser runs, plugin installs, and passthrough to registered binaries skip the host agent's approval prompt. C2 · Approval gates
  3. Registered external binaries run directly on the host, outside the QuickJS sandbox, whose host bridge is also not a complete boundary. C4 · Code-execution isolation
  4. Untrusted page content, the user's logged-in sessions, and unrestricted navigation share one session, so a prompt injection can leak data and act on the user's accounts unattended. C5 · Untrusted input blast radius
  5. Invoking an external CLI that isn't installed auto-runs an unpinned global npm install with no consent, and plugins from any git repo load in-process. C7 · Third-party extensions

Criteria

C1 Identity & least privilege

Minimal 0.25 / 1.00

Webcmd drives a real browser with whichever logged-in profile the agent names, so it acts with the user's full sessions on every site logged into that profile. Profiles are separate cookie jars, which lets a user keep a narrow 'work' or 'social' identity, but nothing inside a profile limits which site or action the agent can reach. Webcmd also passes external tools such as gh and docker the user's full environment and ambient credentials. A hijacked agent therefore carries the user's accounts across services.

C2 Approval gates

Minimal 0.07 / 1.00

Webcmd gives a host agent no approval signal for its most powerful tool: 'browser run' executes any program against any logged-in site and mixes reads and writes with no risk annotation, dry-run, or read-only mode. Site adapter commands must declare read or write, which helps. Worse, the shipped skill pre-approves every 'webcmd' command in Claude Code, so browser runs, plugin installs, and passthrough to docker, gh, or any binary the model registers skip the host's per-call prompt. The only payment safeguard is a sentence in the skill.

C3 Tool & action scoping

Minimal 0.15 / 1.00

The main tool is a general-purpose program runner: any URL, any site, any action in the logged-in browser, limited only by a short denylist of Playwright methods. The separate 'web fetch' command does block private and internal addresses, and adapter commands take typed arguments, but browser run, page.request, and external passthrough have no URL or argument allowlists. Every tool is on by default.

C4 Code-execution isolation

Minimal 0.25 / 1.00

Browser-run programs execute in a QuickJS engine compiled to WebAssembly, with no Node APIs and only injected host functions, which is a real boundary. But the host-side Playwright bridge uses a denylist that is not a strict boundary. Separately, external-CLI passthrough runs any binary the model registers directly on the host, and plugins load in-process. An escape lands in a Node process holding the user's files, network, and browser sessions.

C5 Untrusted input blast radius

Minimal 0.25 / 1.00

Webcmd's job is to read untrusted web pages into an agent that holds the user's logged-in sessions and can navigate anywhere, so all three Rule-of-Two legs are present in one session. Browser-run output is structured JSON with page URL and title kept apart from the result, but nothing marks content as untrusted or restricts what the agent can do after reading it. The skill's 'page content is untrusted' line is a prompt, not a control. A hijacked agent can exfiltrate by navigating to an attacker URL and can post, message, or buy with no human involved.

C6 Memory, context & configuration integrity

Minimal 0.45 / 1.00

Webcmd keeps 'site memory' in a local git repository under the user's home directory and feeds it back to later agents as navigation guidance. Writes go through checks: every fact needs a verification date, size limits apply, candidate observations are rejected if they look like secrets, and git history lets the user roll back. The first visit to a site may pull a seed from webcmd's cloud service, which is written into memory without the same fact validation. Nothing is loaded from the working directory, but memory has no expiry and can steer future actions.

C7 Third-party extensions

Minimal 0.00 / 1.00

Plugins install by cloning the latest commit of any git repository the agent names. Dependencies are installed with scripts disabled, but the plugin then loads in-process with full access. External CLIs are worse: invoking one that isn't installed auto-installs it with an unpinned global 'npm install -g' (install scripts run) and no prompt, and the model can register new CLIs with its own install command. All of this sits behind the skill's blanket pre-approval.

C8 Secrets & sensitive-data protection

Minimal 0.40 / 1.00

Webcmd has no telemetry and redacts secret-looking fields (cookie, token, authorization, and similar) from browser-run results, logs, and traces before they reach the agent. The hosted API key, an opt-in, uses the macOS keychain or a 0600 file. But the high-value secrets are the browser's session cookies, and their protection from browser-run programs is not a complete boundary. By default the first visit to each site also sends the domain to webcmd's cloud seed service. External CLIs inherit the full environment.

C9 Audit & traceability

Minimal 0.28 / 1.00

By default webcmd keeps no durable record of what the agent did in the browser. The daemon holds an in-memory buffer of the last 200 messages, only for failed commands, and any local caller can clear it. Detailed traces of actions, network, and console exist but are off by default and are written only when a command finishes. Site-memory changes are committed to git, which is the only durable record.

C10 Limits & kill switch

Minimal 0.40 / 1.00

Each browser run has a default 30-second wall-clock and CPU limit, a 128 MB memory cap, and an output cap. Ctrl-C cancels the in-flight run in the daemon. But the agent can raise the timeout to any value with --timeout, since there is no ceiling. External passthrough commands have no limit, and the daemon and browser keep running in the background after a command ends.