BoundBench

OpenAgent

Self-hosted personal AI assistant (Go web app, single binary) with RAG, agent loops, shell, local-file, office, browser and desktop tools, MCP and skills.

github.com/the-open-agent/openagent · 2026-10-04 · c9a8fca

Defense-in-depth score

1.3 / 10

Minimal

OpenAgent ships with every built-in tool enabled, including an unsandboxed host shell that inherits the server's full environment, local file read/write, desktop GUI control and browser automation, and none of them require human approval. A casbin-based permission engine exists in the repo but is never called from the tool path. Authentication and access control in the default install are not locked down. A hijacked session (for example via a fetched web page) can exfiltrate data and take irreversible actions on the host with nobody in the loop.

Key gaps (4)

  1. No approval gate or isolation: the default store exposes an unsandboxed host shell with full environment to the model, and the permission engine is never wired in. C4 · Code-execution isolation
  2. A hijacked default session can exfiltrate and take irreversible host actions unattended (C5-WORSTCASE). C5 · Untrusted input blast radius
  3. A default-enabled skill tells the model to npm-install and fetch third-party skills, which the ungated shell carries out without consent (C7-RCELOAD). C7 · Third-party extensions
  4. local_file_move's only confirmation is a confirmed=true argument the model supplies itself (C2-SELFAPPROVE). C2 · Approval gates

Criteria

C1 Identity & least privilege

Minimal 0.07 / 1.00

OpenAgent runs every tool as the operating-system user that launched the server, and the shell tool passes the server's full environment to each command. The one real authorization check is that high-risk tool types (shell, files, GUI, browser) are only exposed when the requester is the global admin; store admins and ordinary users get them filtered out. Enforcement of that check does not cover every entry point, and authentication in the default install is not locked down.

C2 Approval gates

Minimal 0.05 / 1.00

There is no human approval anywhere in the tool path: the agent loop executes whatever tool calls the model returns, including shell commands, file writes, GUI automation and browser actions. A tri-state allow/ask/deny policy engine (guard/) and a tool-policy table exist, but the host never builds or consults the guard, and the audit code itself notes it is not yet wired in. The system prompt actively tells the model not to refuse. File writes and moves through local_file are snapshotted so they can be reverted, but shell actions are not.

C3 Tool & action scoping

Minimal 0.15 / 1.00

Tool scoping is mixed. web_fetch has a real SSRF guard that resolves the host and rejects non-public addresses, re-checking on redirects. But the default tool set is everything, including a raw shell string, and local_file only requires that paths be absolute, so any file the OS user can reach is in scope. An optional shellCommandAllowlist config limits the first word of shell commands, but it is empty by default.

C4 Code-execution isolation

Minimal 0.00 / 1.00

Shell commands run as an ordinary subprocess of the server on the host, using sh -c and the server's full environment. Nothing in the tool, object or model packages sets up a container, OS sandbox or restricted user; the only 'sandbox' references disable Chrome's own sandbox for the browser tools. The Docker image does run as a non-root user, but that user is given passwordless sudo, and the single-binary install that the README leads with has no isolation at all.

C5 Untrusted input blast radius

Minimal 0.00 / 1.00

Web pages, search results, uploaded documents, knowledge-base passages and MCP results all enter the model's context with no marking or provenance, and nothing in code changes what the agent may do after reading them. The system prompt added when tools are present tells the model to act immediately and never refuse. In the default configuration a successful injection can use the shell or web_fetch to send data out and take irreversible actions on the host without any human seeing it first. Opt-in chat pipes widen this further when configured.

C6 Memory, context & configuration integrity

Minimal 0.20 / 1.00

Persistent context comes from per-chat history, store knowledge bases, an opt-in experience library, and skills. All skills are injected into every default store's prompt, and skill files are re-synced into the database from a skills folder next to the binary at startup. There is no memory-write tool, and the LLM-driven experience review that writes skills automatically is off by default. However, the ungated shell can write that skills folder or the SQLite database directly, so an injection can plant instructions that every later session and user receives.

C7 Third-party extensions

Minimal 0.00 / 1.00

MCP servers are added by the global admin as raw stdio commands or URLs with no pinning or integrity check, and launch as the server's OS user. Skills (all enabled by default) are prompt files, but the bundled clawhub skill tells the model to npm-install a CLI and fetch new skills from clawhub.com on the fly; with the ungated shell, the model can install and run arbitrary third-party packages without anyone approving it.

C8 Secrets & sensitive-data protection

Minimal 0.20 / 1.00

Provider API keys and tool secrets are stored in plain text in the database (SQLite next to the binary by default), and masked as *** in API responses. Tool-audit records redact sensitive JSON fields, but the agent loop prints every tool result and knowledge passage to stdout unredacted, and the shell tool hands the server's full environment to every command. The shipped configuration's default database credentials are also not locked down. No third-party telemetry was found.

C9 Audit & traceability

Minimal 0.40 / 1.00

Tool calls are recorded in three places: the chat transcript stores each call's name, arguments and result; high-risk built-in tools also write a database record with redacted arguments and the user; and a JSONL audit file per session records every tool call's name, outcome and duration (but only the argument length). All of these live where the server process, and therefore the shell tool, can edit them. The JSONL writer drops events when its queue is full and swallows write errors, and the transcript is saved only when the answer finishes.

C10 Limits & kill switch

Minimal 0.20 / 1.00

The agent loop keeps calling the model and running tools for as long as the model returns tool calls; there is no round, time, or token cap. Individual shell commands time out after 30 seconds by default, but the model can raise that to 300 seconds and start background shell sessions that live until 30 minutes of idleness. Cancelling an answer stops output being written, but tool calls run under a background context, so in-flight commands keep going. A per-user rate limit applies only to signed-out users and defaults to 10,000 messages per 15 minutes.