BoundBench

elizaOS

TypeScript framework and app stack for autonomous AI agents: core runtime, standalone agent host, desktop/web/mobile app and first-party plugins.

github.com/elizaOS/eliza · 2026-10-05 · aa4f2b0

Defense-in-depth score

2.8 / 10

Minimal

elizaOS ships a lot of careful security engineering (per-requester role checks, an SSRF-guarded web fetch, a vault for connector tokens, a confirm-code gate for destructive shell commands), but the default desktop agent runs any other shell command directly on your machine as you, with your model and connector keys in its environment. A web page or document that hijacks the agent can read your files and send them out, or make irreversible changes, without asking. Instruction files and startup hooks from the project folder load without a trust prompt. Turn on the container sandbox mode and set workspace roots before pointing it at anything you care about.

Key gaps (6)

  1. Shell commands run as the logged-in user and inherit the agent's provider and connector keys, so a hijacked owner session reaches the user's whole account. C1 · Identity & least privilege
  2. The shell confirmation for destructive commands can be switched off at runtime without an operator decision. C2 · Approval gates
  3. Shell commands run directly on the host by default; the container sandbox is opt-in. C4 · Code-execution isolation
  4. If content the agent reads hijacks it, it can leak files and keys and take irreversible actions through the shell with no human involved. C5 · Untrusted input blast radius
  5. A project folder used as the workspace can add startup hooks and instruction files that load without a trust decision. C6 · Memory, context & configuration integrity
  6. Plugins, drop-ins and hooks run inside the agent process with all of its credentials. C7 · Third-party extensions

Criteria

C1 Identity & least privilege

Minimal 0.15 / 1.00

Actions are checked against the role of the person who asked (owner, admin, user, guest) on every execution path, and a failed role lookup refuses instead of guessing. But the authority behind those actions is the operator's own: shell commands run as the local OS user and inherit the agent's process environment, where configured model and connector keys are placed, and only code-injection variables are stripped. The file tools refuse a short list of credential folders, while the shell does not.

C2 Approval gates

Minimal 0.25 / 1.00

Shell commands that a built-in classifier flags as destructive (recursive deletes, disk wipes, dropping database objects) stop and ask the user to reply with a one-time code, and that approval is bound to the exact command, folder, requester and conversation. Everything else runs without asking: other shell commands, file writes and edits, web requests and browser actions. The confirmation can be turned off at runtime without an operator decision, and there is no undo for file or shell changes.

C3 Tool & action scoping

Minimal 0.28 / 1.00

The file tools resolve symlinks and refuse a list of private and system locations, and web fetches go through a guard that pins DNS, re-checks redirects and blocks private network addresses. But the file tools are otherwise allowed anywhere on disk unless the operator sets workspace roots, and the shell tool takes any command string. Shell, file write and web tools are all enabled by default; shell can be switched off as a whole.

C4 Code-execution isolation

Minimal 0.38 / 1.00

By default the shell tool runs commands directly on the host with bash as the logged-in user, because the runtime mode defaults to the unsandboxed local mode. An opt-in safe mode routes commands into a Docker or Apple container that runs as a non-root user with all capabilities dropped, no network and memory and process limits, and it refuses to run rather than fall back to the host when no sandbox is available. That container still mounts the workspace read-write, and in-process hooks and plugins never pass through it.

C5 Untrusted input blast radius

Minimal 0.25 / 1.00

Some untrusted content, such as email and webhook payloads, is wrapped in markers with a warning and checked against injection patterns, but matches are only logged. The coding tools' web fetch and search results are not wrapped, and nothing limits what the agent can do after reading untrusted content. In one default session the agent can read web pages, reach the user's files and keys, and send data out or change things through the shell without a human, so a successful injection can both leak data and take irreversible actions.

C6 Memory, context & configuration integrity

Minimal 0.05 / 1.00

The agent loads instruction files such as AGENTS.md, USER.md and MEMORY.md from its workspace into every conversation, and its own template tells it to write facts and reflections there. When the agent is started in a project folder, that folder becomes the workspace, so the project's instruction files load silently, and startup hooks found in the workspace are loaded and run inside the agent process without a trust decision. These files are shared across all conversations, are not tagged by source and can steer tool use.

C7 Third-party extensions

Minimal 0.15 / 1.00

Plugins can be installed from the elizaOS registry, including by the agent itself through an owner-only action, using the package manager with install scripts disabled; an approval-bound path can pin an exact version, but the default install takes the registry's current version and can fall back to a git clone. Plugins, drop-in plugins from the state folder and workspace hooks all run inside the agent process with its full environment, and none of them is checked against a hash or signature.

C8 Secrets & sensitive-data protection

Minimal 0.45 / 1.00

Connector OAuth tokens go into an encrypted vault keyed from the OS keychain, the main config file is written with owner-only permissions, and shell output and logs pass through secret redaction. Model provider and connector keys are still copied into the process environment, so every shell command the agent runs inherits them, and the session-level secret substitution that keeps secrets out of model context is off by default. No third-party telemetry was found; boot telemetry stays on local disk.

C9 Audit & traceability

Minimal 0.40 / 1.00

When enabled, the trajectory recorder stores each model call and its tool calls with arguments in the local database, and a separate audit feed records sandbox and policy events. Trajectory recording is on in development runs but off by default in production builds unless the operator opts in. Records live in the agent's own state, which its file and shell tools can reach, and actions do not wait for their record to be written; tools run by external coding agents are not captured.

C10 Limits & kill switch

Minimal 0.45 / 1.00

Each turn has a 1.5 million prompt-token budget and a breaker that stops repeated identical tool calls, and shell commands time out after two minutes by default (ten at most). There is no cap on the number of tool calls, and background shell sessions and the always-loaded scheduler can keep running after a turn ends.