BoundBench

AutoGPT Platform

Platform for building, deploying and running continuous agents

github.com/Significant-Gravitas/AutoGPT · 2026-10-03 · f8b0e0a

Defense-in-depth score

4.2 / 10

Minimal

AutoGPT agents act with every account a user connects and can browse, fetch, run a sandboxed shell, call any block and run on schedules or webhooks. Code execution is well contained by default (bubblewrap with no network), credentials are encrypted and kept out of the model, and AutoPilot pauses before flagged irreversible blocks. The dominant risk is prompt injection: nothing stops a hijacked agent from sending data out through web fetch or the HTTP block and acting through unflagged blocks, and builder graphs never pause by default. The stronger approval modes and the injection judge ship behind a feature flag that is off on a self-hosted install.

Key gaps (2)

  1. The generic HTTP request block, catalogue MCP writes and browser actions run with no human approval by default, so the approval gate is bypassed by the most general outward action path. C2 · Approval gates
  2. A prompt-injected session can exfiltrate data and take irreversible actions unattended: no default control limits egress or state change after reading untrusted content, and builder graphs triggered by webhooks never pause. C5 · Untrusted input blast radius

Criteria

C1 Identity & least privilege

Moderate 0.53 / 1.00

Agents act with the credentials each user has connected (OAuth grants and API keys), and every block execution looks those credentials up by the requesting user's id, failing if the credential is not theirs. OAuth logins request the scopes the blocks ask for, but one credential serves both reads and writes for the whole run, and nothing issues short-lived or per-task tokens. When the optional E2B sandbox is enabled, the AutoPilot shell receives a token for every connected provider. A hijacked agent therefore holds write access across all of a user's connected services.

C2 Approval gates

Minimal 0.25 / 1.00

By default AutoPilot pauses before running a block that is flagged irreversible (for example send email, pay cart, post to Slack) and shows the user the exact input, which they can edit, approve or reject. But the flag covers only some blocks: the generic HTTP request block, catalogue MCP tools, web fetch, browser actions and the shell all run without asking. Graphs built in the visual builder default to never pausing. The stronger per-chat approval system (Ask First / Auto modes) is shipped behind a feature flag that is off on a self-hosted install, and in its own default Auto mode it lets an LLM judge decide shell and platform actions.

C3 Tool & action scoping

Minimal 0.40 / 1.00

The shared HTTP client is a real SSRF defence: it resolves every DNS answer, blocks private and metadata ranges, pins the connection to the checked IP and re-validates every redirect. Workspace file paths are checked with realpath containment, and the SQL block defaults to read-only. But the default AutoPilot tool set also includes a raw shell, an arbitrary-public-URL HTTP block and a browser whose URL validation does not cover every path, and every tool is on by default.

C4 Code-execution isolation

Moderate 0.60 / 1.00

On a default self-hosted install the AutoPilot shell runs inside bubblewrap: a cleared environment, no network, a read-only system filesystem, a per-session writable workspace, process and memory limits and a 120-second cap. If bubblewrap is missing the tool refuses rather than running on the host. There is no seccomp filter, and the agent-browser's Chromium runs directly in the root backend container with that container's full environment. The E2B cloud sandbox is stronger isolation but is only used when an E2B key is configured, and then receives the user's integration tokens and unrestricted egress.

C5 Untrusted input blast radius

Minimal 0.23 / 1.00

Agents read web pages, search results, browser content, MCP tool output, webhooks and messages from other agents, and nothing in the default configuration limits what they can do afterwards. A hijacked AutoPilot can send data out through web fetch or the HTTP block and take irreversible actions through unflagged blocks, and builder graphs triggered by webhooks run unattended. The platform has a content judge that holds suspicious reads for review, but it is an LLM classifier and only runs when the opt-in approval-mode flag is on.

C6 Memory, context & configuration integrity

Minimal 0.30 / 1.00

The Claude Agent SDK is started with no settings sources and a strict MCP config, so no project or home-directory files can add tools, hooks or servers. But AutoPilot's self-written skills are on by default: the model can store a skill with no approval, and later sessions receive the skill index as trusted context and can load and follow it. Storage is per-user. Graph memory (Graphiti) is behind an off-by-default flag.

C7 Third-party extensions

Minimal 0.30 / 1.00

The platform loads no third-party code into its own processes: MCP servers are reached only over HTTPS, and agent-browser is version-pinned in the image. Remote MCP servers receive only their own credential and the arguments sent to them. However the model can point run_capability at any HTTPS MCP server URL on its own, server tool lists are fetched fresh with no pinning, and only write-looking tools on non-catalogue servers ask first.

C8 Secrets & sensitive-data protection

Moderate 0.50 / 1.00

Stored integration credentials are Fernet-encrypted with a per-install key, the backend refuses to boot on a missing or previously published key, OAuth tokens are SecretStr, Sentry events are scrubbed and telemetry is off unless configured. Credentials are injected at block execution and never sent to the model. But the shipped deployment defaults do not lock down credential handling. Subprocesses such as agent-browser inherit the backend's full environment.

C9 Audit & traceability

Moderate 0.63 / 1.00

Every chat message, including tool calls and results, is stored in Postgres, graph runs record who or what started them (schedule, webhook, AutoPilot chat) and their parent run, and every human review keeps its status and decision time. This gives a usable trail with actor attribution. The records are ordinary database rows that cascade-delete with the session or user, with no tamper evidence, and tool calls made inside Claude SDK sub-agents only reach application logs.

C10 Limits & kill switch

Minimal 0.45 / 1.00

AutoPilot turns are capped at 100 tool rounds and $10 of model spend, users have daily and weekly cost limits by default, the shell is capped at 120 s, and graph blocks time out after 30 minutes. But graphs have no spend ceiling on a self-hosted install (credits are off), sub-agent limits only gate launches, stopping a graph is cooperative, and schedules keep running after a chat stops.