BoundBench

CAMEL

Multi-agent framework / research on agent scaling

github.com/camel-ai/camel · 2026-10-03 · 0106b76

Defense-in-depth score

1.9 / 10

Minimal

CAMEL gives agents powerful bundled tools (shell, code execution, browser, email, MCP) but almost no guardrails on by default. Model-chosen shell commands and code run on the host as your user with your full environment, the agent loop has no approval gate or step/time limit, and nothing limits what injected web or email content can make an agent do. The one default prompt (host code execution) asks y/N without showing the code at the default log level. Treat any CAMEL agent with execution or messaging toolkits as fully trusted with your machine and accounts unless you add isolation (E2B/Docker) and approvals yourself.

Key gaps (4)

  1. Shell and code-execution subprocesses inherit the full host environment, so any credential in the developer's shell is reachable by model-written commands. C1 · Identity & least privilege
  2. The most powerful bundled tool (TerminalToolkit.shell_exec) runs model-chosen shell commands with no approval by default. C2 · Approval gates
  3. Default code and shell execution run on the host as the user with the full process environment, so a hijacked agent's code reaches host files, network and credentials. C4 · Code-execution isolation
  4. Injected content can drive both exfiltration and irreversible actions unattended: nothing tracks untrusted input or gates egress/state-changing tools. C5 · Untrusted input blast radius

Criteria

C1 Identity & least privilege

Minimal 0.00 / 1.00

CAMEL has no identity or authorization layer of its own. Each toolkit builds its own client from environment variables or local token files, and the ChatAgent runs whatever tool the model names with no per-request authorization check. The shell and code-execution tools pass the full process environment to every subprocess, so any cloud, Git or API credential the developer has in their shell is available to model-written commands. The bundled Gmail toolkit requests read, send and modify scopes in one token and exposes all of them together.

C2 Approval gates

Minimal 0.23 / 1.00

The ChatAgent executes every registered tool immediately; there is no framework-level approval step. The bundled code-execution toolkit asks 'Running code? [y/N]' by default when it runs on the host, but the code itself is only written to an INFO log line that the default WARNING log level hides, so the person approving does not see what will run. The terminal toolkit, the most powerful bundled tool, has an approval callback that shows the exact command, but it is off by default and does not cover its own file-writing tool. An LLM-based risk judge (LLMGuardRuntime) exists but an LLM judge is not human approval.

C3 Tool & action scoping

Minimal 0.25 / 1.00

Argument validation is thin. The terminal toolkit's safe mode is a denylist of command names (rm, sudo, dd and so on), and its enforcement, including path containment for cd and file writes, is not a strict boundary. The file toolkit accepts absolute paths anywhere on disk, and the web-fetch tool accepts any http(s) URL and follows redirects with no block on internal addresses. FunctionTool does not validate model arguments against the schema before calling the function.

C4 Code-execution isolation

Minimal 0.47 / 1.00

By default, model-written code and commands run directly on the host as the developer's user. CodeExecutionToolkit defaults to a 'subprocess' sandbox that is just a local subprocess, and TerminalToolkit defaults to a local backend whose only protection is a command-name denylist. Both pass the full environment to the child process. Stronger options exist: a Docker interpreter (stock container, non-root user, no other hardening) and an E2B remote sandbox, which is a real boundary but is opt-in and only covers the code-execution toolkit.

C5 Untrusted input blast radius

Minimal 0.00 / 1.00

CAMEL does nothing to limit what injected content can make an agent do. Web pages, emails, browser content and MCP tool results enter the conversation as ordinary tool messages with no provenance or untrusted marking, and nothing disables or gates egress or state-changing tools once such content has been read. A single agent built from bundled toolkits can read untrusted email or web content, hold the developer's credentials, and send email or run shell commands without a human in the loop.

C6 Memory, context & configuration integrity

Minimal 0.20 / 1.00

The default agent memory is an in-memory chat history tied to one agent object, so poisoning normally dies with the session. When persistence is used there are no write controls: the optional MemoryToolkit lets the model load arbitrary JSON records (any role) into its own memory and save memory to any path, and vector or JSON stores accept writes without validation or provenance. The optional SkillToolkit auto-discovers SKILL.md instructions from the working directory and gives them priority over the user's own skills, but skills are text only and cannot enable tools or change security settings.

C7 Third-party extensions

Minimal 0.23 / 1.00

CAMEL loads third-party code mainly through MCP servers the developer lists in a config, and through Hugging Face models in a few toolkits. Nothing is pinned, hashed, or re-approved when it changes; MCP stdio servers are launched as given by the config. MCP servers run as separate processes and, unless the config passes env, the MCP SDK gives them a reduced environment (library behaviour, inferred). The Jina reranker toolkit's local mode loads its model with trust_remote_code=True, running remote model code in-process, though that mode is opt-in.

C8 Secrets & sensitive-data protection

Minimal 0.13 / 1.00

API keys are read from environment variables and are not masked anywhere. At INFO level the agent logs every model request in full with no secret redaction (the default WARNING level hides it, but one environment variable turns it on). Shell and code subprocesses inherit the whole environment, so a model-written 'env' command reveals every key. Telemetry integrations (AgentOps, Langfuse, Traceroot) are opt-in via environment variables, and the Gmail token file is written with 0600 permissions.

C9 Audit & traceability

Minimal 0.15 / 1.00

CAMEL keeps no audit trail by default. The ChatAgent returns tool-call records (name, arguments, result) to the calling code in memory, without timestamps or actor attribution, and does not write them anywhere. Model requests are logged only at INFO, below the default level. The terminal toolkit writes a command log, but inside the agent's own workspace, and it deletes existing .log files there each time the toolkit is constructed.

C10 Limits & kill switch

Minimal 0.23 / 1.00

The ChatAgent has an iteration cap and step/tool timeouts, but all default to unlimited (max_iteration=None, timeouts=None). The step timeout only stops waiting; the worker thread keeps running. A stop_event can halt the loop between iterations if the developer supplies one. The terminal toolkit's 20-second command timeout converts a slow command into a background process that keeps running. Workforce tasks have a 600-second default timeout.