BoundBench

smolagents

Barebones library for code-writing agents

github.com/huggingface/smolagents · 2026-10-03 · c30b115

Defense-in-depth score

1.6 / 10

Minimal

As shipped, smolagents runs model-written Python inside your own process with an interpreter its authors say is not a security boundary, and there is no approval step, no taint handling, and no credential scoping. A prompt injection in a fetched web page can read API keys held by tools and send them out through the bundled web tool. The remote sandboxes (E2B, Modal, Blaxel, Docker) are a real improvement but must be opted into. Importing the library also silently loads a .env file, and the bundled web UI's default exposure is not locked down.

Key gaps (5)

  1. load_dotenv() runs at every `import smolagents`, so a .env in the project or working directory can silently redirect the model endpoint and credentials. C6 · Memory, context & configuration integrity
  2. The default LocalPythonExecutor runs model-written code in the developer's own process and is documented as not a security boundary; an escape has the OS user's full authority. C4 · Code-execution isolation
  3. No identity scoping: an executor escape acts with the OS user's full ambient authority and every credential on the host. C1 · Identity & least privilege
  4. A hijacked CodeAgent can read secrets held by tool objects and send them out via arbitrary URL fetch, with no human involved. C5 · Untrusted input blast radius
  5. Hub tools, once trust_remote_code=True is passed, are downloaded unpinned and exec'd inside the agent's process with all its credentials. C7 · Third-party extensions

Criteria

C1 Identity & least privilege

Minimal 0.00 / 1.00

smolagents has no identity or authorization layer of its own. The model API key (for example HF_TOKEN read from the environment) and any tool keys are long-lived credentials loaded from the environment, and tools build their own clients from them. The default code executor runs inside the developer's Python process, so anything that escapes its filters acts with the full authority of the OS user, including every credential on the machine. Nothing narrows authority per tool or per request.

C2 Approval gates

Minimal 0.00 / 1.00

There is no human approval gate anywhere in the framework. Model-written code is executed as soon as it is parsed, and ToolCallingAgent calls tools directly after a type check. Step callbacks run only after a step has executed, and final-answer checks only validate the final answer. Every consequential action a registered tool can take happens unattended.

C3 Tool & action scoping

Minimal 0.23 / 1.00

The default CodeAgent turns every action into model-written Python, and tools called from that code receive whatever arguments the code computes, with no schema check. ToolCallingAgent does check argument names and JSON types against each tool's schema, but there is no value-level validation (no path, URL, or host allowlists). The bundled VisitWebpageTool fetches any URL, including internal addresses. Tools are opt-in per agent (base tools are off by default), but code execution is always part of a CodeAgent.

C4 Code-execution isolation

Moderate 0.50 / 1.00

By default CodeAgent runs model-written code in LocalPythonExecutor, an in-process Python interpreter that restricts imports and dunder access. Its own docstring, the README, and SECURITY.md all say it is not a security sandbox, and an escape lands in the developer's process with all its credentials, files, and network. The framework also ships remote sandboxes (E2B, Modal, Blaxel, Docker); with E2B every code action and tool runs in a remote ephemeral VM, managed agents are refused rather than run locally, and setup failure raises an error instead of falling back. Those sandboxes are opt-in, so the stronger mechanism is capped.

C5 Untrusted input blast radius

Minimal 0.00 / 1.00

Tool results, web pages, and MCP outputs are fed back to the model as 'Observation' messages, which are converted into user-role messages, so untrusted content has the same standing as the user's instructions. Nothing structurally limits a hijacked agent: there is no taint tracking, no approval after untrusted content is read, and no egress restriction. A hijacked CodeAgent can read any secret reachable from its tools and send it out through VisitWebpageTool, while taking whatever irreversible actions its tools allow. The bundled web UI's default exposure is not locked down.

C6 Memory, context & configuration integrity

Minimal 0.00 / 1.00

Importing smolagents always imports remote_executors, which calls load_dotenv() at import time. That silently loads a .env file from the current directory (in notebooks and REPLs) or from the package's parent directories, so a .env in a cloned project can set variables like OPENAI_BASE_URL or HF_TOKEN that redirect where the model client sends the user's key and where its code comes from. Agent memory itself stays in the process and is never written to disk, but the bundled web UI's session isolation is not a complete boundary.

C7 Third-party extensions

Minimal 0.15 / 1.00

Loading a tool or agent from the Hugging Face Hub requires trust_remote_code=True, but once set, the tool's code is downloaded at whatever revision is latest (no pin is required and no hash is checked) and executed with exec() inside the agent's own process. ToolCollection.from_mcp also requires the flag, but MCPClient launches any given MCP server without one. The official docs and docstrings set trust_remote_code=True in routine examples. Model loaders default trust_remote_code to False and remote-executor variables default to safe JSON instead of pickle.

C8 Secrets & sensitive-data protection

Minimal 0.13 / 1.00

Credentials come from environment variables and are stored as plain attributes on model and tool objects. Because the default interpreter allows any non-dunder attribute access on objects it is given, model-written code can read keys held by tools (for example GoogleSearchTool.api_key) or by a managed agent's model client and print or send them. There is no redaction anywhere, and DEBUG verbosity prints full model input messages. On the plus side, the framework sends no telemetry, and the Docker sandbox receives only its own kernel token rather than the host environment.

C9 Audit & traceability

Minimal 0.30 / 1.00

Every step is recorded in a structured in-memory log (the code action, observations, errors, token usage, and timing) and printed to the console at the default verbosity. But the record lives only in the agent's Python process: nothing is written to disk or shipped elsewhere by default, so it is lost when the process ends or crashes. Tool calls made inside a code action are not recorded individually, sub-agent steps live only in the sub-agent's memory, and there is no actor attribution. OpenTelemetry tracing is documented through an external package, not this code.

C10 Limits & kill switch

Minimal 0.33 / 1.00

Agents stop after 20 steps by default, the local interpreter caps each code action at 10 million operations, and a 30-second execution timeout is configured. That timeout is weak: it uses a thread pool inside a with-block, and the code's own docstring says the thread cannot be killed, so long-running code (for example time.sleep from the allowed time module) keeps running. There is no wall-clock, token, or cost cap per run. Managed sub-agents start a fresh step budget on every call, and in a CodeAgent the model can pass max_steps directly when calling a sub-agent. interrupt() only sets a flag that is checked between steps.