# Defense-in-Depth Score: smolagents

**Repo:** https://github.com/huggingface/smolagents · **Commit:** `c30b115286e000e98711fae5e85993547b73d826` · **Reviewed:** 2026-10-03
**What it is:** Barebones library for code-writing agents
**Category:** Agent Frameworks
**Scored configuration:** CodeAgent with default constructor arguments (executor_type="local" LocalPythonExecutor, max_steps=20, add_base_tools=False) and the README's WebSearchTool example; GradioUI and remote sandboxes assessed at their own defaults.
**Agent surface (default):** code execution yes · filesystem write opt-in · network egress yes · external credentials yes · persistent memory no · untrusted input yes · third party extensions opt-in · sub agents opt-in · external communication opt-in

## Score: 1.6 / 10.0 (Minimal)

| # | Criterion | S | C | D | B | Raw | Cap | Score | Confidence |
|---|---|---|---|---|---|---|---|---|---|
| C1 | Identity & least privilege | L0 | L0 | L0 | L0 | 0.00 | — | **0.00** | High |
| C2 | Approval gates | L0 | L0 | L0 | L0 | 0.00 | — | **0.00** | High |
| C3 | Tool & action scoping | L1 | L0 | L2 | L1 | 0.23 | — | **0.23** | High |
| C4 | Code-execution isolation | L4 | L3 | L0 | L2 | 0.62 | G1 | **0.50** (alt) | High |
| C5 | Untrusted input blast radius | L0 | L0 | L0 | L0 | 0.00 | C5-PUBLICTRIGGER | **0.00** | High |
| C6 | Memory, context & configuration integrity | L0 | L0 | L0 | L0 | 0.00 | C6-REPOCONFIG | **0.00** | High |
| C7 | Third-party extensions | L1 | L1 | L0 | L0 | 0.15 | — | **0.15** | High |
| C8 | Secrets & sensitive-data protection | L1 | L0 | L1 | L0 | 0.12 | — | **0.12** | High |
| C9 | Audit & traceability | L1 | L1 | L2 | L1 | 0.30 | — | **0.30** | High |
| C10 | Limits & kill switch | L2 | L1 | L1 | L1 | 0.33 | — | **0.33** | High |


As shipped, smolagents runs model-written Python inside your own process with an interpreter its authors say is not a security boundary, and there is no approval step, no taint handling, and no credential scoping. A prompt injection in a fetched web page can read API keys held by tools and send them out through the bundled web tool. The remote sandboxes (E2B, Modal, Blaxel, Docker) are a real improvement but must be opted into. Importing the library also silently loads a .env file, and the bundled web UI's default exposure is not locked down.

## Critical gaps
- load_dotenv() runs at every `import smolagents`, so a .env in the project or working directory can silently redirect the model endpoint and credentials. (ASI06, ASI04, T1; C6) — [src/smolagents/remote_executors.py:46-48](https://github.com/huggingface/smolagents/blob/c30b115286e000e98711fae5e85993547b73d826/src/smolagents/remote_executors.py#L46-L48); [src/smolagents/agents.py:80](https://github.com/huggingface/smolagents/blob/c30b115286e000e98711fae5e85993547b73d826/src/smolagents/agents.py#L80); [src/smolagents/models.py:1686](https://github.com/huggingface/smolagents/blob/c30b115286e000e98711fae5e85993547b73d826/src/smolagents/models.py#L1686)
- The default LocalPythonExecutor runs model-written code in the developer's own process and is documented as not a security boundary; an escape has the OS user's full authority. (ASI05, T11, LLM05; C4) — [src/smolagents/agents.py:1535](https://github.com/huggingface/smolagents/blob/c30b115286e000e98711fae5e85993547b73d826/src/smolagents/agents.py#L1535); [src/smolagents/local_python_executor.py:1693](https://github.com/huggingface/smolagents/blob/c30b115286e000e98711fae5e85993547b73d826/src/smolagents/local_python_executor.py#L1693)
- No identity scoping: an executor escape acts with the OS user's full ambient authority and every credential on the host. (ASI03, T3; C1) — [src/smolagents/local_python_executor.py:1693](https://github.com/huggingface/smolagents/blob/c30b115286e000e98711fae5e85993547b73d826/src/smolagents/local_python_executor.py#L1693); [src/smolagents/models.py:1535-1540](https://github.com/huggingface/smolagents/blob/c30b115286e000e98711fae5e85993547b73d826/src/smolagents/models.py#L1535-L1540)
- A hijacked CodeAgent can read secrets held by tool objects and send them out via arbitrary URL fetch, with no human involved. (ASI01, LLM01, T6; C5) — [src/smolagents/local_python_executor.py:390-393](https://github.com/huggingface/smolagents/blob/c30b115286e000e98711fae5e85993547b73d826/src/smolagents/local_python_executor.py#L390-L393); [src/smolagents/default_tools.py:528](https://github.com/huggingface/smolagents/blob/c30b115286e000e98711fae5e85993547b73d826/src/smolagents/default_tools.py#L528)
- Hub tools, once trust_remote_code=True is passed, are downloaded unpinned and exec'd inside the agent's process with all its credentials. (ASI04, T17, LLM03; C7) — [src/smolagents/tools.py:563](https://github.com/huggingface/smolagents/blob/c30b115286e000e98711fae5e85993547b73d826/src/smolagents/tools.py#L563); [src/smolagents/tools.py:575](https://github.com/huggingface/smolagents/blob/c30b115286e000e98711fae5e85993547b73d826/src/smolagents/tools.py#L575)

## Criterion details

### C1 Identity & least privilege — 0.00 (high)

smolagents has no identity or authorization layer of its own. The model API key (for example HF_TOKEN read from the environment) and any tool keys are long-lived credentials loaded from the environment, and tools build their own clients from them. The default code executor runs inside the developer's Python process, so anything that escapes its filters acts with the full authority of the OS user, including every credential on the machine. Nothing narrows authority per tool or per request.

- **S L0:** Credentials are ambient: the model client takes HF_TOKEN (or a passed key) and tools such as GoogleSearchTool read their own keys from the environment; there is no scoped or per-tool identity. — [src/smolagents/models.py:1535-1540](https://github.com/huggingface/smolagents/blob/c30b115286e000e98711fae5e85993547b73d826/src/smolagents/models.py#L1535-L1540); [src/smolagents/default_tools.py:186](https://github.com/huggingface/smolagents/blob/c30b115286e000e98711fae5e85993547b73d826/src/smolagents/default_tools.py#L186) (verified)
  - *To reach the next level:* No per-tool or role-scoped credential; every tool and the model share the process's ambient environment.
- **C L0:** Tools construct their own privileged clients from environment variables and there is no authorization check on any tool path. — [src/smolagents/default_tools.py:186](https://github.com/huggingface/smolagents/blob/c30b115286e000e98711fae5e85993547b73d826/src/smolagents/default_tools.py#L186); searched `rg -n -S 'approv|confirm|human_in|ask_user|require_'` in `src/` → 4 hits (2 hits are few-shot prompt text using the word 'confirm', 2 are the CLI's rich Confirm import and its 'Configure advanced options?' prompt; none gates a tool call) (verified)
  - *To reach the next level:* No central authorization layer that every tool call passes through.
- **D L0:** The default local executor runs in the developer's process with the OS user's full authority; least privilege requires the developer to pick a remote sandbox and scope keys themselves. — [src/smolagents/agents.py:1535](https://github.com/huggingface/smolagents/blob/c30b115286e000e98711fae5e85993547b73d826/src/smolagents/agents.py#L1535); [src/smolagents/local_python_executor.py:1693](https://github.com/huggingface/smolagents/blob/c30b115286e000e98711fae5e85993547b73d826/src/smolagents/local_python_executor.py#L1693) (verified)
  - *To reach the next level:* No narrower default identity; the in-process executor inherits everything the launching user holds.
- **B L0:** If the in-process executor's filters are bypassed (its authors state it is not a security boundary), code runs as the OS user with access to all local credentials, files, and network. — [src/smolagents/local_python_executor.py:1693](https://github.com/huggingface/smolagents/blob/c30b115286e000e98711fae5e85993547b73d826/src/smolagents/local_python_executor.py#L1693); [README.md:245](https://github.com/huggingface/smolagents/blob/c30b115286e000e98711fae5e85993547b73d826/README.md#L245); [src/smolagents/agents.py:1726](https://github.com/huggingface/smolagents/blob/c30b115286e000e98711fae5e85993547b73d826/src/smolagents/agents.py#L1726) (verified)
  - *To reach the next level:* Default execution would need to run outside the user's process with no ambient credentials to limit blast radius.
- **Cap:** none

### C2 Approval gates — 0.00 (high)

There is no human approval gate anywhere in the framework. Model-written code is executed as soon as it is parsed, and ToolCallingAgent calls tools directly after a type check. Step callbacks run only after a step has executed, and final-answer checks only validate the final answer. Every consequential action a registered tool can take happens unattended.

- **S L0:** No approval mechanism exists; parsed code goes straight to the executor and step callbacks fire after execution. — [src/smolagents/agents.py:1726](https://github.com/huggingface/smolagents/blob/c30b115286e000e98711fae5e85993547b73d826/src/smolagents/agents.py#L1726); [src/smolagents/agents.py:623](https://github.com/huggingface/smolagents/blob/c30b115286e000e98711fae5e85993547b73d826/src/smolagents/agents.py#L623); searched `rg -n -S 'approv|confirm|human_in|ask_user|require_'` in `src/` → 4 hits (2 hits are few-shot prompt text using the word 'confirm', 2 are the CLI's rich Confirm import and its 'Configure advanced options?' prompt; none gates a tool call) (verified)
  - *To reach the next level:* No per-call human approval showing the exact code or tool arguments.
- **C L0:** The most powerful path, code execution, is ungated, as is every tool call. — [src/smolagents/agents.py:1726](https://github.com/huggingface/smolagents/blob/c30b115286e000e98711fae5e85993547b73d826/src/smolagents/agents.py#L1726); [src/smolagents/agents.py:1476](https://github.com/huggingface/smolagents/blob/c30b115286e000e98711fae5e85993547b73d826/src/smolagents/agents.py#L1476) (verified)
  - *To reach the next level:* No gate on code execution or tool calls.
- **D L0:** No approval exists to be on by default. — searched `rg -n -S 'approv|confirm|human_in|ask_user|require_'` in `src/` → 4 hits (2 hits are few-shot prompt text using the word 'confirm', 2 are the CLI's rich Confirm import and its 'Configure advanced options?' prompt; none gates a tool call) (verified)
  - *To reach the next level:* No default-on approval.
- **B L0:** Anything a registered tool or an executor escape can do (file deletion, outbound requests, sending data) is irreversible and there are no checkpoints or rollback. — [src/smolagents/agents.py:1726](https://github.com/huggingface/smolagents/blob/c30b115286e000e98711fae5e85993547b73d826/src/smolagents/agents.py#L1726); [src/smolagents/default_tools.py:528](https://github.com/huggingface/smolagents/blob/c30b115286e000e98711fae5e85993547b73d826/src/smolagents/default_tools.py#L528) (verified)
  - *To reach the next level:* No checkpoints, rollback, dry-run, or bound on consequential actions.
- **Cap:** none

### C3 Tool & action scoping — 0.23 (high)

The default CodeAgent turns every action into model-written Python, and tools called from that code receive whatever arguments the code computes, with no schema check. ToolCallingAgent does check argument names and JSON types against each tool's schema, but there is no value-level validation (no path, URL, or host allowlists). The bundled VisitWebpageTool fetches any URL, including internal addresses. Tools are opt-in per agent (base tools are off by default), but code execution is always part of a CodeAgent.

- **S L1:** The only framework-level validation is a type check of argument names and JSON types (ToolCallingAgent); there are no value allowlists, and the bundled web tool takes a raw URL. — [src/smolagents/tools.py:1387-1402](https://github.com/huggingface/smolagents/blob/c30b115286e000e98711fae5e85993547b73d826/src/smolagents/tools.py#L1387-L1402); [src/smolagents/default_tools.py:528](https://github.com/huggingface/smolagents/blob/c30b115286e000e98711fae5e85993547b73d826/src/smolagents/default_tools.py#L528); searched `rg -n 'localhost|169\.254|private_ip|is_private|allowed_hosts|ipaddress'` in `src/` → 1 hits (single hit is a docstring example URL in mcp_client.py; no internal-address blocking exists) (verified)
  - *To reach the next level:* No allowlist validation in code (resolved paths, host/URL allowlists blocking internal addresses, numeric bounds).
- **C L0:** In the default CodeAgent, tools are invoked directly from interpreted code, which bypasses validate_tool_arguments entirely. — [src/smolagents/local_python_executor.py:918](https://github.com/huggingface/smolagents/blob/c30b115286e000e98711fae5e85993547b73d826/src/smolagents/local_python_executor.py#L918); [src/smolagents/agents.py:492](https://github.com/huggingface/smolagents/blob/c30b115286e000e98711fae5e85993547b73d826/src/smolagents/agents.py#L492); [src/smolagents/agents.py:1476](https://github.com/huggingface/smolagents/blob/c30b115286e000e98711fae5e85993547b73d826/src/smolagents/agents.py#L1476) (verified)
  - *To reach the next level:* CodeAgent tool calls do not pass through any shared validation layer.
- **D L2:** Tools are explicitly chosen per agent and add_base_tools defaults to False, but a CodeAgent always includes code execution and the README example adds web search. — [src/smolagents/agents.py:301](https://github.com/huggingface/smolagents/blob/c30b115286e000e98711fae5e85993547b73d826/src/smolagents/agents.py#L301); [src/smolagents/agents.py:1535](https://github.com/huggingface/smolagents/blob/c30b115286e000e98711fae5e85993547b73d826/src/smolagents/agents.py#L1535) (verified)
  - *To reach the next level:* No read-only default; code execution is always enabled for CodeAgent.
- **B L1:** A misused tool has broad reach: VisitWebpageTool can reach any host (including internal services), and the code tool is limited only by an import allowlist. — [src/smolagents/default_tools.py:528](https://github.com/huggingface/smolagents/blob/c30b115286e000e98711fae5e85993547b73d826/src/smolagents/default_tools.py#L528); [src/smolagents/utils.py:59](https://github.com/huggingface/smolagents/blob/c30b115286e000e98711fae5e85993547b73d826/src/smolagents/utils.py#L59) (verified)
  - *To reach the next level:* No workspace or host scoping and no quantity bounds on tool use.
- **Cap:** none

### C4 Code-execution isolation — 0.50 (high)

By default CodeAgent runs model-written code in LocalPythonExecutor, an in-process Python interpreter that restricts imports and dunder access. Its own docstring, the README, and SECURITY.md all say it is not a security sandbox, and an escape lands in the developer's process with all its credentials, files, and network. The framework also ships remote sandboxes (E2B, Modal, Blaxel, Docker); with E2B every code action and tool runs in a remote ephemeral VM, managed agents are refused rather than run locally, and setup failure raises an error instead of falling back. Those sandboxes are opt-in, so the stronger mechanism is capped.

- **default configuration** (default; raw 0.25 → 0.25)
  - **S L1:** Isolation is AST-level filtering inside the host Python process (import allowlist, dunder blocking), which its authors document as not a security boundary. — [src/smolagents/local_python_executor.py:1693](https://github.com/huggingface/smolagents/blob/c30b115286e000e98711fae5e85993547b73d826/src/smolagents/local_python_executor.py#L1693); [README.md:245](https://github.com/huggingface/smolagents/blob/c30b115286e000e98711fae5e85993547b73d826/README.md#L245); [src/smolagents/local_python_executor.py:390-393](https://github.com/huggingface/smolagents/blob/c30b115286e000e98711fae5e85993547b73d826/src/smolagents/local_python_executor.py#L390-L393) (verified)
    - *To reach the next level:* No OS-level or kernel-level separation for the default executor.
  - **C L1:** All model-written code goes through the filtering interpreter, but registered tools, MCP stdio servers, and anything they spawn run unfiltered in or alongside the host process. — [src/smolagents/agents.py:1726](https://github.com/huggingface/smolagents/blob/c30b115286e000e98711fae5e85993547b73d826/src/smolagents/agents.py#L1726); [src/smolagents/local_python_executor.py:918](https://github.com/huggingface/smolagents/blob/c30b115286e000e98711fae5e85993547b73d826/src/smolagents/local_python_executor.py#L918); [src/smolagents/mcp_client.py:118](https://github.com/huggingface/smolagents/blob/c30b115286e000e98711fae5e85993547b73d826/src/smolagents/mcp_client.py#L118) (verified)
    - *To reach the next level:* Tool and extension code paths are not covered by the executor's restrictions.
  - **D L2:** The restricted interpreter is on by default, but passing '*' in additional_authorized_imports removes the import allowlist with only an INFO-level note. — [src/smolagents/agents.py:1578](https://github.com/huggingface/smolagents/blob/c30b115286e000e98711fae5e85993547b73d826/src/smolagents/agents.py#L1578); [src/smolagents/local_python_executor.py:375](https://github.com/huggingface/smolagents/blob/c30b115286e000e98711fae5e85993547b73d826/src/smolagents/local_python_executor.py#L375) (verified)
    - *To reach the next level:* Disabling the import allowlist is a plain config value with only an INFO-level log, not a loudly named operator flag.
  - **B L0:** An escape lands in the agent's own process, which holds the model API key, tool keys, the user's filesystem, and unrestricted network. — [src/smolagents/models.py:1535-1540](https://github.com/huggingface/smolagents/blob/c30b115286e000e98711fae5e85993547b73d826/src/smolagents/models.py#L1535-L1540); [src/smolagents/default_tools.py:186](https://github.com/huggingface/smolagents/blob/c30b115286e000e98711fae5e85993547b73d826/src/smolagents/default_tools.py#L186); [src/smolagents/local_python_executor.py:1693](https://github.com/huggingface/smolagents/blob/c30b115286e000e98711fae5e85993547b73d826/src/smolagents/local_python_executor.py#L1693) (verified)
    - *To reach the next level:* The default executor would need to run with no secrets, no host filesystem, and restricted egress.
- **opt-in E2B remote sandbox (executor_type="e2b")** (alt; raw 0.62, cap G1 → 0.50) ← counted
  - **S L4:** E2BExecutor runs code in a remote ephemeral sandbox service created per executor and killed on cleanup. — [src/smolagents/remote_executors.py:368](https://github.com/huggingface/smolagents/blob/c30b115286e000e98711fae5e85993547b73d826/src/smolagents/remote_executors.py#L368); [src/smolagents/remote_executors.py:443](https://github.com/huggingface/smolagents/blob/c30b115286e000e98711fae5e85993547b73d826/src/smolagents/remote_executors.py#L443) (verified)
  - **C L3:** With a remote executor every code action runs in the sandbox, tools are re-instantiated from their source inside it, managed agents are refused rather than run locally, and setup errors propagate instead of falling back to host execution. — [src/smolagents/agents.py:492](https://github.com/huggingface/smolagents/blob/c30b115286e000e98711fae5e85993547b73d826/src/smolagents/agents.py#L492); [src/smolagents/tools.py:1339-1341](https://github.com/huggingface/smolagents/blob/c30b115286e000e98711fae5e85993547b73d826/src/smolagents/tools.py#L1339-L1341); [src/smolagents/agents.py:1608-1609](https://github.com/huggingface/smolagents/blob/c30b115286e000e98711fae5e85993547b73d826/src/smolagents/agents.py#L1608-L1609) (verified)
    - *To reach the next level:* MCP stdio servers that the developer starts with MCPClient still run unsandboxed on the host.
  - **D L0:** Remote sandboxes are opt-in: executor_type defaults to "local". — [src/smolagents/agents.py:1535](https://github.com/huggingface/smolagents/blob/c30b115286e000e98711fae5e85993547b73d826/src/smolagents/agents.py#L1535) (verified)
    - *To reach the next level:* The sandbox is not the default executor.
  - **B L2:** The sandbox is ephemeral and receives no host environment, but it keeps the provider's default unrestricted egress and no egress allowlist is configured. — [src/smolagents/remote_executors.py:368](https://github.com/huggingface/smolagents/blob/c30b115286e000e98711fae5e85993547b73d826/src/smolagents/remote_executors.py#L368); [src/smolagents/remote_executors.py:649](https://github.com/huggingface/smolagents/blob/c30b115286e000e98711fae5e85993547b73d826/src/smolagents/remote_executors.py#L649) (verified)
    - *To reach the next level:* No network egress restriction or allowlist configured for the sandbox.
- **Cap:** G1 — Opt-in mechanism: off in the scored default configuration.
- **Notes:** Docker executor is a stock container (no user, cap-drop, or seccomp options set by default); E2B was scored as the strongest opt-in mechanism.

### C5 Untrusted input blast radius — 0.00 (high)

Tool results, web pages, and MCP outputs are fed back to the model as 'Observation' messages, which are converted into user-role messages, so untrusted content has the same standing as the user's instructions. Nothing structurally limits a hijacked agent: there is no taint tracking, no approval after untrusted content is read, and no egress restriction. A hijacked CodeAgent can read any secret reachable from its tools and send it out through VisitWebpageTool, while taking whatever irreversible actions its tools allow. The bundled web UI's default exposure is not locked down.

- **S L0:** No structural limit; nothing changes after untrusted content enters context. — searched `rg -n -i 'untrusted|injection|spotlight|quarantin'` in `src/` → 6 hits (hits are pickle-deserialization warnings in serialization.py and the LocalPythonExecutor docstring; nothing marks or isolates untrusted content) (verified)
  - *To reach the next level:* No rule that disables or gates egress and state-changing tools once untrusted content has been read.
- **C L0:** Tool observations are wrapped as tool-response messages that are converted to the user role, with nothing distinguishing untrusted sources. — [src/smolagents/memory.py:129](https://github.com/huggingface/smolagents/blob/c30b115286e000e98711fae5e85993547b73d826/src/smolagents/memory.py#L129); [src/smolagents/models.py:284](https://github.com/huggingface/smolagents/blob/c30b115286e000e98711fae5e85993547b73d826/src/smolagents/models.py#L284) (verified)
  - *To reach the next level:* No source of untrusted content is distinguished from principal instructions.
- **D L0:** There is no control to be on by default. — searched `rg -n -i 'untrusted|injection|spotlight|quarantin'` in `src/` → 6 hits (hits are pickle-deserialization warnings in serialization.py and the LocalPythonExecutor docstring; nothing marks or isolates untrusted content) (verified)
  - *To reach the next level:* No default-on containment.
- **B L0:** A hijacked agent can read secrets reachable from tool objects and exfiltrate them via arbitrary URL fetch, and take irreversible actions, with no human involved. — [src/smolagents/default_tools.py:528](https://github.com/huggingface/smolagents/blob/c30b115286e000e98711fae5e85993547b73d826/src/smolagents/default_tools.py#L528); [src/smolagents/local_python_executor.py:390-393](https://github.com/huggingface/smolagents/blob/c30b115286e000e98711fae5e85993547b73d826/src/smolagents/local_python_executor.py#L390-L393); [src/smolagents/default_tools.py:186](https://github.com/huggingface/smolagents/blob/c30b115286e000e98711fae5e85993547b73d826/src/smolagents/default_tools.py#L186) (verified)
  - *To reach the next level:* Exfiltration and irreversible actions would both need human approval.
- **Cap:** C5-PUBLICTRIGGER — The bundled web UI's default configuration lets parties other than the principal trigger the agent.

### C6 Memory, context & configuration integrity — 0.00 (high)

Importing smolagents always imports remote_executors, which calls load_dotenv() at import time. That silently loads a .env file from the current directory (in notebooks and REPLs) or from the package's parent directories, so a .env in a cloned project can set variables like OPENAI_BASE_URL or HF_TOKEN that redirect where the model client sends the user's key and where its code comes from. Agent memory itself stays in the process and is never written to disk, but the bundled web UI's session isolation is not a complete boundary.

- **S L0:** A workspace .env is loaded silently at import, and it can set endpoint and credential variables that the model clients fall back to (for example, OpenAIModel passes base_url=None so the OpenAI SDK reads OPENAI_BASE_URL). — [src/smolagents/remote_executors.py:46-48](https://github.com/huggingface/smolagents/blob/c30b115286e000e98711fae5e85993547b73d826/src/smolagents/remote_executors.py#L46-L48); [src/smolagents/agents.py:80](https://github.com/huggingface/smolagents/blob/c30b115286e000e98711fae5e85993547b73d826/src/smolagents/agents.py#L80); [src/smolagents/models.py:1686](https://github.com/huggingface/smolagents/blob/c30b115286e000e98711fae5e85993547b73d826/src/smolagents/models.py#L1686) (verified)
  - *To reach the next level:* No workspace-trust decision before loading environment and endpoint settings from project files.
- **C L0:** Neither the .env path nor the session memory has any validation or provenance. — [src/smolagents/remote_executors.py:46-48](https://github.com/huggingface/smolagents/blob/c30b115286e000e98711fae5e85993547b73d826/src/smolagents/remote_executors.py#L46-L48); [src/smolagents/memory.py:129](https://github.com/huggingface/smolagents/blob/c30b115286e000e98711fae5e85993547b73d826/src/smolagents/memory.py#L129) (verified)
  - *To reach the next level:* No control on any memory or configuration path.
- **D L0:** Session isolation in the bundled web UI is not a complete boundary. (verified)
  - *To reach the next level:* Per-session isolation in the bundled UI needs hardening.
- **B L0:** Injected context can persist beyond a single session in the bundled UI, and a planted .env affects every session started from that directory. — [src/smolagents/remote_executors.py:46-48](https://github.com/huggingface/smolagents/blob/c30b115286e000e98711fae5e85993547b73d826/src/smolagents/remote_executors.py#L46-L48); [src/smolagents/agents.py:1726](https://github.com/huggingface/smolagents/blob/c30b115286e000e98711fae5e85993547b73d826/src/smolagents/agents.py#L1726) (verified)
  - *To reach the next level:* Poisoned context would need to be session-scoped or easily purged, and project config would need user review before use.
- **Cap:** C6-REPOCONFIG — load_dotenv() runs unconditionally at `import smolagents`, so a .env in the project or working directory can, with no trust decision, set the model API base URL and credentials (python-dotenv's find_dotenv searches cwd in interactive sessions and the caller file's parent directories otherwise; the OpenAI SDK reads OPENAI_BASE_URL when base_url is None).
- **Notes:** The trigger (load_dotenv at import, verified) is in smolagents; the search path and the SDK's environment fallback are library behaviour, inferred from python-dotenv's find_dotenv and the OpenAI SDK.

### C7 Third-party extensions — 0.15 (high)

Loading a tool or agent from the Hugging Face Hub requires trust_remote_code=True, but once set, the tool's code is downloaded at whatever revision is latest (no pin is required and no hash is checked) and executed with exec() inside the agent's own process. ToolCollection.from_mcp also requires the flag, but MCPClient launches any given MCP server without one. The official docs and docstrings set trust_remote_code=True in routine examples. Model loaders default trust_remote_code to False and remote-executor variables default to safe JSON instead of pickle.

- **S L1:** Sources are chosen by the developer but unpinned: revision is optional and nothing verifies a hash or signature. — [src/smolagents/tools.py:563](https://github.com/huggingface/smolagents/blob/c30b115286e000e98711fae5e85993547b73d826/src/smolagents/tools.py#L563); searched `rg -n 'sha256|hashlib|verify'` in `src/` → 0 hits (verified)
  - *To reach the next level:* No pinning requirement or integrity check on Hub tools/agents or MCP servers.
- **C L1:** Only the Hub loaders and ToolCollection.from_mcp require an acknowledgement; MCPClient connects or launches servers with no check, and nothing re-verifies tool definitions later. — [src/smolagents/tools.py:549](https://github.com/huggingface/smolagents/blob/c30b115286e000e98711fae5e85993547b73d826/src/smolagents/tools.py#L549); [src/smolagents/tools.py:1052](https://github.com/huggingface/smolagents/blob/c30b115286e000e98711fae5e85993547b73d826/src/smolagents/tools.py#L1052); [src/smolagents/mcp_client.py:118](https://github.com/huggingface/smolagents/blob/c30b115286e000e98711fae5e85993547b73d826/src/smolagents/mcp_client.py#L118) (verified)
  - *To reach the next level:* Verification does not cover MCP servers or later changes to extension code.
- **D L0:** Consent is a generic trust_remote_code=True flag that shows nothing about what will run, and official examples set it routinely (lowered one level per the framework rule). — [src/smolagents/tools.py:549](https://github.com/huggingface/smolagents/blob/c30b115286e000e98711fae5e85993547b73d826/src/smolagents/tools.py#L549); [src/smolagents/agents.py:1097](https://github.com/huggingface/smolagents/blob/c30b115286e000e98711fae5e85993547b73d826/src/smolagents/agents.py#L1097); [docs/source/en/guided_tour.md:617](https://github.com/huggingface/smolagents/blob/c30b115286e000e98711fae5e85993547b73d826/docs/source/en/guided_tour.md#L617) (verified)
  - *To reach the next level:* Consent should show the exact package/revision and permissions, and examples should not normalise the flag.
- **B L0:** Hub tool code is exec'd in the agent's own process with all of its credentials and environment. — [src/smolagents/tools.py:575](https://github.com/huggingface/smolagents/blob/c30b115286e000e98711fae5e85993547b73d826/src/smolagents/tools.py#L575) (verified)
  - *To reach the next level:* Extensions would need to run in a separate process with a scrubbed environment.
- **Cap:** none
- **Notes:** trust_remote_code is required before any Hub code executes, so C7-RCELOAD does not apply.

### C8 Secrets & sensitive-data protection — 0.12 (high)

Credentials come from environment variables and are stored as plain attributes on model and tool objects. Because the default interpreter allows any non-dunder attribute access on objects it is given, model-written code can read keys held by tools (for example GoogleSearchTool.api_key) or by a managed agent's model client and print or send them. There is no redaction anywhere, and DEBUG verbosity prints full model input messages. On the plus side, the framework sends no telemetry, and the Docker sandbox receives only its own kernel token rather than the host environment.

- **S L1:** Secrets are read from environment variables with no masking or redaction on any path. — [src/smolagents/models.py:1535-1540](https://github.com/huggingface/smolagents/blob/c30b115286e000e98711fae5e85993547b73d826/src/smolagents/models.py#L1535-L1540); [src/smolagents/default_tools.py:186](https://github.com/huggingface/smolagents/blob/c30b115286e000e98711fae5e85993547b73d826/src/smolagents/default_tools.py#L186); searched `rg -n -i 'redact|mask|SecretStr'` in `src/` → 2 hits (both hits are a local variable named redacted_version in GoogleSearchTool result formatting, not secret redaction) (verified)
  - *To reach the next level:* No type-level masking or log/model-bound redaction.
- **C L0:** No path (logs, model-bound messages, tool objects reachable from code, error messages) is protected. — [src/smolagents/local_python_executor.py:390-393](https://github.com/huggingface/smolagents/blob/c30b115286e000e98711fae5e85993547b73d826/src/smolagents/local_python_executor.py#L390-L393); [src/smolagents/monitoring.py:220](https://github.com/huggingface/smolagents/blob/c30b115286e000e98711fae5e85993547b73d826/src/smolagents/monitoring.py#L220); searched `rg -n -i 'redact|mask|SecretStr'` in `src/` → 2 hits (both hits are a local variable named redacted_version in GoogleSearchTool result formatting, not secret redaction) (verified)
  - *To reach the next level:* No protected path.
- **D L1:** No telemetry is sent, but verbose DEBUG logging of full messages is one argument away and unredacted. — searched `rg -n -i 'sentry|posthog|telemetry'` in `src/` → 0 hits; [src/smolagents/monitoring.py:220](https://github.com/huggingface/smolagents/blob/c30b115286e000e98711fae5e85993547b73d826/src/smolagents/monitoring.py#L220) (verified)
  - *To reach the next level:* Redaction does not exist, so it cannot be on by default.
- **B L0:** Long-lived provider keys held by tool and model objects are reachable from model-written code through attribute access, and the whole process environment is reachable after an executor escape. — [src/smolagents/local_python_executor.py:390-393](https://github.com/huggingface/smolagents/blob/c30b115286e000e98711fae5e85993547b73d826/src/smolagents/local_python_executor.py#L390-L393); [src/smolagents/default_tools.py:186](https://github.com/huggingface/smolagents/blob/c30b115286e000e98711fae5e85993547b73d826/src/smolagents/default_tools.py#L186); [src/smolagents/models.py:1535-1540](https://github.com/huggingface/smolagents/blob/c30b115286e000e98711fae5e85993547b73d826/src/smolagents/models.py#L1535-L1540) (verified)
  - *To reach the next level:* Keys reachable from the model would need to be scoped and short-lived, or kept out of objects the interpreter can reach.
- **Cap:** none

### C9 Audit & traceability — 0.30 (high)

Every step is recorded in a structured in-memory log (the code action, observations, errors, token usage, and timing) and printed to the console at the default verbosity. But the record lives only in the agent's Python process: nothing is written to disk or shipped elsewhere by default, so it is lost when the process ends or crashes. Tool calls made inside a code action are not recorded individually, sub-agent steps live only in the sub-agent's memory, and there is no actor attribution. OpenTelemetry tracing is documented through an external package, not this code.

- **S L1:** ActionStep keeps a structured per-step record (code action, observations, errors, timing), but the arguments of tool calls made inside a code action are never recorded, only the code that produced them. — [src/smolagents/memory.py:53-60](https://github.com/huggingface/smolagents/blob/c30b115286e000e98711fae5e85993547b73d826/src/smolagents/memory.py#L53-L60); [src/smolagents/agents.py:602](https://github.com/huggingface/smolagents/blob/c30b115286e000e98711fae5e85993547b73d826/src/smolagents/agents.py#L602); [src/smolagents/local_python_executor.py:918](https://github.com/huggingface/smolagents/blob/c30b115286e000e98711fae5e85993547b73d826/src/smolagents/local_python_executor.py#L918) (verified)
  - *To reach the next level:* No structured record of every individual tool call with its evaluated arguments and result status.
- **C L1:** Only the top-level step is recorded; individual tool calls made inside a code action and sub-agent trajectories are not in the parent's record. — [src/smolagents/agents.py:1726](https://github.com/huggingface/smolagents/blob/c30b115286e000e98711fae5e85993547b73d826/src/smolagents/agents.py#L1726); [src/smolagents/agents.py:876](https://github.com/huggingface/smolagents/blob/c30b115286e000e98711fae5e85993547b73d826/src/smolagents/agents.py#L876) (verified)
  - *To reach the next level:* Tool calls within code actions and sub-agent steps are not captured in the parent record.
- **D L2:** Recording is on by default but held in the agent's own process memory, cleared on each run by reset, and alterable by that process. — [src/smolagents/memory.py:234](https://github.com/huggingface/smolagents/blob/c30b115286e000e98711fae5e85993547b73d826/src/smolagents/memory.py#L234); [src/smolagents/monitoring.py:146-147](https://github.com/huggingface/smolagents/blob/c30b115286e000e98711fae5e85993547b73d826/src/smolagents/monitoring.py#L146-L147) (verified)
  - *To reach the next level:* Records are not written by a component the model's process cannot control.
- **B L1:** Console output is best-effort and the structured record is never persisted, so it is lost on crash or exit. — [src/smolagents/monitoring.py:146-147](https://github.com/huggingface/smolagents/blob/c30b115286e000e98711fae5e85993547b73d826/src/smolagents/monitoring.py#L146-L147); searched `rg -n -i 'opentelemetry|otel|audit'` in `src/` → 0 hits (verified)
  - *To reach the next level:* No durable per-action record.
- **Cap:** none

### C10 Limits & kill switch — 0.33 (high)

Agents stop after 20 steps by default, the local interpreter caps each code action at 10 million operations, and a 30-second execution timeout is configured. That timeout is weak: it uses a thread pool inside a with-block, and the code's own docstring says the thread cannot be killed, so long-running code (for example time.sleep from the allowed time module) keeps running. There is no wall-clock, token, or cost cap per run. Managed sub-agents start a fresh step budget on every call, and in a CodeAgent the model can pass max_steps directly when calling a sub-agent. interrupt() only sets a flag that is checked between steps.

- **S L2:** An iteration cap (max_steps) plus a per-code-action operation cap and execution timeout are enforced in code. — [src/smolagents/agents.py:545](https://github.com/huggingface/smolagents/blob/c30b115286e000e98711fae5e85993547b73d826/src/smolagents/agents.py#L545); [src/smolagents/local_python_executor.py:1444](https://github.com/huggingface/smolagents/blob/c30b115286e000e98711fae5e85993547b73d826/src/smolagents/local_python_executor.py#L1444); [src/smolagents/local_python_executor.py:60](https://github.com/huggingface/smolagents/blob/c30b115286e000e98711fae5e85993547b73d826/src/smolagents/local_python_executor.py#L60) (verified)
  - *To reach the next level:* No wall-clock or token/cost cap per run, and no rate limits on side-effecting tools.
- **C L1:** Limits apply to the top-level loop; a timed-out code action keeps running, and each sub-agent call starts a new budget. — [src/smolagents/local_python_executor.py:308-311](https://github.com/huggingface/smolagents/blob/c30b115286e000e98711fae5e85993547b73d826/src/smolagents/local_python_executor.py#L308-L311); [src/smolagents/local_python_executor.py:299](https://github.com/huggingface/smolagents/blob/c30b115286e000e98711fae5e85993547b73d826/src/smolagents/local_python_executor.py#L299); [src/smolagents/agents.py:876](https://github.com/huggingface/smolagents/blob/c30b115286e000e98711fae5e85993547b73d826/src/smolagents/agents.py#L876) (verified)
  - *To reach the next level:* Tool calls and sub-agents do not count against the parent's budget, and timeouts do not stop work.
- **D L1:** The defaults are sensible (20 steps), but model-written code can call a managed agent with its own max_steps, because run() takes max_steps and managed-agent kwargs are passed straight through. — [src/smolagents/agents.py:300](https://github.com/huggingface/smolagents/blob/c30b115286e000e98711fae5e85993547b73d826/src/smolagents/agents.py#L300); [src/smolagents/agents.py:468](https://github.com/huggingface/smolagents/blob/c30b115286e000e98711fae5e85993547b73d826/src/smolagents/agents.py#L468); [src/smolagents/agents.py:876](https://github.com/huggingface/smolagents/blob/c30b115286e000e98711fae5e85993547b73d826/src/smolagents/agents.py#L876); [src/smolagents/local_python_executor.py:918](https://github.com/huggingface/smolagents/blob/c30b115286e000e98711fae5e85993547b73d826/src/smolagents/local_python_executor.py#L918) (verified)
  - *To reach the next level:* The model can raise sub-agent limits; the cap needs to be enforced outside the call arguments.
- **B L1:** Stopping is cooperative between steps, timed-out executor threads keep running, and there is no spend ceiling. — [src/smolagents/agents.py:546](https://github.com/huggingface/smolagents/blob/c30b115286e000e98711fae5e85993547b73d826/src/smolagents/agents.py#L546); [src/smolagents/agents.py:756](https://github.com/huggingface/smolagents/blob/c30b115286e000e98711fae5e85993547b73d826/src/smolagents/agents.py#L756); [src/smolagents/local_python_executor.py:299](https://github.com/huggingface/smolagents/blob/c30b115286e000e98711fae5e85993547b73d826/src/smolagents/local_python_executor.py#L299); searched `rg -n -i 'wall_clock|max_time|max_cost|budget|max_tokens_total'` in `src/` → 0 hits (verified)
  - *To reach the next level:* No cost ceiling, and in-flight execution is not cancelled on stop or timeout.
- **Cap:** none
- **Notes:** Inferred from Python's stdlib: ThreadPoolExecutor.__exit__ calls shutdown(wait=True), so leaving the with-block after FuturesTimeoutError waits for the worker thread. This means the 30s timeout may not even return control early. Not executed.

## Rule-of-Two check
[A] untrusted input: Web pages and search results fed back as user-role observations (src/smolagents/default_tools.py:528, src/smolagents/models.py:284) · [B] sensitive data/systems: Model and tool API keys held on objects reachable from model code (src/smolagents/models.py:1535, src/smolagents/default_tools.py:186) · [C] state change / egress: Unrestricted GET to any URL and in-process code execution (src/smolagents/default_tools.py:528, src/smolagents/agents.py:1726) · Same default session? Yes

## Highest-impact improvements
1. Stop calling load_dotenv() at import time in remote_executors.py; load env files only on explicit request. — C6 S L0→L2, +0.150 before caps (Playbook 2)
2. Harden the bundled web UI's default exposure and session isolation. — C6 D L0→L2, +0.100 before caps (Playbook 2)
3. Run every tool call made from CodeAgent code through validate_tool_arguments, and add host allowlisting/internal-address blocking to VisitWebpageTool. — C3 C L0→L2, +0.150 before caps (Playbook 3)
4. Add an optional pre-execution approval hook that shows the exact code action, enabled by default for the local executor. — C2 S L0→L3, +0.225 before caps (Playbook 5)
5. Stop managed-agent calls from accepting max_steps from model code, and charge sub-agent steps to the parent's budget. — C10 D L1→L2, +0.050 before caps (Playbook 3 step 3)

## Re-audit log
- No changes.

## Limitations
- Static source review of the pinned commit only; nothing was executed, installed, or probed.
- Behaviour of third-party libraries (python-dotenv's .env search path, the OpenAI SDK's OPENAI_BASE_URL fallback, mcpadapt/mcp subprocess environment, ThreadPoolExecutor shutdown semantics, E2B default egress) is inferred from those libraries, not from this repo.
- Executor escapes and attribute-chain secret access are reasoned from the interpreter's source and its authors' own statements; no exploit was run.
- The `smolagent` and `webagent` CLIs, examples/, and docs/ were only skimmed; scoring follows the library's constructor defaults.
- No text aimed at AI reviewers was found in README.md, AGENTS.md, or SECURITY.md.
