BoundBench

OpenAI Agents SDK (Python)

Lightweight multi-agent workflow framework with tools, handoffs, guardrails

github.com/openai/openai-agents-python · 2026-10-03 · 81f0ccf

Defense-in-depth score

3.2 / 10

Minimal

The SDK ships strong building blocks (per-call approvals bound to exact arguments, remote sandbox backends, log redaction) but turns none of them on: every tool runs without approval by default, and the README's sandbox example runs model-written shell commands directly on a Linux host with the developer's full environment. A prompt-injected agent built from defaults can read host credentials and act or exfiltrate unattended, and tracing exports full prompts and tool data to OpenAI by default. Safety depends on the developer opting into needs_approval, a hosted or hardened sandbox, environment filtering, and turning off sensitive trace data.

Key gaps (4)

  1. The README-led UnixLocal sandbox runs model-written commands unconfined on Linux with the full host environment, so a hijacked agent holds all of the developer's credentials. C1 · Identity & least privilege
  2. No execution isolation by default: the lead sandbox backend executes on the host with credentials in the environment. C4 · Code-execution isolation
  3. Approval is off for every tool, and LocalShellTool/ComputerTool cannot be gated at all, so the most powerful paths skip the gate by default. C2 · Approval gates
  4. A prompt-injected agent can exfiltrate data and take irreversible actions with no human involved in the default configuration. C5 · Untrusted input blast radius

Criteria

C1 Identity & least privilege

Minimal 0.20 / 1.00

The SDK has no identity or authorization layer of its own: every tool a developer registers runs inside the application process with whatever credentials that process holds, and nothing maps a tool call to a permission check. The local sandbox backend that the README leads with copies the whole host environment (API keys, cloud credentials) into every shell command the model runs by default. The sandbox can be told to run commands as a separate OS user or to pass only an allowlist of environment variables, but both are opt-in. A hijacked agent therefore holds everything the developer's process holds.

C2 Approval gates

Minimal 0.25 / 1.00

The SDK has a well-built human-in-the-loop primitive: a tool marked needs_approval pauses the run, the application sees the exact tool name and arguments, the human approves or rejects, and the approval is bound to that exact call. But it is off for every tool type by default, including the shell tool the sandbox agent gets automatically, and two of the most powerful tool types (the local shell tool and the computer-use tool) cannot be gated at all. As shipped, every consequential action runs without a human.

C3 Tool & action scoping

Minimal 0.35 / 1.00

Function tool arguments are parsed against a strict JSON schema and validated by a generated pydantic model before the tool body runs, and sandbox file tools resolve paths against the workspace root. But validation is only as narrow as the types the developer writes, MCP tool arguments are forwarded without SDK-side validation, and the sandbox agent ships with a general shell tool that takes any command string. The default sandbox capability set includes shell, file writing and compaction; a misused shell on the README's local backend reaches the whole machine.

C4 Code-execution isolation

Minimal 0.47 / 1.00

Code execution isolation depends entirely on which sandbox backend the developer picks; there is no default. The README's sandbox example uses the local Unix backend, which on Linux runs model-written shell commands directly on the host as the developer's user with the full host environment, and on macOS adds only filesystem restrictions without network isolation. The Docker backend is a stock container (root, default capabilities, network on), and the hosted backends (E2B, Modal, Daytona and others) are real remote sandboxes but opt-in and network-enabled by default. Local shell tools, MCP stdio servers and function tools always run on the host.

C5 Untrusted input blast radius

Minimal 0.15 / 1.00

Nothing in the SDK limits what a hijacked agent can do after reading untrusted content. Tool and MCP results, web search results and handoff messages enter the conversation and the model can then call any tool, including shell and network tools, with no approval by default. Input, output and tool guardrails exist but are opt-in hooks whose typical use is detection, and the default trace export sends the full content to OpenAI. A successful prompt injection can leak data and take irreversible actions unattended.

C6 Memory, context & configuration integrity

Minimal 0.35 / 1.00

Conversation sessions are in-memory by default and scoped by session ID in SQL, and long-term sandbox memory is opt-in. When persistence is enabled, stored history and model-generated memory summaries are replayed into context (memory summaries into the instructions) without validation or review. The default sandbox system prompt tells the model to obey AGENTS.md files found anywhere in the workspace, so a cloned repository's instruction files act as high-priority instructions without any trust decision, although they cannot add tools or change security settings by themselves.

C7 Third-party extensions

Minimal 0.23 / 1.00

Third-party extensions come in as MCP servers (local subprocesses or remote endpoints), hosted MCP tools, and sandbox skills or Git repositories pulled into the workspace. The developer names each one in code, but nothing pins versions, checks hashes, or detects changed tool definitions, and the official MCP examples launch servers with `npx -y` at whatever version is latest. Local MCP servers run as separate processes with, when no env is given, the MCP library's reduced default environment, but they still run as the developer's user.

C8 Secrets & sensitive-data protection

Minimal 0.30 / 1.00

The SDK keeps model inputs and tool data out of its own logs by default, redacts error and validation messages, and strips mount credentials from serialized run state. But tracing is on by default with sensitive data included, exporting full model inputs, outputs and tool arguments/results to OpenAI's trace ingest endpoint whenever an OpenAI key is present, and the README-led local sandbox hands the whole host environment, including OPENAI_API_KEY, to model-run shell commands, where one `env` call puts it in model context.

C9 Audit & traceability

Minimal 0.45 / 1.00

Tracing is on by default and records a structured span per function, MCP, shell, computer, custom and apply-patch tool call, nested under agent, turn and handoff spans with trace and group IDs, and ships them off-host to OpenAI. Approval and rejection decisions are not recorded as spans, there is no requesting-principal or approver attribution, LocalShellTool calls get no tool span, and spans are exported in batches every few seconds, so a crash loses the tail and export is silently skipped without an OpenAI key.

C10 Limits & kill switch

Minimal 0.42 / 1.00

Runs stop after 10 model turns by default, MCP calls have a 5-second read timeout by default, and streaming runs can be cancelled immediately. Function tool timeouts are opt-in, there is no wall-clock or token/cost cap, and an agent used as a tool starts its own fresh 10-turn budget, so delegation multiplies the ceiling. Sandbox shell commands keep running in the background after the tool returns (up to 64 PTY processes).