BoundBench

AutoGen

Microsoft's Python and .NET framework for building multi-agent LLM applications (AgentChat, Core, Extensions), now in maintenance mode.

github.com/microsoft/autogen · 2026-10-04 · 027ecf0

Defense-in-depth score

2.8 / 10

Minimal

AutoGen gives developers building blocks but almost no safety defaults: tool calls run as soon as the model requests them, code-execution approval is an opt-in callback, and teams have no turn or cost limit unless one is configured. Code runs in a stock Docker container with full network and a read-write workspace, falling back to the host with the full environment when Docker is missing. In the flagship Magentic-One team, a malicious web page can steer unapproved code execution and data exfiltration with no human in the loop.

Key gaps (2)

  1. Code execution, the most powerful action, runs with no approval gate by default: approval_func defaults to None and PythonCodeExecutionTool has no hook. C2 · Approval gates
  2. A hijacked Magentic-One team can read local files (FileSurfer at cwd, WebSurfer file://), exfiltrate over the open web and run unapproved code, all unattended. C5 · Untrusted input blast radius

Criteria

C1 Identity & least privilege

Minimal 0.05 / 1.00

AutoGen has no identity or authorization layer of its own: every tool, code executor and MCP server runs with whatever authority the host Python process has, and nothing checks a request against the person who asked for it. The host-process code executor copies the full process environment, including model API keys, into every script it runs. The Docker executor does not pass the environment, which keeps host credentials out of the default sandbox, but nothing scopes the credentials developers hand to tools. The optional distributed gRPC runtime listens on an unauthenticated, unencrypted port.

C2 Approval gates

Minimal 0.28 / 1.00

Ordinary tool calls made by AssistantAgent, including MCP tools and the Python code-execution tool, run as soon as the model asks for them; there is no approval hook on that path. CodeExecutorAgent offers an optional approval callback that sees the exact code and can reject it, and it warns when none is set, but it is off by default and covers only that one agent. The Magentic-One team runs code without approval unless the developer supplies the callback, and the docs show using another LLM as the approver.

C3 Tool & action scoping

Minimal 0.45 / 1.00

Every function tool validates its arguments against a typed Pydantic schema before running, which catches malformed calls but is not an allowlist. A few bundled tools add real scoping (the file surfer confines paths with realpath containment), while others are wide open: the web surfer visits any URL including file:// and internal addresses, the HTTP tool formats model input straight into the path, and the code-execution tool takes arbitrary code. MCP tool arguments are forwarded without local validation. AssistantAgent starts with no tools, but the official Magentic-One preset enables web, file and code execution together.

C4 Code-execution isolation

Minimal 0.47 / 1.00

Model-written code runs either on the host (LocalCommandLineCodeExecutor, which copies the full environment, API keys included) or in a stock Docker container that runs as root with default capabilities, full network and the workspace mounted read-write. The helper used by Magentic-One prefers Docker but quietly falls back to the host executor, with only a Python warning, when Docker is missing. An experimental Azure Container Apps executor offers a stronger remote sandbox but must be chosen explicitly. MCP stdio servers always run on the host.

C5 Untrusted input blast radius

Minimal 0.00 / 1.00

Nothing in AutoGen limits what a hijacked agent can do after reading untrusted content. Tool results, web pages and other agents' messages enter the conversation with the same standing as the user's instructions, and CodeExecutorAgent by default executes code blocks from any participant's message. In the documented Magentic-One team, a web page the browser agent reads can steer the coder into writing code that runs without approval and with full network access, and the browser itself can open file:// URLs and send data to any site. Leaking data and taking destructive action can both happen with no human involved.

C6 Memory, context & configuration integrity

Minimal 0.25 / 1.00

Memory is opt-in and the framework never writes to it on the model's behalf, but anything a developer stores is re-injected as a system message with no provenance marking, so poisoned memory carries system-level weight in every later run. Persistent stores such as ChromaDB default to one shared collection with no per-user namespace. AutoGen does not auto-load instruction or settings files from the working directory, and component configs are only loaded when the developer calls load_component.

C7 Third-party extensions

Minimal 0.30 / 1.00

Developers add MCP servers explicitly in code, which avoids silent installs, but AutoGen launches whatever command it is given with no pinning, integrity check or detection of changed tool definitions. Component configs can only name provider classes from AutoGen's own namespaces, a useful allowlist, but it can be widened by an environment variable and its enforcement does not cover every path, and a FunctionTool config runs arbitrary embedded Python via exec. The Docker executor pulls the unpinned python:3-slim image. MCP stdio servers run on the host as the same user; in-process adapters such as LangChain tools share the agent's process.

C8 Secrets & sensitive-data protection

Minimal 0.30 / 1.00

Model API keys come from environment variables and are typed as SecretStr in config models, which keeps them out of reprs and serialized configs. There is no redaction anywhere else: when event logging is enabled, every LLM call is logged with the full prompt and response, and the host-process code executor gives model-written code a copy of the full environment. AutoGen sends no telemetry of its own; tracing is OpenTelemetry and is a no-op unless the developer configures a provider.

C9 Audit & traceability

Minimal 0.35 / 1.00

Each run returns a structured stream of tool-call request and execution events attributed to the agent that made them, and function tools log a ToolCallEvent with arguments and result. None of this is persisted unless the developer attaches a logging handler or an OpenTelemetry provider; by default the record exists only in memory. MCP calls produce trace spans but not event-log entries, approvals are not logged as distinct events, and nothing records the requesting human.

C10 Limits & kill switch

Minimal 0.30 / 1.00

A single AssistantAgent stops after one round of tool calls by default and code executors time out after 60 seconds, killing the process. Multi-agent teams, however, default to no turn limit and no termination condition, so a round-robin or selector team runs until something external stops it. Turn, token and wall-clock limits exist as opt-in termination conditions checked between messages; nested teams and agent-as-tool start fresh budgets, and there are no rate limits on side-effecting tools.