BoundBench

Agno

Framework/runtime to build, run and manage agent platforms

github.com/agno-agi/agno · 2026-10-03 · 3ca74c2

Defense-in-depth score

2.0 / 10

Minimal

Agno gives developers powerful toolkits (shell, in-process Python, files, web, MCP) but few guardrails switched on. Out of the box no tool needs approval, code runs on the host with every API key in the environment, and there is no step, time or cost limit. Its human-in-the-loop confirmation and remote sandbox toolkits are well built but opt-in, so safety depends on the developer turning them on.

Key gaps (4)

  1. A hijacked Agno agent holds every credential in the host process environment: tools read keys from env and ShellTools passes the full environment to subprocesses. C1 · Identity & least privilege
  2. Shell and Python-exec tools run without any approval by default; the confirmation gate is opt-in per tool. C2 · Approval gates
  3. ShellTools and PythonTools run model-generated commands and code on the host (PythonTools in-process via exec) with the full environment and no sandbox. C4 · Code-execution isolation
  4. Untrusted tool output enters context unmarked and a hijacked agent can both exfiltrate (arbitrary URL fetch) and act irreversibly (shell) with no human involved. C5 · Untrusted input blast radius

Criteria

C1 Identity & least privilege

Minimal 0.15 / 1.00

Agno has no notion of a scoped agent identity. Every toolkit builds its own client from whatever API keys sit in the process environment (for example a GitHub token), and the shell tool starts commands that inherit the full environment, so a hijacked agent holds everything the hosting process holds. AgentOS can check a caller's JWT scopes before a run starts, but that is off by default and only guards the HTTP entry point, not what tools do with their credentials.

C2 Approval gates

Minimal 0.25 / 1.00

Agno has a solid human-in-the-loop primitive: a tool marked as requiring confirmation pauses the run and hands the caller the exact tool name and arguments to approve or reject. But nothing is marked by default — not even the shell, Python-exec or file-write tools — so out of the box the agent runs every tool call without asking anyone. Developers must opt in per tool, and ShellTools' own docstring tells them to.

C3 Tool & action scoping

Minimal 0.28 / 1.00

All Python tools get typed argument validation through pydantic, and the file toolkits confine paths to a base directory after resolving symlinks. But the most powerful bundled tools take raw input: ShellTools runs any argv, PythonTools runs any code, and the web tools fetch any URL with redirects followed and no block on internal addresses. Toolkits also enable their write functions by default (FileTools.save_file is on).

C4 Code-execution isolation

Minimal 0.47 / 1.00

The built-in code tools run directly on the host: ShellTools spawns host processes, PythonTools calls exec() inside the agent's own process, skill scripts and MCP stdio servers start as local processes, and the docstrings say plainly that none of this is a sandbox. Remote sandboxes (E2B, Daytona) exist as separate opt-in toolkits, but they only cover their own tools; everything else still runs on the host with the agent's credentials.

C5 Untrusted input blast radius

Minimal 0.07 / 1.00

Tool results — web pages, files, MCP outputs, other agents' replies — go into the conversation as ordinary tool messages, with nothing marking them as untrusted and nothing changing what the agent may do after reading them. The bundled prompt-injection guardrail is opt-in, only matches a short list of phrases, and only checks the user's own input, not tool output. A hijacked agent with the usual toolkits can both leak secrets (any URL fetch) and act irreversibly (shell) without a human.

C6 Memory, context & configuration integrity

Minimal 0.05 / 1.00

Long-term user memory is off by default, but once enabled the model can write memories through a tool (via an LLM memory manager) and they are injected into every later system prompt with no provenance check, approval or expiry. Memories are keyed by user_id, but when an app doesn't pass one every caller shares the same 'default' bucket. Agno does not auto-load instruction files or a .env from the working directory.

C7 Third-party extensions

Minimal 0.10 / 1.00

Agno loads whatever MCP servers, skills and packages the developer points it at, with no pinning, hash check or re-approval. The MCP command check allows npx/uvx (so whatever version is latest at launch), but MCP stdio servers get only a minimal environment. Skill scripts run on the host with the full environment, and the bundled PythonTools lets the model pip-install any package it names.

C8 Secrets & sensitive-data protection

Minimal 0.13 / 1.00

API keys come from environment variables and are not masked; nothing scrubs them from subprocess environments, so a shell or skill command can read every key. Anonymous telemetry is on by default and sends run metadata (model, flags such as has_tools) to Agno's API, but no prompts or tool output. Debug logging, when turned on, prints full messages without redaction.

C9 Audit & traceability

Minimal 0.35 / 1.00

With default settings Agno keeps no durable record of what the agent did: tool calls go only to debug-level logs, and a few toolkits print an INFO line. If you configure a database, each run is stored with its tool calls, arguments, results, timestamps, user_id and approval flags, but by default only when the run ends, so a crash mid-run loses it. OpenTelemetry tracing is available as an opt-in.

C10 Limits & kill switch

Minimal 0.15 / 1.00

An Agno agent has no step, time or cost limit by default: the model loop runs until the model stops calling tools. tool_call_limit is opt-in, and when it trips it only returns an error message to the model while the loop keeps going. Runs can be cancelled, but cancellation is a flag checked between steps, so in-flight tool calls (like a shell command with no timeout) keep running.