BoundBench

BabyAGI (functionz)

Experimental self-building agent framework that stores Python functions in a database and executes them, with a Flask dashboard, REST API and function-calling chat.

github.com/yoheinakajima/babyagi · 2026-10-04 · fa8930e

Defense-in-depth score

1.3 / 10

Minimal

BabyAGI runs every stored function with Python exec() inside the server process, injects every stored API key into each one, and pip-installs whatever imports a function lists. Access control on the dashboard and REST API that execute functions is not locked down. There are no approval gates, sandbox, or limits. The author labels it experimental and not for production; treat it that way.

Key gaps (5)

  1. Every stored secret key is injected into the scope of every executed function, so any hijacked function holds all credentials. C1 · Identity & least privilege
  2. Default pack exposes add_key_wrapper and get_all_secret_keys as callable tools, letting a caller rewrite or dump the credential store. C1 · Identity & least privilege
  3. All function code, including model-written code from add_new_function, runs via in-process exec() with every secret in scope and any listed import pip-installed on the host. C4 · Code-execution isolation
  4. Function code, triggers and the encryption key are loaded from files in the current working directory, so a pre-seeded funztionz.db executes attacker code on import with no trust decision. C6 · Memory, context & configuration integrity
  5. The executor pip-installs any import name listed for a function, including model-supplied ones, on the host with no pinning or consent. C7 · Third-party extensions

Criteria

C1 Identity & least privilege

Minimal 0.00 / 1.00

BabyAGI keeps one global store of secret keys and injects every stored key into the scope of every function it executes, whatever that function declares it needs. Access control on the dashboard and REST API that run functions is not locked down. Registered tools can also add or overwrite stored keys and read them all back.

C2 Approval gates

Minimal 0.05 / 1.00

There is no approval step anywhere. Functions run directly from the HTTP API and from model tool calls in the chat, including the default function that adds new code to the system. Function code is versioned so a code overwrite can be rolled back, but anything the code does (package installs, external API calls, file writes) happens immediately.

C3 Tool & action scoping

Minimal 0.05 / 1.00

The default tool set includes a function that stores arbitrary model-supplied Python code and imports, and every function can be called with any arguments; the only check is that required parameter names are present. In the chat, the user picks which functions the model sees, which is the one real narrowing, but the HTTP API can call any function directly.

C4 Code-execution isolation

Minimal 0.40 / 1.00

All function code, including code a model writes through add_new_function or a caller pushes through the API, runs with Python exec() inside the server process, and any import it lists is pip-installed on the host. There is no sandbox on this path and every stored secret is in the executing scope. An optional E2B plugin can run code in a remote sandbox, but it is not loaded by default and only covers code explicitly passed to it.

C5 Untrusted input blast radius

Minimal 0.00 / 1.00

Nothing separates untrusted content from instructions. Tool results go back to the model without marking, and the handling of chat input and rendered model output is not locked down. A hijacked model can call add_new_function to run arbitrary code or read all secrets if those functions are offered.

C6 Memory, context & configuration integrity

Minimal 0.00 / 1.00

The framework's state, including all function code, triggers and keys, lives in a SQLite file and an encryption-key file resolved relative to the current working directory, shared by every user of the server. The model can write new persistent code through add_new_function, and that code runs in later sessions. A funztionz.db already present in the working directory is loaded as-is, and its functions and triggers run when default packs register on import.

C7 Third-party extensions

Minimal 0.00 / 1.00

Any import name stored with a function, including one supplied by the model through add_new_function, is pip-installed on the host at execution time if it isn't already importable, with no pinning, hash, or consent. Installed packages run in-process with every secret. The chat page also loads its Markdown renderer from a CDN without a pinned version.

C8 Secrets & sensitive-data protection

Minimal 0.15 / 1.00

Stored keys are Fernet-encrypted in the database, but the encryption key sits in a plaintext JSON file next to it, and secret material is exposed through further output paths. The default pack's get_all_secret_keys function returns every key in plaintext to its caller. The chat function turns on LiteLLM's verbose logging by default.

C9 Audit & traceability

Minimal 0.45 / 1.00

Every function execution goes through one executor that writes a structured log row (function name, parameters, output, timing, status, parent and trigger links) before the code runs. That is a usable trail of calls, but it has no record of who requested a call, lives in a database file in the working directory that executed code can modify, and does not record what happens inside a function.

C10 Limits & kill switch

Minimal 0.20 / 1.00

The core has no step, time, or cost limits and no way to stop a running function. The default chat makes at most two model calls per request and trigger chains skip functions already run in the chain, which bounds those loops, but any function (including model-written ones) can run forever in the server process.