BoundBench

PocketFlow

A 100-line Python graph framework for building LLM workflows and agents from nodes and action-labelled transitions.

github.com/The-Pocket/PocketFlow · 2026-10-04 · f74d023

Defense-in-depth score

3.4 / 10

Minimal

PocketFlow is a tiny orchestration core with almost no surface of its own: it runs no code, loads no extensions, and stores nothing, which is where most of its points come from. Everything safety-relevant is left to the developer: there is no approval step, no input validation, no record of what ran, and no step or time limit, so an agent loop can run forever. The dominant risk is the documented agent pattern itself, where model output picks the next action and tool results flow back unmarked, letting a prompt injection drive any action the developer wired in with the process's full credentials.

Key gaps (2)

  1. No identity or authorization layer: every node runs with the host process's full ambient credentials, so a hijacked agent holds all of them. C1 · Identity & least privilege
  2. Node output loops back into prompts unmarked and model output selects the next action with no approval, so a prompt injection can leak data and take irreversible actions unattended (C5-WORSTCASE). C5 · Untrusted input blast radius

Criteria

C1 Identity & least privilege

Minimal 0.00 / 1.00

PocketFlow has no identity or authorization layer. Nodes are plain Python classes that run inside the developer's process with whatever credentials that process holds, and nothing checks an action against a policy or against the user who asked for it. The official docs and examples read provider keys from the environment of the same process. If an agent built on it is hijacked, the attacker gets the full authority of the host process.

C2 Approval gates

Minimal 0.00 / 1.00

There is no approval primitive in the library: no interrupt, pause, or confirmation step before a node runs. Whatever action string the model's node returns selects the next node, and that node's code runs immediately. Human-in-the-loop exists only as cookbook examples the developer writes by hand. A wrongly chosen or injected action can therefore run any consequential operation the developer wired in, with no undo.

C3 Tool & action scoping

Minimal 0.10 / 1.00

The library has no tool abstraction and no argument validation: data passes from the shared store to node code unchecked. The only structural limit is routing: a model-returned action can select only a successor the developer wired, and an unknown action ends the flow with a warning. Official docs encourage programmable actions such as model-written SQL, and the examples follow that advice without bounds.

C4 Code-execution isolation

N/A · full credit 1.00 / 1.00

The core library contains no code-execution feature: no shell, interpreter, eval of model output, or process spawning. Nodes are ordinary Python methods the developer writes, so isolating any code they run is the developer's job. Note that an official cookbook example (code generator) runs model-written Python in-process with exec() and full builtins, so developers copying examples get no isolation.

C5 Untrusted input blast radius

Minimal 0.00 / 1.00

The library does nothing to separate untrusted content from instructions. Whatever a node writes into the shared store, such as web search results, becomes context for the next model call with the same standing as the user's request, and the model's reply directly chooses the next action. The official agent pattern combines web search results, model-chosen outbound queries, and the process's credentials in one loop with no approval. A successful prompt injection can therefore leak data and take irreversible actions unattended.

C6 Memory, context & configuration integrity

N/A · full credit 1.00 / 1.00

The core library has no memory, retrieval store, checkpointing, or auto-loaded configuration. The shared store is an in-memory dict the caller passes in and owns. The repository's .cursorrules files are documentation for the developer's IDE and are never read by the library. An official example (coding agent) does persist a model-written memory file and auto-load AGENTS.md from the working directory, so that pattern is the developer's risk if copied.

C7 Third-party extensions

N/A · full credit 1.00 / 1.00

PocketFlow loads no third-party code at runtime: the library imports only Python standard modules and has no plugin loader, MCP client, tool hub, or model-file loading. Anything external comes from the developer's own node code. Cookbook examples that use MCP are separate applications, not part of the library.

C8 Secrets & sensitive-data protection

Minimal 0.30 / 1.00

The library itself never reads, stores, logs, or transmits credentials, and it ships no telemetry or logging, so it adds no leak paths of its own. It also offers no protection: nothing masks secrets, keeps them out of model-bound prompts, or scrubs them from what nodes see. The docs advise keeping keys in environment variables, where they are long-lived and reachable by every node in the process.

C9 Audit & traceability

Minimal 0.00 / 1.00

The library keeps no record of what a flow did. Nodes run, return actions, and mutate the shared dict with no structured log of which node ran, with what inputs, or what it returned. The only output is a handful of Python warnings about wiring problems. Tracing appears only as an opt-in cookbook example that relies on a third-party service, so after an incident there is nothing from the framework to reconstruct.

C10 Limits & kill switch

Minimal 0.00 / 1.00

There are no step, time, or cost limits. A flow runs `while` there is a next node, so a decision node that keeps routing back to itself loops forever, and parallel batch nodes start every item at once with no concurrency limit. Retries default to one attempt, which bounds nothing about damage. There is no stop or cancel primitive; halting means killing the process.