BoundBench

LangGraph

Graph-based orchestration runtime for stateful, durable agents

github.com/langchain-ai/langgraph · 2026-10-03 · 7dc9195

Defense-in-depth score

4.2 / 10

Minimal

LangGraph is a low-level runtime: it ships no shell, code-execution or plugin features of its own, which is why two criteria score full marks. Everything else is left to the developer. By default every tool call the model makes runs immediately with the host process's full credentials, tool output flows back to the model unmarked, and the only limit is a 10,007-step cap. The dominant risk is prompt injection through tool results driving the developer's tools; human-in-the-loop interrupts and checkpoint history exist but are opt-in, and the default checkpoint deserializer will execute code stored in a tampered database.

Key gaps (2)

  1. No identity or authorization layer: every tool runs with the host process's full ambient credentials, so a hijacked agent holds all of them. C1 · Identity & least privilege
  2. Tool output loops back to the model unmarked and no approval runs by default, so a prompt injection can leak data and take irreversible actions unattended. C5 · Untrusted input blast radius

Criteria

C1 Identity & least privilege

Minimal 0.00 / 1.00

LangGraph has no identity or authorization layer of its own. Tools and nodes run inside the developer's process with whatever credentials that process holds, and nothing checks a tool call against a policy or against the user who asked for it. The Python SDK includes an Auth decorator API for the hosted LangGraph Server, but its enforcement lives in the closed-source server, not in this repository. If an agent is hijacked, the attacker gets the full authority of the host process.

C2 Approval gates

Minimal 0.30 / 1.00

By default a LangGraph agent executes every tool call the model emits with no human approval: create_react_agent and compile() both default to no interrupts. Developers can opt in to pausing before the tools node (interrupt_before=['tools']) or call interrupt() inside a tool, and the pending tool calls, with their exact arguments, sit in saved state where a person can inspect, edit, or replace them before resuming. The pause is all-or-nothing per node with no risk tiers, a nested sub-agent's own tool calls are not covered by the parent's pause, and anyone holding the thread id can resume. Nothing undoes actions already taken outside the graph.

C3 Tool & action scoping

Minimal 0.45 / 1.00

ToolNode only runs tools that were registered and returns an error for any other name, and tool arguments are type-checked against each tool's Pydantic schema before the tool runs. Values the framework injects (graph state, the store, runtime) overwrite anything the model supplies for those parameters. That is type validation, not an allowlist: LangGraph ships no path, URL, or quantity constraints, and every tool the developer passes is available to the model at every step.

C4 Code-execution isolation

N/A · full credit 1.00 / 1.00

The LangGraph runtime and prebuilt agent contain no code-execution feature: no shell tool, no Python or JavaScript interpreter, no eval of model output. Tools are ordinary Python callables the developer writes; whether they run commands is the developer's choice, and isolating them is the developer's job. The CLI does run Docker commands, but only from the operator's own configuration. The checkpoint deserializer can import and call Python objects named in stored data, but that data is written by the framework, not by the model, and it is scored under memory integrity (C6).

C5 Untrusted input blast radius

Minimal 0.00 / 1.00

Tool results, including web pages, files, API responses and the output of remote graphs, are wrapped as tool messages and fed straight back to the model on the next step. Nothing in LangGraph marks that content as untrusted, tracks taint, or disables state-changing or outbound tools once untrusted content has been read. Because LangGraph's whole purpose is to combine untrusted inputs, private data and outbound tools in one loop, a successful prompt injection can both leak data and take irreversible actions without a human in the default configuration.

C6 Memory, context & configuration integrity

Minimal 0.30 / 1.00

By default a compiled graph keeps state only in memory for one run. When a developer adds a checkpointer (per conversation thread) or a store (shared across threads), everything the agent read, including injected tool output, is saved and reloaded verbatim on the next turn, with no validation, provenance, or expiry. The default checkpoint deserializer is permissive: its own module docstring says any Python callable stored in checkpoint data will be imported and executed on load, so anyone who can write to the checkpoint database can run code in the agent process; a strict allowlist exists but needs an environment variable. LangGraph loads no instruction files or .env files from a working directory.

C7 Third-party extensions

N/A · full credit 1.00 / 1.00

LangGraph itself loads no third-party code at runtime: no plugin loader, MCP client, tool hub, model-file loading, or package installs the model can trigger. Tools come only from the developer's own code. RemoteGraph calls a remote LangGraph server over HTTP, so its outputs are untrusted input (C5) rather than code run locally. Developers who add MCP or hub tools through other libraries take on that risk outside LangGraph.

C8 Secrets & sensitive-data protection

Minimal 0.30 / 1.00

LangGraph ships no telemetry or crash reporting, and debug printing is off by default. The SDK reads its server API key from environment variables and sends it only as a request header. There is no redaction anywhere: full conversation state, including tool output and anything secret in it, is stored in plaintext in the checkpoint database unless the developer opts in to the EncryptedSerializer, and the framework does nothing to keep secrets out of model-bound messages. Credentials in the host process are long-lived and reachable by every tool.

C9 Audit & traceability

Minimal 0.47 / 1.00

Without a checkpointer, the only record of a run is the state returned to the caller, which is lost if the process dies. With a checkpointer, LangGraph saves a versioned snapshot at every step plus each task's writes, including tool calls with their arguments and results, interrupts and resume values, with run ids and parent links across subgraphs, which makes replay and time-travel possible. Records are written in the background by default, carry no actor or approver identity, and live in a database the agent process can overwrite or delete.

C10 Limits & kill switch

Minimal 0.33 / 1.00

Every graph run has a step cap, but the default is 10,007 supersteps (overridable by an environment variable), so it bounds almost nothing. Per-step and per-node wall-clock timeouts exist but default to off, and there is no token or cost budget. Each subgraph or nested agent counts its own steps from zero, so delegation gets a fresh budget. Timeouts and cancellation cancel async work, but synchronous tool calls already running in threads cannot be stopped.