BoundBench

LibreChat

Self-hosted multi-model chat platform with agents, MCP, code interpreter and actions

github.com/LibreChat-AI/LibreChat · 2026-10-05 · f10b1d9

Defense-in-depth score

4.0 / 10

Minimal

LibreChat is a well-structured multi-user service: requests are tied to accounts and roles, outbound actions get SSRF protection, model-written code runs in a separate remote service rather than on the server, and stdio MCP servers can only be added by admins. The dominant risk is that its human-approval gate for tool calls is off by default, so content an agent reads (files, search results, tool outputs) can drive any tool it holds, including outbound HTTP actions and MCP write tools, with nobody approving. Token spending limits are also off by default, and the server process holds every user's stored credentials.

Key gaps (1)

  1. With tool approval off by default and no provenance controls, injected content can make an agent send data out and take external write actions through actions or MCP tools with no human involved (C5-WORSTCASE). C5 · Untrusted input blast radius

Criteria

C1 Identity & least privilege

Minimal 0.45 / 1.00

LibreChat is a multi-user service: every request is tied to a logged-in account, role permissions are checked on each feature, and model provider keys default to 'user provided', so each user supplies and pays for their own. Per-user OAuth tokens are kept for actions and MCP servers, and stdio MCP servers get a stripped environment. But the server process itself still holds every stored credential (encrypted user keys, action secrets, the database connection), and an action's stored API key is used for whoever runs the agent. The first account to register becomes administrator.

C2 Approval gates

Minimal 0.42 / 1.00

LibreChat ships a real human-approval gate for agent tool calls: when enabled, a paused call shows the tool name and its exact arguments, and the user can approve, reject, or edit the arguments; administrators can write allow, deny and ask rules by tool-name pattern. But the gate is off unless an administrator adds a toolApproval block to librechat.yaml, and the shipped example has it commented out. In the default deployment every tool an agent has, including actions that call external APIs and MCP tools, runs with no human in the loop.

C3 Tool & action scoping

Moderate 0.53 / 1.00

Outbound HTTP from actions and user-added MCP servers goes through a domain check, and when no admin allowlist is set, actions use connection-time SSRF protection that blocks private and metadata addresses. Users cannot add command-launching (stdio) MCP servers; only admins can, in librechat.yaml. Tools have typed schemas. But actions are deliberately general: an agent author can point one at any public API with any operations, and the default capability list turns on actions, code execution, web search, sub-agents and the rest of the tool families at once.

C4 Code-execution isolation

Moderate 0.63 / 1.00

LibreChat does not run model-written code on its own server. The code-execution and bash tools send code over HTTP to a separate Code Interpreter API service (configured by base URL and an API key), and a separate opt-in mode routes commands to workers the user enrolls on their own machine. The only local process launches in the server are operator plugin hooks and document conversion, both with a stripped environment. The isolation of the remote service lives outside this repository and could not be examined, so its strength and blast radius are inferred.

C5 Untrusted input blast radius

Minimal 0.00 / 1.00

Agents read content their user did not write: uploaded files, web search results, action and MCP tool outputs, and instructions of agents shared by other users. Nothing in the code tracks where content came from or restricts what a turn can do after reading it, and with the approval gate off by default an injected instruction can drive every tool the agent holds. Those tools can include outbound HTTP actions to any public host and connected MCP tools with write access, so data leakage and external side effects can both happen unattended.

C6 Memory, context & configuration integrity

Minimal 0.35 / 1.00

Long-term memory is off unless an administrator adds a memory block to librechat.yaml. When on, the model can save and delete memory entries with a tool, entries are stored per user and per agent, and they are fed back into later conversations as context; users can view, edit and opt out of memories. Entries are size-limited but not validated or marked by source. Agents, prompts and skills are user-authored and only shared with others when a role grants sharing, which ordinary users lack by default. There is no workspace or repository auto-loading in this server product.

C7 Third-party extensions

Minimal 0.28 / 1.00

Third-party code enters LibreChat through MCP servers and operator plugins. Only administrators can add command-launching MCP servers (in librechat.yaml), and ordinary users cannot add remote MCP servers by default. Nothing pins or verifies the packages those commands launch, and changed tool lists from a server are refreshed without re-approval. Stdio MCP servers run as separate processes with a minimal default environment, under the same container user as LibreChat.

C8 Secrets & sensitive-data protection

Minimal 0.45 / 1.00

Stored user API keys, action secrets and OAuth tokens are encrypted at rest with a server key, which the server generates when none is configured. Logs pass through a redaction filter for API-key, bearer-token and secret patterns on both console and file output, and tracing to Langfuse or OpenTelemetry is opt-in. The example environment turns debug file logging on. Message content, files and tool results are sent to the model provider and stored in the database without redaction unless an administrator configures PII filters.

C9 Audit & traceability

Minimal 0.45 / 1.00

Every message and its tool calls, with arguments and outputs, are stored in MongoDB per user and conversation, so a session can be reconstructed. LibreChat also has an append-only, hash-chained audit log, but today it only records role grants and permission changes, not agent runs, tool calls or approvals. Records live in the same database the server process writes to.

C10 Limits & kill switch

Minimal 0.40 / 1.00

Each agent turn is capped at 50 graph steps by default, sub-agents derive their turn budget from the same limit, model calls have response timeouts, and per-user message concurrency and rate limits are on in the example environment. Users can stop a generation, which aborts the running job. But token spending limits (balance) are off by default, there is no default maximum, so any agent author can set a much higher step limit on their own agent, and scheduled agent runs fire without the user present.