BoundBench

AnythingLLM

Local-first all-in-one AI app with agents, MCP tools and RAG

github.com/Mintplex-Labs/anything-llm · 2026-10-05 · feb04ca

Defense-in-depth score

2.6 / 10

Minimal

AnythingLLM ships a real per-call approval step for its file, email and calendar tools, keeps opt-in filesystem access inside an allowlisted folder, and blocks Community Hub code downloads unless the operator enables them. The default agent can still store anything in the workspace's shared vector memory and fetch arbitrary URLs without asking, so injected content can persist across users and send document contents out. MCP servers and imported skills run with the server's full credentials, and MCP tools are never gated.

Key gaps (2)

  1. Imported agent skills are loaded into the server process and MCP servers inherit the server's full environment, so a malicious extension gets every stored provider key. C7 · Third-party extensions
  2. The single most powerful default actions (memory store and arbitrary-URL fetch) and all MCP tools are outside the approval gate by design (C2-POWERBYPASS). C2 · Approval gates

Criteria

C1 Identity & least privilege

Minimal 0.45 / 1.00

The agent runs inside the AnythingLLM server and uses the server's own credentials: every LLM, embedding and search provider key from the instance's .env file, plus any connector it has been given. Built-in tools are scoped to the workspace the chat belongs to, for example document search and memory only touch that workspace's vector namespace. MCP servers are launched with the server's full environment, and imported skills run inside the server process. The shipped Docker setup starts in single-user mode, and until the operator sets a password every request is treated as the instance owner.

C2 Approval gates

Minimal 0.25 / 1.00

AnythingLLM has a real per-call approval step: built-in tools that write files, create documents, send or draft email, change calendar events or create scheduled jobs pause and show the user the tool's arguments with approve and reject buttons, and API sessions without a human deny such calls. But the gate is opt-in per tool, so the default memory-store tool, the web scraper, the SQL agent, agent flows and every MCP tool run without asking. An operator environment variable can auto-approve any or all skills, users can tick 'always allow' per skill, and scheduled jobs auto-approve everything. There is no undo for agent actions.

C3 Tool & action scoping

Minimal 0.45 / 1.00

Tools are typed with JSON schemas, and the opt-in filesystem tools have solid path containment that resolves symlinks against an allowlisted root. The default web scraper accepts any URL; its address check only rejects some private IP literals, allows loopback, ignores hostnames, and is described in its own source as a convenience rather than a security control. The opt-in SQL tool sends raw model-written SQL, relying on its description to keep it read-only. Extension tools get no shared validation layer.

C4 Code-execution isolation

Minimal 0.10 / 1.00

AnythingLLM has no shell or code-interpreter tool, and the default skill set never executes model-written code. The one model-driven execution path is the opt-in SQL agent, which sends model-written SQL straight to an operator-configured database with no isolation or read-only enforcement. Imported skills and MCP servers are third-party code and are rated under extensions. The Docker image runs as a non-root user but the shipped compose file adds the SYS_ADMIN capability.

C5 Untrusted input blast radius

Minimal 0.05 / 1.00

Nothing in the code treats untrusted content differently from the user's own instructions: web pages, uploaded documents and tool or MCP results are fed back to the model as ordinary function results. The default skill set lets the agent read workspace documents and fetch any URL without approval, so a hijacked agent can send private document content out through a URL. Irreversible actions such as sending email are opt-in and gated by approval.

C6 Memory, context & configuration integrity

Minimal 0.05 / 1.00

The default memory tool lets the model store any text into the workspace's vector database with no approval or validation. That text is stored like a document chunk, without a workspace document record, and is then retrieved into later agent sessions and into ordinary chats for every user of the workspace. The separate per-user memories feature is off by default. Isolation is per workspace, not per user.

C7 Third-party extensions

Minimal 0.25 / 1.00

Nothing third-party is enabled by default, and only admins can add extensions. Community Hub downloads are disabled unless the operator sets an environment variable, are then limited to items the hub marks verified, are shown for code review, and arrive inactive. But imported skills are loaded with require() into the server process with all its credentials, there is no hash or signature check on what is downloaded, and MCP servers are user-chosen commands started with the server's full environment.

C8 Secrets & sensitive-data protection

Minimal 0.25 / 1.00

Provider and connector keys live in the server's plaintext .env file, which the compose file mounts into the container. The settings API returns only whether a key is set, not its value. Agent tool calls are logged to the console with their full arguments, nothing masks secrets in logs or model-bound content, and anonymous telemetry, which includes tool names, is on by default. MCP servers inherit every key in the environment when enabled.

C9 Audit & traceability

Minimal 0.38 / 1.00

Every tool call the agent makes, including MCP and imported-skill calls, is printed to the server console with its arguments, and the chat record stores the prompt, final reply and requesting user. Neither is a structured audit trail: tool calls are not saved with the chat, approvals and rejections are not durably recorded, and agent activity is not written to the instance's event log. Scheduled jobs are the exception and keep a per-run trace of tool calls.

C10 Limits & kill switch

Minimal 0.35 / 1.00

Each agent response may chain at most 10 tool calls by default (configurable), after which the model must answer. Stopping the session or closing the socket aborts in-flight LLM requests and stops further turns, but running tool calls are not cancelled. There is no token or spend budget and no wall-clock limit on interactive sessions; scheduled jobs have a five-minute default timeout.