BoundBench

MaxKB

Open-source platform for building enterprise RAG agents and agentic workflows with tools, MCP servers and skills.

github.com/1Panel-dev/MaxKB · 2026-10-05 · 2c7c8c9

Defense-in-depth score

3.1 / 10

Minimal

MaxKB runs all agent and tool code under a dedicated low-privilege account with a cleared environment and a process and network blocking library, which is more care than most agent platforms take. But nothing an agent does needs a person's approval, every published agent gets an unauthenticated public chat link by default, and any agent with a tool also gets a shell with open internet access. The dominant risk is a hijacked public agent leaking knowledge-base data or acting through its tools unattended.

Key gaps (2)

  1. A hijacked agent can leak knowledge-base data and take irreversible actions through its tools with no person involved, and public-link users reach it without logging in. C5 · Untrusted input blast radius
  2. Secret material is not kept out of model-bound requests in the default agent flow. C8 · Secrets & sensitive-data protection

Criteria

C1 Identity & least privilege

Minimal 0.33 / 1.00

Agents act with whatever credentials the builder attaches to each tool, MCP server or nested agent, and those credentials are the same for every person who chats with the agent. Nothing checks a tool call against the chat user who caused it, and each published agent gets a public, unauthenticated chat link by default. Code the agent runs is dropped to a separate low-privilege account with a cleared environment, which narrows what that code inherits, but the platform itself runs as root in a container that also holds the database. The install starts with a documented default administrator password.

C2 Approval gates

Minimal 0.05 / 1.00

There is no approval step for anything an agent does. When an agent has any tool attached, the platform builds it with a shell, file tools, a sub-agent tool and the attached MCP servers, custom tools and nested agents, and explicitly disables interruption for the file tools; no tool call waits for a person. Workflow form nodes collect input from the chat user, who may be the untrusted party, so they do not act as an approval gate. Consequential actions through MCP servers and custom tools happen without a person in the loop.

C3 Tool & action scoping

Minimal 0.45 / 1.00

Some tool inputs are handled carefully: shell commands are split into simple commands and each one is re-quoted before it runs, the file tools are confined to a temporary directory, and the platform's URL fetcher only connects to public addresses and refuses redirects. But the agent's main tools are general purpose (a shell that can run arbitrary Python with internet access, arbitrary MCP servers, builder-written code), and MCP configuration is only checked for its transport. Attaching any single tool to an agent also hands it the shell, file and sub-agent tools.

C4 Code-execution isolation

Moderate 0.53 / 1.00

Every code path the model or a builder can reach (the shell tool, custom tool code, custom tools served over MCP, the remote-MCP proxy and the web crawler) runs as a dedicated low-privilege account with a cleared environment and a preloaded library that blocks new processes, restricted syscalls and connections to a configured host list. This is careful work, but it is a user-and-library boundary inside the same container as the application (running as root), PostgreSQL and Redis, not a separate container or kernel-level sandbox. The sandbox is on by default and switched off by a single environment variable. The shell path has no memory or CPU limits and outbound internet access stays open.

C5 Untrusted input blast radius

Minimal 0.00 / 1.00

Nothing limits what injected content can make an agent do. Agents read public chat users' messages (each agent gets an unauthenticated public link by default), uploaded and crawled documents in the knowledge base, MCP tool results and other agents' replies, and none of it is marked or treated differently from instructions. In the same session an agent can reach knowledge-base data and tool credentials, send data out through the shell's internet access or its tools, and take actions through MCP servers and custom tools, with no person involved.

C6 Memory, context & configuration integrity

Minimal 0.25 / 1.00

Knowledge bases are shared by everyone who chats with an agent, and they can be filled from crawled websites, uploaded files and knowledge workflows, so poisoned content persists and is retrieved for every user, including for agents that hold tools. Long-term memory is off by default; when enabled it is an LLM-written summary per agent and chat user, stored without validation and substituted into the system prompt. Memory is namespaced per agent and chat user in the database queries. There is no workspace or repository configuration that the agent auto-loads.

C7 Third-party extensions

Minimal 0.42 / 1.00

Tools from the official tool store are downloaded over HTTPS from a single allowlisted vendor host, deserialized with a class allowlist and added switched off, so a builder has to enable them. MCP servers are any URL a builder enters, unpinned, and their tool definitions are fetched fresh every session with no change detection. Third-party code runs in the same sandbox account with a cleared environment, but there is no per-extension isolation and no hash or signature check on store downloads.

C8 Secrets & sensitive-data protection

Minimal 0.25 / 1.00

Model and tool credentials are encrypted at rest and masked in the UI, but the decryption key pair lives in the same database. Sandboxed processes get a cleared environment, and the remote-MCP proxy deliberately strips exception text so headers and URLs do not leak. There is no telemetry and the image logs at INFO, though the code defaults to DEBUG when unset and debug logs are not redacted. Secret material is not kept out of model-bound requests in the default agent flow, which caps this criterion.

C9 Audit & traceability

Minimal 0.38 / 1.00

Calls to library tools are saved with their input and output, chat records keep each conversation and its node details per chat user, and administrative actions go to an audit log with user and IP. But shell commands, file operations and sub-agent calls are only visible inside the answer text or debug logs, tool records carry no actor, and everything lives in the application database that the platform process can change. Records are written after the action runs.

C10 Limits & kill switch

Minimal 0.45 / 1.00

Agent runs are capped at 100 graph steps by default, tool code times out after an hour, the shell tool after two minutes, and public-link users have a daily request cap. There is no token or cost cap, sub-agents start a fresh step budget, background shell jobs can outlive the request, and the chat API has no stop endpoint that cancels in-flight work.