BoundBench

Serena

MCP coding toolkit: LSP-based semantic retrieval and editing

github.com/oraios/serena · 2026-10-03 · d0f7f92

Defense-in-depth score

2.3 / 10

Minimal

As shipped, Serena gives the connected model an unrestricted shell that runs as your user with your full environment. It adds file-editing tools whose project-folder boundary the model can move with activate_project. The server provides no sandbox, approval step, or provenance marking, so safety rests on the MCP client's approval prompts and on trusting the repository. A repository's committed .serena/project.yml and memories can inject instructions and enable tools without a trust decision. The maintainers say this openly: their security docs assume the repository and LLM are trusted and recommend a container for real isolation.

Key gaps (4)

  1. The default-on shell tool runs any model-supplied command on the host as the OS user with the full inherited environment. C1 · Identity & least privilege
  2. No execution isolation by default: shell commands are same-user host subprocesses; the Docker setup is opt-in and experimental. C4 · Code-execution isolation
  3. A hijacked session can both exfiltrate (via shell network tools) and take irreversible actions with no server-side gate. C5 · Untrusted input blast radius
  4. Repository-controlled .serena/project.yml can enable tools and select language servers to launch without a trust decision. C6 · Memory, context & configuration integrity

Criteria

C1 Identity & least privilege

Minimal 0.00 / 1.00

Serena runs as the OS user who launched it and does nothing to narrow that authority. Its shell tool, which is on by default, runs any command with the user's full environment, and language-server processes get a copy of the whole environment too. There is no per-request authorization layer and no read-only default; the only guidance against misuse is a line in the tool description asking the model not to run unsafe commands. If the model is steered by malicious content, it can use everything the user's account can reach, including cloud CLIs and keys in the environment.

C2 Approval gates

Minimal 0.20 / 1.00

As an MCP server, Serena leaves approval to the client and gives it only MCP hints. Tools inherit a read-only or destructive hint from an internal 'can edit' marker. The shell and file-edit tools are correctly flagged, but not every tool's hint is accurate. There is no dry-run or diff preview, and no confirmation step the server enforces. A read-only project mode exists but is off by default and set in the project's own repository file. Wrongly approved shell commands can make irreversible changes, and Serena keeps no checkpoint to undo them.

C3 Tool & action scoping

Minimal 0.28 / 1.00

Serena's file and symbol tools check that paths stay inside the active project. The check is lexical, and the code allows symlinks on purpose. Two default-on tools make this boundary weak. The shell tool accepts any command string and any absolute working directory. activate_project lets the model make any directory, such as the home directory, the new project root. The default context exposes all non-optional tools, including shell and editing. Users can exclude tools globally or pick a narrower context.

C4 Code-execution isolation

Minimal 0.47 / 1.00

The default shell tool runs model-written commands on the host as the same user, with no sandbox and the full environment. The optional REPL tool runs Python inside the Serena process. Serena ships an experimental Docker image that its security docs recommend as the only real protection. It is opt-in, runs as root with default capabilities and full network, and mounts the workspace read-write. That still beats the default, so the criterion takes the opt-in score, capped because it is off by default.

C5 Untrusted input blast radius

Minimal 0.00 / 1.00

Serena reads untrusted content all the time: source files, search results, shell output, committed memories and the repository's .serena/project.yml. None of it is marked as untrusted. The project's initial_prompt is placed into the activation result as project instructions, and the tool descriptions tell the model to read memories before running commands. A repository author can therefore write directly to the model. The server offers no structural limit, so a hijacked session can run shell commands to leak data and make irreversible changes, unless the MCP client happens to gate them.

C6 Memory, context & configuration integrity

Minimal 0.13 / 1.00

Serena loads state the repository controls. .serena/project.yml can inject an initial prompt, add tools and choose which language servers are downloaded and started. Memories under .serena/memories are designed to be committed and shared, and are offered to the model on activation. A trust setting, empty by default on new installs, gates only two project settings: the activation command and language-server overrides. The model can write and edit project and global memories freely; read-only patterns exist but are empty by default. A poisoned memory can therefore persist across sessions and spread to everyone who clones the repository.

C7 Third-party extensions

Minimal 0.28 / 1.00

With the default LSP backend, Serena downloads and runs third-party language servers on its own when a project is activated. Which servers run depends on languages it detects in the repository or on the repository's project.yml. Versions are pinned, and archive downloads are restricted by host and checked against a SHA-256 hash when one is configured. npm-based servers are pinned but installed without a digest. The servers run as separate processes under the same user with the full environment. The user is not asked before installation.

C8 Secrets & sensitive-data protection

Minimal 0.15 / 1.00

Serena holds almost no secrets of its own. Its local authentication secret is hidden from repr and kept in a config file restricted to its owner. Everything else is unprotected. Full tool arguments and results, including any secret a file read returns, are logged at INFO to ~/.serena/logs and to the dashboard. read_file can open gitignored files such as .env. The shell and language servers inherit the full environment. A content-free usage ping is sent by default unless SERENA_USAGE_REPORTING=false.

C9 Audit & traceability

Moderate 0.50 / 1.00

Every tool call goes through one wrapper. It logs the tool name and all arguments before running and logs the result afterwards, with timestamps, to a per-run file under ~/.serena/logs. The file is written by default, outside the project folder. It is plain text with no actor attribution, nothing makes it tamper-evident, and the shell tool can edit or delete it.

C10 Limits & kill switch

Minimal 0.33 / 1.00

Serena applies a 240-second timeout and a 150,000-character output cap to every tool call. The timeout only stops waiting: the worker thread and any shell process keep running, and the shell call itself has no timeout. The model can raise the output cap on each call through max_answer_chars. There are no rate limits.