BoundBench

Kimi Code CLI

Moonshot AI's terminal coding agent (TUI, print mode, ACP and web server) that reads and edits code, runs shell commands, fetches web pages and delegates to subagents.

github.com/MoonshotAI/kimi-code · 2026-10-05 · 21406fb

Defense-in-depth score

2.7 / 10

Minimal

Kimi Code asks before running shell commands in its default mode and shows the exact command, but it has no sandbox: an approved command runs as you, with your full environment and every credential in it. Web fetches, web searches, reads anywhere on disk and file edits inside a git repository are approved automatically, so a prompt-injected session can send data out without a prompt, and the approval gate can be bypassed in the default configuration. Treat untrusted repositories and web content as able to act with your account, and run it in a container or VM if that matters.

Key gaps (3)

  1. The approval gate can be bypassed in the default configuration. C2 · Approval gates
  2. Shell commands and MCP servers run directly on the host as the user, with the full environment; there is no sandbox. C4 · Code-execution isolation
  3. A hijacked session can send data out through auto-approved web fetches, and an approval-free path to consequential actions exists in the default configuration. C5 · Untrusted input blast radius

Criteria

C1 Identity & least privilege

Minimal 0.05 / 1.00

Kimi Code runs as the logged-in user with no identity of its own and no narrowing of that user's authority. Shell commands and MCP servers are started with the CLI's complete environment, so every cloud key, token and SSH agent socket the user has is available to them. The only thing standing between a hijacked agent and that authority is the per-call approval prompt for shell commands.

C2 Approval gates

Minimal 0.25 / 1.00

The default "Always Ask" mode runs an ordered policy chain before every tool call: reads, web fetches, web searches and subagent launches are approved automatically, file edits inside a git repository are approved automatically, and everything else (shell commands, MCP tools, writes elsewhere) falls through to a prompt that shows the exact command or file content. A parsed check of shell commands escalates a short list of destructive commands, and file edits are snapshotted so they can be undone. The approval gate can be bypassed in the default configuration, which caps this criterion. Switching to the looser modes needs a flag, a slash command or a config value, and the config value is applied silently.

C3 Tool & action scoping

Minimal 0.40 / 1.00

The web fetcher is well built: it only allows http(s), refuses private and loopback addresses after DNS resolution, re-checks every redirect and pins the resolved address. File tools normalize paths and require an absolute path to touch anything outside the working directory, but they then allow it, so reads anywhere on disk are possible. The shell tool takes a raw command string, and Bash, Write and web tools are all on by default, though each can be disabled in the user config.

C4 Code-execution isolation

Minimal 0.00 / 1.00

There is no sandbox. Approved shell commands run as `$SHELL -c` directly on the host as the user, MCP stdio servers are launched the same way, and both receive the CLI's full environment. No container, OS sandbox profile or isolated runtime exists anywhere in the engine, so once a command is approved, it can reach every file, credential and network destination the user can.

C5 Untrusted input blast radius

Minimal 0.13 / 1.00

Kimi Code does not distinguish untrusted content: web pages, search results, repository files and MCP results enter the conversation with the same standing as the user's request, and there is no taint tracking. What limits a hijacked session is the approval prompt on shell commands, but web fetches, web searches and reads of any file outside the sensitive-file list are approved automatically, so data can leave without a human seeing it. An approval-free path to consequential actions also exists in the default configuration, which makes the worst case both data exposure and an unattended state change.

C6 Memory, context & configuration integrity

Minimal 0.17 / 1.00

Settings come only from the user's own config file, and the TUI asks the user to trust a folder before anything else starts, listing project MCP servers (with their commands), project skills, extra directories and AGENTS.md files; declining exits. After that one decision, AGENTS.md and project skills load silently as instructions, and because file edits inside a git repository are approved automatically, the agent can change those instruction files without a prompt and the change is picked up in later sessions. There is no long-term memory store.

C7 Third-party extensions

Minimal 0.30 / 1.00

No plugin or MCP server is enabled on a fresh install. Plugins are installed by the user from a marketplace or any GitHub repository, and the installed archive is pinned to the commit it resolved to; MCP servers are whatever command the user (or a trusted project) configures, with no pinning or integrity check. The trust prompt shows the exact command of project MCP servers, but plugin and MCP processes run as the user with the CLI's full environment.

C8 Secrets & sensitive-data protection

Minimal 0.35 / 1.00

OAuth tokens are stored as 0600 JSON files under ~/.kimi-code (a keyring option appears in the config schema but no keyring code exists), and an API key can sit in config.toml. Log records pass through a key- and pattern-based redactor, telemetry strings are scrubbed of URLs, paths and token-shaped strings, and the file tools refuse to read .env files and private keys. But shell commands and MCP servers inherit every secret in the user's environment, nothing redacts tool output before it goes to the model, and telemetry is on by default (opt out with KIMI_DISABLE_TELEMETRY).

C9 Audit & traceability

Moderate 0.63 / 1.00

Every agent, including subagents, writes an append-only `wire.jsonl` record under ~/.kimi-code/sessions with each message, tool call and result, every approval decision, permission-mode changes and subagent runs. Records are flushed with a durable write per batch, and sessions can be replayed from them. The record sits outside the project folder but in a directory the user's own shell commands can edit, and it carries no separate approver identity or tamper protection.

C10 Limits & kill switch

Minimal 0.40 / 1.00

A per-turn step cap exists but is unset by default, which means unlimited, and there is no token or cost cap unless the user sets a goal budget. Shell commands time out after 60 seconds by default (300 maximum), but a foreground command that hits its timeout is moved to the background rather than killed, and background commands can run for up to 24 hours and are not stopped when the turn is interrupted. Interrupting a turn kills foreground commands by process group.