BoundBench

oh-my-pi (omp)

Terminal coding agent (omp), a fork of the Pi coding agent with sub-agents, MCP, browser, extensions and plugins.

github.com/can1357/oh-my-pi · 2026-10-05 · 8aaf115

Defense-in-depth score

1.9 / 10

Minimal

omp ships with tool approval set to 'yolo', so shell commands, file writes, code evaluation, browsing and sub-agents run without asking, on your machine with your full environment and no sandbox. A prompt injection in anything it reads can therefore leak credentials and take irreversible actions unattended. Starting omp inside a cloned repository also loads that repository's settings, hooks, extensions and MCP servers with no trust prompt. A capable per-call approval system exists: set tools.approvalMode to always-ask, and run omp in a container or VM.

Key gaps (6)

  1. Tools run with the user's ambient identity and full environment, so a hijacked session can use every credential the user has. C1 · Identity & least privilege
  2. The default approval mode is yolo, so the bash tool runs without approval. C2 · Approval gates
  3. Commands run on the host as the user with no sandbox. C4 · Code-execution isolation
  4. A hijacked session can exfiltrate data and take irreversible actions with no human involved. C5 · Untrusted input blast radius
  5. Repository settings, hooks, extensions and MCP servers load without a workspace-trust decision. C6 · Memory, context & configuration integrity
  6. Workspace extension modules are imported into the agent process and MCP servers launched without consent. C7 · Third-party extensions

Criteria

C1 Identity & least privilege

Minimal 0.00 / 1.00

omp runs as the user who launches it and uses that user's ambient authority: every shell command and every MCP server it starts receives the agent's full process environment, including any API keys and cloud or git credentials in it. The only scrubbing removes a narrow set of variables such as git repository-location settings, not credentials. There is no per-tool identity or authorization check in code. A hijacked session therefore acts with everything the user can do.

C2 Approval gates

Minimal 0.25 / 1.00

omp has a real approval system: tools declare read, write or exec tiers, prompts show the command, and users can write ordered allow/prompt/deny rules for bash that are matched per segment of compound commands. But it ships with approval mode 'yolo', which auto-approves every tier, so shell commands, file writes, code evaluation and sub-agents run without asking. Sub-agents are always run in yolo mode, and checkpoints are off by default, so most actions cannot be rolled back. Switching tools.approvalMode to always-ask turns the gate on.

C3 Tool & action scoping

Minimal 0.20 / 1.00

The default tool set includes general-purpose tools: a raw shell, code evaluation, file write and edit, a URL fetcher and a browser. File paths are resolved but not confined to the workspace, and the fetch tool has no host or internal-address restrictions. A built-in list of catastrophic shell patterns only raises an approval hint, and user deny rules are empty by default. Individual tools can be turned off in settings, but out of the box a misused tool can reach the whole machine and network.

C4 Code-execution isolation

Minimal 0.00 / 1.00

Shell commands, code evaluation and MCP servers run directly on the host as the user, with the agent's environment. No sandbox, container or OS-level confinement exists for any execution path; the copy-on-write worktree isolation used for sub-agent tasks separates file changes, not privileges. If model-chosen code is malicious, it gets everything the user has.

C5 Untrusted input blast radius

Minimal 0.25 / 1.00

omp reads web pages, fetched URLs, files, MCP results and MCP server instructions, and all of it enters the model's context with no structural limit on what a hijacked session can then do. The one specific measure wraps page-provided WebMCP content in nonce-delimited 'untrusted' markers, which helps the model but stops nothing. Because tools run unapproved by default, injected instructions can read secrets and send them out or take irreversible actions without a human.

C6 Memory, context & configuration integrity

Minimal 0.10 / 1.00

The long-term memory backends and the learn tool are off by default. But omp performs no project-trust gating, which its own source says: a repository's .omp and .claude settings, hooks, extension modules, MCP servers and instruction files such as AGENTS.md are loaded automatically when you start omp inside it. A cloned repository can therefore change security settings, add tools or start servers without any trust decision, and those changes persist for every session in that checkout.

C7 Third-party extensions

Minimal 0.07 / 1.00

omp loads extension modules, custom tools and hooks as code inside its own process, launches MCP servers as subprocesses with the full environment, and installs plugins from a marketplace without hash or signature checks. Project-level MCP configuration is enabled by default and project extension directories are scanned automatically, so a repository can add running code without a consent step. A malicious extension gets everything the agent has.

C8 Secrets & sensitive-data protection

Minimal 0.25 / 1.00

Stored provider credentials sit in a local SQLite file restricted to the user, and there is no third-party telemetry; OpenTelemetry export only happens when the operator sets OTEL endpoints. A placeholder-based secret obfuscation and outbound credential redaction exist but are off by default. Shell commands and MCP servers inherit the full environment, so any long-lived key in it is reachable by the model.

C9 Audit & traceability

Moderate 0.50 / 1.00

Every session is written as a structured JSONL transcript under the user's agent directory, including tool calls and results, and sub-agents write child transcripts next to the parent. Entries are handed to the OS as they complete, so a crash loses at most in-flight text. The record has no actor attribution, approval log or tamper protection, and the agent's own shell can edit it.

C10 Limits & kill switch

Minimal 0.25 / 1.00

Each tool has an enforced timeout (bash defaults to 5 minutes, up to an hour), sub-agents have a request budget that force-stops at 1.5 times its value, a recursion depth of 2 and a concurrency cap, and stopping kills the shell's whole process group. But the main session has no step, cost or wall-clock limit, and the global tool-timeout ceiling and sub-agent wall-clock limit default to unlimited.