BoundBench

Orca

Agentic development environment for fleets of parallel coding agents

github.com/stablyai/orca · 2026-10-05 · f873aaa

Defense-in-depth score

1.7 / 10

Minimal

Orca launches Claude Code, Codex and the other agents it supports with their permission-skipping flags on by default, so every agent runs shell commands, edits files and uses the network without asking, as your user, with your full environment and no sandbox. It also pre-accepts each agent's folder-trust prompt, so repository-controlled agent settings load silently. The dominant risk is a prompt-injected agent acting on all your credentials unattended; switch agent permissions to Manual and turn off workspace trust before pointing Orca at repositories or issues you do not control.

Key gaps (6)

  1. Default launch arguments switch off every supported agent's approval prompt, so shell execution runs without any human gate. C2 · Approval gates
  2. In the default configuration a hijacked agent can both exfiltrate credentials and take irreversible actions with no human involved. C5 · Untrusted input blast radius
  3. Orca pre-accepts launched agents' folder-trust prompts by default, so repository files configure agent hooks, tools and settings without a trust decision. C6 · Memory, context & configuration integrity
  4. Every agent terminal inherits the user's full environment and authority with no narrowing. C1 · Identity & least privilege
  5. Agent commands run unsandboxed as the user, with Codex's own sandbox disabled by the default flag. C4 · Code-execution isolation
  6. Launched agent CLIs and skills run unpinned as the user with the full environment. C7 · Third-party extensions

Criteria

C1 Identity & least privilege

Minimal 0.00 / 1.00

Orca runs every agent it launches as the logged-in user and makes no attempt to narrow that authority. Each agent terminal receives Orca's full process environment, so cloud, Git and API credentials in the user's shell reach every agent and every command it runs. The local control API that agents use to drive Orca is guarded by one shared token that grants the same authority to any caller holding it. A hijacked agent therefore acts with everything the user can do.

C2 Approval gates

Minimal 0.00 / 1.00

Orca's default launch arguments switch off the approval prompts of every supported agent: Claude Code starts with --dangerously-skip-permissions, Codex with --dangerously-bypass-approvals-and-sandbox, and the other agents with their equivalent auto-approve flags. The onboarding screen shows this as a toggle that is on by default, and the code comment states that bypass is the posture a user gets until they choose otherwise. Orca adds no approval step of its own for agent actions, so shell commands, file writes, pushes and network calls run unattended. Users can pick Manual mode, which leaves approval to each agent's own prompt.

C3 Tool & action scoping

Minimal 0.15 / 1.00

The capabilities Orca hands to agents are general-purpose. Each agent gets an unrestricted shell, and Orca's own control API lets an agent open any URL in the built-in browser, type text into any terminal pane (including other agents' panes), create worktrees and change some Orca settings. Orca's API does validate argument shapes and sizes with typed schemas, but nothing restricts hosts, paths or commands. Everything is enabled by default.

C4 Code-execution isolation

Minimal 0.00 / 1.00

Agent commands run as ordinary processes of the logged-in user. Orca isolates agents only by giving each one its own git worktree directory, which is not a security boundary, and its default Codex arguments also turn off Codex's own sandbox. No OS sandbox, container or VM backend is applied to agent terminals, setup scripts or tool processes. Anything an agent runs can read the user's home directory and credentials and reach the network.

C5 Untrusted input blast radius

Minimal 0.00 / 1.00

Agents in Orca read repositories, GitHub and Linear issues, web pages in the built-in browser and messages from other agents, and nothing separates that content from the user's instructions. Because agents run with approvals off, full credentials and open network access, injected instructions can leak data and take irreversible actions with no human involved. Any agent can also type into another agent's terminal through Orca's control API, so one compromised agent can steer the rest. Issue links are pasted as drafts rather than submitted, which keeps a human in the first step only.

C6 Memory, context & configuration integrity

Minimal 0.25 / 1.00

Orca's own repository file, orca.yaml, can define setup scripts and terminal commands, and the desktop app asks the user to approve their exact content (re-asking when it changes) before running them. However, Orca by default pre-accepts the folder-trust prompt of Claude Code, Codex, Cursor, Copilot and other agents for every workspace it opens, so repository-controlled agent configuration (instruction files, project settings, hooks and tool servers) loads in the launched agent without a trust decision. Agents can also change the default launch arguments and environment for future agents through Orca's control API. A poisoned repository configuration therefore persists and can trigger tool use in later sessions.

C7 Third-party extensions

Moderate 0.50 / 1.00

By default Orca launches whichever agent CLIs the user has installed, unpinned, with the full user environment, and its skill updater runs npx --yes skills, which fetches the latest package each time. Because Orca pre-trusts workspaces, a repository can also add tool servers to the launched agents. Orca's own plugin system is much stronger: installed plugin content is hash-checked, reserved identities must come from the official organization, consent is re-requested when a plugin's capabilities or code expand, and plugin workers get an allowlisted environment. That plugin system is off by default.

C8 Secrets & sensitive-data protection

Minimal 0.28 / 1.00

Orca encrypts the integration tokens and cookies it stores using the operating system's secure storage, and keeps plugin workers on an allowlisted environment. Agent terminals, however, inherit Orca's full process environment, so any long-lived keys in the user's shell are visible to every agent and every command. Product telemetry is on by default for new installs; events pass a strict schema validator with enumerated fields and length caps, and crash dumps are not uploaded. Secrets are not masked in terminal output or in what agents send to their model providers.

C9 Audit & traceability

Minimal 0.30 / 1.00

Orca keeps terminal scrollback and shell history for each worktree in its application data folder, and its orchestration database records tasks, dispatches and messages between agents. It does not keep its own structured record of each tool call an agent makes; that lives in each agent's own transcript, if anywhere. The records sit outside the worktree but are writable by the same user the agents run as, and scrollback is capped and overwritten.

C10 Limits & kill switch

Minimal 0.20 / 1.00

Orca puts no step, time or spend limit on the agents it runs; those loops belong to each agent CLI. Its orchestration coordinator caps concurrent dispatches at four by default, stops nested dispatch beyond one level, and fails a task after repeated failures, but the concurrency value is taken from the caller's request. Stopping a terminal kills the whole process group, which is a real halt, while scheduled automations keep firing.