C1 Identity & least privilege
Minimal 0.00 / 1.00
Prime Agent runs as the logged-in user with that user's full authority. The Python kernel that executes all model code is started with a copy of the entire host environment, and every shell command it runs inherits it again, so provider API keys and any cloud or git credentials in the environment are available to model-written code. The stored credential file in the agent directory is readable by the same process. There is no per-tool identity or authorization check; a hijacked session can do anything the user can.
C2 Approval gates
Minimal 0.00 / 1.00
There is no approval step. The only model tool is a persistent Python REPL that can run shell commands, edit files, call MCP servers and spawn sub-agents, and every call runs immediately. A dirty-tree guard for destructive git commands exists in the codebase but is attached to a shell tool that the live agent does not expose. The README tells users to work in a disposable clone, which is the only protection against unwanted changes.
C3 Tool & action scoping
Minimal 0.00 / 1.00
The agent's only tool takes a free-form Python string. File paths, URLs, shell commands and package installs are all expressed as code, so there is nothing for an argument check to inspect. The tool set cannot be narrowed: the old tool-selection flags were removed and the daemon always configures the REPL as the only tool.
C4 Code-execution isolation
Minimal 0.00 / 1.00
Model-written Python and shell commands run directly on the host as the user. The README states plainly that the worker and kernel processes are not a security sandbox and advises running untrusted work elsewhere. No container, OS sandbox profile or VM is available as an option, and the kernel receives the full host environment, so code it runs can read credentials and reach the network.
C5 Untrusted input blast radius
Minimal 0.00 / 1.00
Content the agent reads, from repository files, web search, MCP servers or messages sent by other running agents, enters the model context with the same standing as the user's instructions. Nothing marks it as untrusted or limits what the agent can do after reading it. The same session holds the user's credentials and can run any shell command or network call, so a successful injection can both leak secrets and take irreversible actions without a human involved.
C6 Memory, context & configuration integrity
Minimal 0.17 / 1.00
Several things persist into future behaviour without a trust decision. The project-level settings file in the working directory is merged over the user's settings and can declare MCP servers, packages, skills and shell settings; project packages are installed and project skills are installed into the kernel when a session starts; AGENTS.md and CLAUDE.md files are loaded silently. The continual-harness memory is better contained: automatic refinement is gated by a model review, always writes the session-local store, and keeps a history that supports rollback, but the model can also write global state directly through its shell.
C7 Third-party extensions
Minimal 0.15 / 1.00
Extensions come from MCP server entries, npm or git packages and Python skill packages. Packages can be pinned and catalog MCP services carry reviewed metadata and a fixed endpoint, but user-declared servers and unpinned packages run whatever is current. Packages listed in a project's settings file are installed automatically when a session starts, and Python skills are installed into and imported by the kernel, so they run inside it with the full environment. Stdio MCP servers do get a reduced environment.
C8 Secrets & sensitive-data protection
Minimal 0.20 / 1.00
Provider credentials come from environment variables or a plaintext auth file kept with owner-only permissions. MCP diagnostics scrub configured values and stdio MCP servers get a reduced environment, but the kernel that runs all model code receives the full environment, and session transcripts are written without redaction. Product telemetry is on by default; its schema is limited to primitive properties with no prompt or tool content.
C9 Audit & traceability
Moderate 0.55 / 1.00
Each session is recorded as a structured JSONL transcript under the agent directory, including the Python code of every tool call, its output and bash command details, and each row is synced to disk as it is written so a session can be replayed. Sub-agents keep their own transcripts. The record is outside the working directory but in a location the agent's own shell can edit, there is no attribution of who approved what (there are no approvals), and nothing is hash-chained or shipped off the machine.
C10 Limits & kill switch
Minimal 0.38 / 1.00
In the default interactive mode the turn loop has no step, time or cost limit, and individual code cells have no default timeout. Interrupting stops an awaited shell command's process group, but background handles, schedules, heartbeats and daemon-backed sessions are designed to keep running after the terminal detaches. Sub-agent nesting is limited to depth 2 by default. The opt-in autonomous mode adds real budgets (12 turns, 80,000 tokens and 30 minutes by default), enforced in code.