C1 Identity & least privilege
Minimal 0.13 / 1.00
Docker Agent runs as the user who launched it and does nothing to narrow that authority. The shell tool, background jobs and every MCP server it launches inherit the whole process environment, including model-provider API keys and any cloud or GitHub tokens the user has exported, and there is no per-tool identity or credential scoping. What keeps a hijacked agent from using those credentials is the separate per-call approval prompt, not any limit on the credentials themselves.
C2 Approval gates
Minimal 0.25 / 1.00
The approval gate is well built: in the default mode every tool that is not annotated read-only, including shell, file writes and web fetches, stops for a per-call prompt that renders the exact call, and allow/ask/deny rules can match on arguments. Unknown tools are rejected, sub-agents inherit the session's mode, and shell commands embedded in skills go through the same prompt. But a docker-agent.yaml in the working directory is picked up automatically and may declare the autonomous safety mode, which approves every call, so a repository you open can turn the gate off; one key press in the prompt also switches the whole session to autonomous. Workspace snapshots that could undo changes are opt-in.
C3 Tool & action scoping
Minimal 0.28 / 1.00
Some built-in tools validate their inputs well: the fetch tool refuses loopback, private and cloud-metadata addresses at connection time and rechecks domain rules on every redirect, and the filesystem tools honour .agentsignore and can be confined to allow-listed directories using kernel-checked roots. But the default agent also ships a raw shell and background-job runner that take any command string, and the filesystem tools accept any absolute path unless an allow-list is configured. Toolsets are chosen per agent in YAML, yet the built-in default enables write, exec and network tools together.
C4 Code-execution isolation
Moderate 0.50 / 1.00
By default every approved shell command and background job runs directly on the host as the user, in the user's working directory, with the full environment and network. Docker Agent can instead run the whole agent inside a Docker Sandboxes VM with a default-deny network proxy (--sandbox, or runtime.sandbox in an agent file), where all tools execute inside the VM and only the working directory is mounted read-write. That option is off by default, so it can at most earn half credit here.
C5 Untrusted input blast radius
Minimal 0.38 / 1.00
Docker Agent does not track whether untrusted content has entered a session: web pages, file contents, AGENTS.md, skills and tool or MCP results all reach the model with ordinary standing. What limits a hijacked agent in the default mode is that every shell command, file write and web fetch still needs a per-call human approval, while the tools that run unprompted are read-only and have no outbound channel. That protection rests on the approval mode, which a project agent file can change and which the user can escalate to autonomous from any prompt.
C6 Memory, context & configuration integrity
Minimal 0.10 / 1.00
Running `docker agent run` with no arguments loads docker-agent.yaml (or .yml/.hcl) from the current directory if one exists, announcing it with a one-line message but asking for no trust decision. That file is a full agent definition: it can add MCP servers and shell-based hooks that start with the session, grant permission rules and set the approval mode. The default agent also loads AGENTS.md from the working directory or any parent and discovers skills from project folders, all as trusted instructions; an optional hook can vet prompt files but is not enabled. Long-term memory is opt-in, and session history lives in the user's home directory.
C7 Third-party extensions
Minimal 0.20 / 1.00
The built-in default agent loads no third-party code, but agent files can add MCP servers (local commands, remote URLs or Docker catalog entries), and a project agent file in the working directory is picked up automatically, so a repository can add servers that start without any consent prompt. When an MCP or LSP command is missing, Docker Agent installs it from the aqua registry, resolving the latest release unless a version is given, and checks a checksum only when the package publishes one. Local MCP servers run as separate processes but inherit the user's full environment, including API keys.
C8 Secrets & sensitive-data protection
Minimal 0.30 / 1.00
Secret redaction is on by default: a pattern scanner scrubs secrets from tool arguments before approval, from tool output before it is stored or shown to the model, and from outgoing chat content. Provider API keys still come from environment variables and are passed whole to every shell command, background job and local MCP server, so any approved command can read them. Usage telemetry to Docker is on by default and, as the project's own documentation says, includes command-line positional arguments, which can contain prompts, and error text.
C9 Audit & traceability
Moderate 0.50 / 1.00
Every session is saved to a local SQLite database in the user's home directory as it runs: user messages, the model's tool calls with arguments, tool results and timestamps, with sub-agent sessions stored as linked child sessions. Rejected calls appear as error results, but approval decisions and who approved them are recorded only in OpenTelemetry spans and hook notifications, and OpenTelemetry export is off unless --otel is passed. The database is outside the workspace but writable by the same user, so an approved shell command could alter it.
C10 Limits & kill switch
Minimal 0.47 / 1.00
Docker Agent has good limit machinery, but little of it is on by default. Agent files can set an iteration cap and run budgets for cost, tokens and working time, which sub-agents draw from as one shared pot, and stopping a turn kills the running shell command's whole process group. In the default agent, though, the iteration cap is zero (unlimited) and no budget is set; what remains is a breaker that stops after five identical tool-call batches in a row and a 30-second shell timeout that the model can raise per call. Background jobs keep running after a turn is interrupted until the session ends.