C1 Identity & least privilege
Minimal 0.28 / 1.00
Model-driven tools run inside throwaway containers that hold no credentials and run as a non-root user, and the API checks project ownership on every agent-task route. The backend itself is broadly privileged, though: it mounts the Docker socket, holds the database, LLM and Git credentials, and a few host-side tools run with that authority. Default account and key handling in the shipped deployment is not locked down, so least privilege depends entirely on operator hardening.
C2 Approval gates
Minimal 0.05 / 1.00
There is no approval step. The verification agent chooses its own sandbox commands, code, and HTTP requests and executes them immediately, and the same holds for scanner tools. Results are limited by the container, but outbound HTTP requests to arbitrary hosts and all sandbox actions proceed unattended.
C3 Tool & action scoping
Minimal 0.25 / 1.00
Tools take typed arguments, but validation is thin. The command tool relies on a first-word prefix check that allows interpreters, shells and curl, the HTTP tool accepts any URL, and argument and path handling in several tools is not a strict boundary. Only the file read, list and search tools check path containment. Every tool is enabled for its agent by default.
C4 Code-execution isolation
Moderate 0.57 / 1.00
Model-written code and commands run in a fresh Docker container per call: non-root user, read-only root filesystem, no-new-privileges, a list of dropped capabilities, memory and CPU limits, and network off by default; if Docker is unavailable the tools fail rather than run on the host. The gaps are that the capability drop is a 12-item list rather than all capabilities, there is no explicit seccomp profile or process limit, two environment variables can switch protections off silently, and the HTTP tool and scanners attach a bridge network with unrestricted egress. The backend that launches the containers mounts the Docker socket.
C5 Untrusted input blast radius
Minimal 0.00 / 1.00
Repository files, scanner output and retrieved code are all appended to the conversation as ordinary Observation messages, with no provenance marking, no tainting and no change in what tools are allowed after untrusted content is read. A hostile repository being audited can therefore steer the agent, which can fetch arbitrary URLs and send arbitrary-method requests from the sandbox without any human involved.
C6 Memory, context & configuration integrity
Minimal 0.45 / 1.00
The only persistent store is a per-project vector index built from the audited repository and queried through RAG tools, with file path and line shown for each hit. The model has no memory-write tool and the backend does not auto-load instruction or settings files from the audited repository. Index contents are still unvalidated repository text that persists across runs and feeds tool-using agents.
C7 Third-party extensions
Minimal 0.30 / 1.00
DeepAudit loads no plugins or MCP servers and no model files that execute code. It does fetch scanner rule packs from a public registry at run time inside the scanner containers, unpinned, and the model can supply the rule source. Those containers hold no credentials and mount the project read-only, which limits what a malicious rule source could do, though egress is unrestricted.
C8 Secrets & sensitive-data protection
Minimal 0.30 / 1.00
Per-user provider keys, Git tokens and SSH private keys are encrypted at rest, and sandbox containers receive no secrets. However key management and API exposure of server-level credentials are not locked down, and nothing redacts secrets in audited code before it is sent to the LLM provider. Debug output prints key suffixes.
C9 Audit & traceability
Moderate 0.55 / 1.00
Each tool call is recorded as an event with arguments, a truncated result, duration and timestamp in the application database, written by the backend rather than by anything the model controls. Events carry the task but no separate approver or agent identity fields beyond the phase, outputs are cut to short previews, and a failed write is only logged while the action continues.
C10 Limits & kill switch
Minimal 0.45 / 1.00
Each agent loop has a fixed iteration cap, and tools and sub-agent runs have enforced timeouts. The task-level timeout is accepted and stored but never read by the runner, the token budget is defined but not used, and the model chooses the sandbox timeout. Sub-agents can be re-dispatched and get fresh counters each time. Cancel sets a flag and cancels the asyncio task.