C1 Identity & least privilege
Minimal 0.35 / 1.00
The agent's GUI actions land inside a sandbox guest that, by default, holds no credentials of the user's, which keeps the model away from the operator's own accounts. But the framework has no authorization layer of its own: developer-registered function tools run in the host Python process with whatever the process can reach, and the guest itself runs as root. Access control on the optional playground server is not locked down.
C2 Approval gates
Minimal 0.05 / 1.00
There is no human approval step anywhere in the agent loop: every click, keystroke, URL visit and registered function call the model emits is executed immediately. When OpenAI's computer-use model raises its own pending safety checks (for example, a suspected malicious instruction on screen), the framework acknowledges them automatically; the confirmation hook is left as a TODO. The only thing limiting consequences is that the default sandbox is thrown away at the end, which does not undo anything the agent did on the web.
C3 Tool & action scoping
Minimal 0.10 / 1.00
The tools are general by design: arbitrary keystrokes and text into a full desktop, clicks anywhere, and a browser tool that visits any URL. The only argument check is that the call's parameter names match the Python signature; values (text, keys, URLs, coordinates) are passed through unvalidated, and every action is enabled by default. Damage is scoped to the sandbox guest, but inside it the agent can do anything a root user at the keyboard can.
C4 Code-execution isolation
Moderate 0.57 / 1.00
This is the project's strongest area. The documented computer is a sandbox that, by default, runs under gVisor when the container engine has it, and Cua's own Mac runtime ships with gVisor built in; it mounts no host directories, publishes ports only on 127.0.0.1, sets CPU and memory limits and is destroyed when the run ends. Weak points: if gVisor is missing the SDK falls back to a plain runc container with only a log warning, the guest runs as root with unrestricted outbound network, and developer-registered function tools run on the host.
C5 Untrusted input blast radius
Minimal 0.00 / 1.00
A computer-use agent reads whatever is on screen, including web pages an attacker controls, and the framework gives that content the same standing as the user's request. Nothing marks it as untrusted, and nothing limits what the agent may do after seeing it: it can browse to any URL (an exfiltration channel) and submit forms or messages without a person in the loop. The model provider's own safety checks are acknowledged automatically.
C6 Memory, context & configuration integrity
N/A · full credit 1.00 / 1.00
The agent SDK keeps no memory between runs: no memory store, vector index or saved summaries are read back into the model's context, and it auto-loads no instruction or settings files from the environment it operates on. The agent acts inside a sandbox and cannot write the host files that configure it. The shipped CLI and Gradio UI call load_dotenv() on the host, which matters only if the user's own project directory is untrusted.
C7 Third-party extensions
Minimal 0.20 / 1.00
By default the agent calls a hosted model and loads no third-party code. Local Hugging Face models are downloaded unpinned on first use, and trust_remote_code defaults to off in the API. However, the Moondream3 loop hard-codes trust_remote_code=True for an unpinned repository, and the shipped CLI turns it on for every model, so selecting those paths runs remote Python inside the agent's own process with all its keys.
C8 Secrets & sensitive-data protection
Minimal 0.25 / 1.00
API keys come from environment variables or constructor arguments and are never placed in the model's context or the sandbox. Product telemetry (PostHog and OpenTelemetry to Cua) is on by default but sends only counts, model names and action types, not prompts or screenshots. There is no secret redaction (trajectory saving also does not keep credentials out), and the PII anonymization callback is an unimplemented stub.
C9 Audit & traceability
Minimal 0.40 / 1.00
An optional trajectory saver writes a structured, per-turn record of every model call, computer action, function call and screenshot to disk, flushed as each action completes. It is off unless the developer passes trajectory_dir, has no actor attribution or approval records (there are no approvals), and is written to a local directory the host process can alter. The telemetry that is on by default is product analytics, not an audit trail.
C10 Limits & kill switch
Minimal 0.15 / 1.00
The agent loop runs until the model stops producing actions; there is no step cap, no wall-clock limit and, by default, no cost cap. An optional budget callback stops the loop once accumulated provider cost exceeds a set amount, and each model request times out after 120 seconds. There is no cancel primitive beyond abandoning the async generator, so a runaway agent loops and spends indefinitely in the default configuration.