C1 Identity & least privilege
Minimal 0.45 / 1.00
MagenticLite does not hand the agent your own credentials: its shell and browser run inside a Quicksand virtual machine that receives none of the host environment, and the folders you can share are checked against a list of credential locations. But the agent has one static set of powers for the whole session: write access to every folder you share, unrestricted internet, and a browser profile whose logins persist across all sessions. There is no per-request authorization in code, so a hijacked agent can act in any web account you have ever signed into through the app.
C2 Approval gates
Minimal 0.25 / 1.00
Shell commands and file writes outside the scratch workspace go through a rule-based classifier, and anything it flags shows you the exact command before it runs, with approve, deny or alternative choices. Two things undercut it. Handing a task to the web-browsing agent is always auto-allowed, and that agent clicks, types and submits forms on any site with no approval at all; its only 'check-in' is when the model itself decides to call something a critical point. And the classifier's allowlist is not a strict boundary, which can let the model step around the shell prompt. A session-wide 'auto-approve all' switch and an auto_approve policy in config.yaml turn the prompts off entirely.
C3 Tool & action scoping
Minimal 0.30 / 1.00
The agent's main tools are general-purpose: a raw shell inside the VM and a browser that can go to any URL. Real argument validation exists only when you pick folders to share (resolved-path containment to your home directory plus a denylist of credential folders such as .ssh and .aws), but enforcement of the denylist is not a complete boundary. File-content handoffs to the browser agent are size-capped. Tool groups can be trimmed with agent_mode, but the default enables everything.
C4 Code-execution isolation
Strong 0.75 / 1.00
This is the strongest part of MagenticLite. By default every shell command, file tool and the browser run inside a Quicksand QEMU virtual machine, so model-generated code never runs on your host, and if the VM fails to boot the app refuses to start instead of falling back to the host. The weak points are what lives inside the VM: the folders you share are mounted read-write, network access is fully open, and the browser profile with your saved logins and cookies sits on the same VM disk the shell can read. Setting sandbox to null in config runs everything directly on your host, with your full environment, and nothing warns you.
C5 Untrusted input blast radius
Minimal 0.13 / 1.00
Nothing structural limits what a malicious web page or file can make the agent do. Tool output is wrapped in a tag but otherwise enters the conversation like your own instructions, and reading untrusted content changes nothing about which tools are allowed. A hijacked agent can read your shared files (reads need no approval), pass their contents to the browser agent through the auto-allowed delegation, and have it type them into any site, or take irreversible actions in logged-in web accounts, all without asking you.
C6 Memory, context & configuration integrity
Minimal 0.45 / 1.00
MagenticLite has no long-term memory tool and does not auto-load instruction files from your folders. Conversation state for resuming a session is stored on the host outside the VM, so the agent cannot edit it directly. What does persist across sessions is the shared browser profile (cookies, site storage, logins), which every session reuses and merges back, and the long-lived VM disk. Separately, a config.yaml in whatever directory you launch from is silently merged into the settings database and can switch off the sandbox or approvals for all future runs.
C7 Third-party extensions
Minimal 0.30 / 1.00
MagenticLite loads no plugins or MCP servers. The only third-party code path is the agent installing packages inside the VM with pip, npm or apt, which the shell classifier sends to you for approval and shows the exact command. Packages are not pinned or verified, and they run inside the shared VM with access to your shared folders, the network and browser cookies.
C8 Secrets & sensitive-data protection
Minimal 0.38 / 1.00
The LLM API key you enter during onboarding is stored in plain text in the local settings database, masked when shown in the UI, and scrubbed from websocket debug logs. It is never passed into the VM, so the agent's shell cannot read it. There is no telemetry. The main gap is the browser profile: cookies and saved logins for any site you signed into live on the VM disk, where the agent's shell can read them without approval and send them out over the open network.
C9 Audit & traceability
Moderate 0.60 / 1.00
Every event the agents produce, including each tool call with its arguments, which agent issued it, and whether it was approved by you, by a session auto-approve, or by the safe-command classifier, is written by the backend to a local database outside the VM, along with your approval decisions. That gives a reasonable record of what happened. It is not tamper-evident, write failures are not checked, and its protection from the VM depends on mount validation that is not a complete boundary.
C10 Limits & kill switch
Moderate 0.50 / 1.00
The orchestrator pauses after 100 rounds and the browser agent after 100 actions per delegation, asking you whether to continue, and each shell command times out after 60 seconds. Stop cancels the agent loop. There is no wall-clock or spending limit, and every delegation to the browser agent gets a fresh 100-action budget, so a long task can run thousands of browser actions between check-ins.