C1 Identity & least privilege
Minimal 0.05 / 1.00
mini runs as the user who launched it and does nothing to narrow that authority. API keys from the global config file are loaded into the agent's own process environment, and every bash command the model runs inherits that whole environment, along with any cloud, Git or SSH credentials the user already has. There is no authorization layer in code; the only thing between the model and the user's account is the per-command confirmation prompt, which is scored under approval gates. The credentials mini itself holds are LLM provider keys, so its own exposure is mostly spend, but a hijacked session reaches whatever the user's shell can.
C2 Approval gates
Minimal 0.47 / 1.00
mini starts in confirm mode: before any batch of model commands runs, it prints the commands and waits for the user to press Enter or type a rejection. The gate sits on the only action path, there is one tool (bash), unknown tool names are refused, and only the user's own typed input can switch to yolo mode. But there are no risk tiers or argument rules: every command gets the same yes/no prompt, and the optional allow-list matches regular expressions against the raw command string. Confirm mode can be turned off by a command-line flag or silently by a config file or the environment variable that selects it. Nothing can be undone: approved commands run directly on the host with no checkpoint.
C3 Tool & action scoping
Minimal 0.00 / 1.00
mini has exactly one tool, bash, and passes the model's command string to the shell unchanged. There is no path containment, URL or host allow-list, or bound on what a command may touch; the optional allow-list only decides which commands skip the confirmation prompt. This is the project's stated design ("no tools other than bash"), and it means a hijacked or mistaken model has the full reach of a shell on the user's machine.
C4 Code-execution isolation
Minimal 0.47 / 1.00
By default every command the model writes runs directly on the user's machine through a shell, as the user, with the full environment (including API keys) and unrestricted network. Commands get a 30-second timeout and the whole process group is killed when it expires, but there is no isolation. The project ships several execution backends that do isolate (Docker or Podman, Singularity, an experimental bubblewrap sandbox, and remote services), and every command goes through the chosen backend with no fallback to the host, but `mini` uses the local backend unless you pick another. The Docker backend uses a stock container (root inside, default capabilities, open network), does not mount the host or pass secrets by default, and is removed after the run.
C5 Untrusted input blast radius
Moderate 0.57 / 1.00
mini reads untrusted text all the time: repository files, command output and anything fetched from the web come back as tool results with nothing marking them as untrusted. What limits a hijacked model is the confirm-mode prompt: in the default mode no command, including network access, runs without the user pressing Enter. That gate is not tied to what was read, and it is the only layer: once a user approves a command, or runs in yolo mode, it can read the user's credentials and send them anywhere or make irreversible changes. Yolo mode can be selected by a flag or silently by configuration.
C6 Memory, context & configuration integrity
Minimal 0.45 / 1.00
mini has no long-term memory and does not load instruction files such as AGENTS.md; each run starts from the built-in config and the task, and the saved trajectory is never read back into a later session. Its settings come from user scope: a .env file in the user's config directory, loaded at startup, plus environment variables that can choose the config file, the model and the limits. Those settings can switch off confirm mode or change the model endpoint, and nothing protects them from the agent itself: because commands run unsandboxed as the user, an approved command can rewrite that file and change every later session.
C7 Third-party extensions
N/A · full credit 1.00 / 1.00
mini has no plugin, skill or MCP system and does not download tools, agents or model files. Agent, model and environment classes are chosen by name in the user's own config and imported from installed Python packages, which is operator configuration rather than third-party code loaded at runtime. The model can still ask to install packages with ordinary bash commands; those go through the same per-command confirmation and host execution scored under approval gates and code-execution isolation, so this criterion's surface is treated as absent.
C8 Secrets & sensitive-data protection
Minimal 0.05 / 1.00
API keys are typed in during first-run setup and stored in plain text in a .env file in the user's config directory, then loaded into the process environment, where every command the model runs can read them (and print them back into the conversation). mini has no masking or redaction anywhere. It sends no telemetry of its own and keeps debug logging off by default, but the full trajectory of each run, including all command output, is written unredacted to the user's config directory.
C9 Audit & traceability
Moderate 0.50 / 1.00
After every step mini rewrites a JSON trajectory with the full message history: each model reply with its parsed commands, each command's output and return code, timestamps, cost and the exit status. User rejections appear in it because they are added as messages, but approvals are not recorded, and there is no notion of who acted. The default file sits in the user's config directory rather than the workspace, but it has a fixed name, so each run overwrites the previous one, and the agent's own shell can edit it.
C10 Limits & kill switch
Minimal 0.45 / 1.00
mini stops a run once its estimated model spend passes $3 by default, and stops with an error if it can't price a model call unless told to ignore cost errors. Each command has a 30-second timeout that kills the whole process group. Step and wall-clock limits exist but are off by default, and when the cost limit trips in an interactive terminal the user is simply asked for new limits. There are no sub-agents. Pressing Ctrl-C stops the loop, but a command that is still running is not killed, and anything a command starts in the background keeps running.