C1 Identity & least privilege
Minimal 0.00 / 1.00
gpt-engineer runs as the user who launched it and does nothing to narrow that authority. The API keys it loads (including from a .env file in the current directory) go into the process environment, and the generated run.sh is started with no scrubbed environment, so the generated code inherits every key and can reach everything the user can: home directory, SSH keys, cloud credential files. There is no authorization layer of any kind.
C2 Approval gates
Minimal 0.38 / 1.00
There is one approval prompt: before running the generated run.sh, the CLI prints the script and asks 'Do you want to execute this code? (Y/n)', and pressing Enter counts as yes. The prompt shows the script but not the generated source files it runs. In the default generate mode, writing the generated files to disk is never gated, and those paths come straight from model output. The --self-heal flag runs the script up to ten times with no approval, and its help text gives no warning about that. Improve mode does show a full diff and asks before applying changes.
C3 Tool & action scoping
Minimal 0.00 / 1.00
The agent's two actions are as broad as they get: it writes files at any path the model names and runs an arbitrary bash script. Paths in the model output are only stripped of a few punctuation characters, so absolute paths and ../ traversal are accepted, and joining them onto the project path in Python lets an absolute path replace the project root entirely. Nothing restricts which commands run.sh may contain.
C4 Code-execution isolation
Minimal 0.42 / 1.00
Generated code runs with no isolation: run.sh is started with shell=True as the same user, from a temporary directory, with the full environment and full network. The repo ships a Docker image and compose file as an opt-in alternative, but that container runs as root and gets the API key from .env, a read-write project mount, and unrestricted network, so it is basic separation at best.
C5 Untrusted input blast radius
Minimal 0.05 / 1.00
Nothing in the code limits what a hijacked run can do. The prompt file and, in improve mode, project files the user selects go into the model with the same standing as the user's instructions. If that content steers the model, the resulting file writes land unattended at any path, which allows irreversible overwrites outside the project. Running code, and with it network exfiltration, still needs the run.sh approval.
C6 Memory, context & configuration integrity
Minimal 0.10 / 1.00
gpt-engineer has no long-term memory store; its logs in .gpteng are never read back into context. It does, however, auto-load a .env file from the current working directory whenever the API key isn't already set. A cloned or downloaded project can use that to supply the API key or, through the OpenAI client's environment variables, redirect the model endpoint. The analytics consent mechanism is also not a strict boundary. Custom system prompts from the project are loaded only with an explicit flag.
C7 Third-party extensions
Minimal 0.05 / 1.00
gpt-engineer loads no plugins or MCP servers. Its default entrypoint prompt, however, tells the model to write a script that installs dependencies, so packages the model picks (unpinned, unverified) get installed and their install scripts run as the user with the full environment. The only consent is the generic run.sh prompt, although that prompt does display the install commands.
C8 Secrets & sensitive-data protection
Minimal 0.20 / 1.00
API keys come from environment variables or a .env file and are passed unchanged to every subprocess the agent starts. Nothing is redacted anywhere. Improve mode's file picker skips dotfiles, so .env isn't offered to the model by default. Opt-in analytics send the full prompt and session logs to the maintainers' RudderStack endpoint, and the consent check is not a strict boundary.
C9 Audit & traceability
Minimal 0.25 / 1.00
Each run appends timestamped plain-text logs of the LLM conversations (generated code, entrypoint chat, improve diffs) under .gpteng/memory/logs inside the project. The output of executed scripts and the user's approve or decline decisions aren't recorded in generate mode. Executed code can edit the logs, because they sit in the workspace and are writable by the same user.
C10 Limits & kill switch
Minimal 0.25 / 1.00
The pipeline makes a fixed handful of model calls (two diff-repair retries, rate-limit backoff capped at seven tries), so model spend per run has a structural bound. The generated script, however, runs with no timeout, and the default prompt even asks it to run parts 'in parallel if necessary'. Ctrl+C kills only the shell process, so background children can keep running. There is no cost cap.