C1 Identity & least privilege
Minimal 0.05 / 1.00
Access control on the backend API is not locked down. Every tool runs with the operator's own OpenAI key, Kaggle key, unauthenticated MongoDB/Redis, and any third-party plugin API keys users saved, all as root in the backend container. There is no per-request authorization of tool use.
C2 Approval gates
Minimal 0.00 / 1.00
There is no human approval step anywhere in the backend. Model-generated Python and SQL run as soon as the model emits them, plugin API calls (including Zapier actions and a Netlify deploy endpoint) fire directly, and the web agent returns click and type actions for the user's live browser without a confirmation step in the server. The most powerful action, arbitrary code execution, is ungated by default.
C3 Tool & action scoping
Minimal 0.25 / 1.00
The data agent's main tools accept arbitrary model-written Python and SQL with no argument validation. Plugins are narrower: each one only calls fixed endpoints on a fixed host, and the model's chosen endpoint must exist in the plugin's table; web-agent actions are type-checked. Python is selected by default in the shipped frontend, and the tool set is chosen per request by the client rather than restricted to read-only.
C4 Code-execution isolation
Minimal 0.40 / 1.00
In the default configuration (CODE_EXECUTION_MODE=local, set in both app.py and docker-compose.yml) model-generated Python runs through IPython's run_cell in the backend's own worker process. A code-safety check exists but is not a strict boundary. The code therefore runs as root in the backend container with the OpenAI key in its environment, full network access, unauthenticated MongoDB/Redis, and write access to the application source and every user's files. An optional Docker code-interpreter mode sends code to a separate kernel container, but it is commented out in the compose file, uses an image not in this repo, adds NET_ADMIN/SYS_PTRACE capabilities, and shares the backend data volume.
C5 Untrusted input blast radius
Minimal 0.00 / 1.00
The agents read plenty of content their user did not write: uploaded files and databases, Kaggle datasets, third-party plugin API responses, and arbitrary web pages in the web agent. None of it is marked or treated differently; tool results go straight back into the model context. A hijacked session can run arbitrary Python with network access and the OpenAI key, call write-capable plugins, and drive the user's browser, with no human in the loop; frontend output rendering is not hardened either.
C6 Memory, context & configuration integrity
Minimal 0.07 / 1.00
Conversation history, including tool outputs that may carry injected text, is stored in MongoDB/Redis and reloaded as chat memory each time the conversation continues, with no validation or provenance. Uploaded and generated files persist in the backend data folder and can be reused as grounding data later. Per-user isolation of stored history and files is not a complete boundary. There are no auto-loaded instruction or config files from a workspace.
C7 Third-party extensions
Minimal 0.07 / 1.00
The 232 plugins are vendored Python modules shipped in the repository and loaded in-process with importlib; they are part of the pinned code rather than downloaded at runtime, though each one sends data to a third-party API. At startup the plugin selector downloads the hkunlp/instructor-large embedding model from Hugging Face with no pinned revision or hash, and (through the sentence-transformers/InstructorEmbedding loader) deserializes its weights in the backend process. Runtime pip installs of helper packages are also unpinned. Nothing runs third-party code in a separate or confined process.
C8 Secrets & sensitive-data protection
Minimal 0.15 / 1.00
Secrets come from environment variables and files: the OpenAI key is set in the compose file, the Kaggle key is written into the image at build time, and third-party plugin API keys are stored in plaintext in Redis/MongoDB. The backend API's handling of stored plugin keys is not locked down. Full request bodies and generated code are logged at DEBUG/TRACE by default, and the only protection is loguru's diagnose=False on file sinks. Model-run code can read every secret from the process environment.
C9 Audit & traceability
Minimal 0.33 / 1.00
Each finished chat turn is written to MongoDB with the intermediate steps (tool used, generated code, results) and the final answer, and loguru writes runtime logs. This gives a usable transcript, but it is written only when the stream completes, carries no real actor identity, and lives in an unauthenticated database and log directory that the model's own code can modify.
C10 Limits & kill switch
Moderate 0.50 / 1.00
The agent loop stops after 5 iterations by default and runs in a child process that is terminated after a period without output (90 s for the data agent) or when the user hits stop. There is no wall-clock limit on the loop itself, no token or cost budget, no rate limiting, and the web agent's step loop is driven by the browser extension with no server-side cap. Termination kills the worker process but not anything it spawned.