C1 Identity & least privilege
Minimal 0.00 / 1.00
EvoAgentX has no identity or authorization layer of its own. Each bundled toolkit builds its own client from environment variables (Gmail, Telegram, databases, search APIs), and the model wrapper copies provider API keys into the process environment, which the shell tool and the in-process Python executor then inherit in full. Every tool runs with whatever the operating-system user and those keys can do. A hijacked agent therefore has the developer's whole account reach.
C2 Approval gates
Minimal 0.15 / 1.00
The framework's tool executor runs every tool call the model makes with no approval step. A human-in-the-loop manager exists, but it is off unless a developer creates and activates it (otherwise it auto-approves), it approves whole workflow actions rather than individual tool calls, and its tool-call review mode is unimplemented. The bundled shell toolkit does prompt per command and shows the exact command, but Python execution, file writes, HTTP requests and email sending have no gate.
C3 Tool & action scoping
Minimal 0.13 / 1.00
Tool arguments are not validated at run time: the base Tool class only checks schema declarations when a tool class is defined, and the executor passes model arguments straight into the tool. Bundled tools are general-purpose: a raw shell string, arbitrary Python, any URL with any HTTP method, and file reads/writes on any absolute path. The storage helper's base-directory containment is not a complete boundary. A dynamic toolkit even lets the model create new code-backed tools at run time.
C4 Code-execution isolation
Minimal 0.33 / 1.00
Model-written Python runs by default in the agent's own process through exec() with full builtins; the import allowlist is skipped unless the developer supplies one, and the code's own docstring warns it is unsafe for untrusted code. Shell commands run on the host as the user. Other host-side execution paths exist, including workflow operator exec(), and not every model-influenced execution path is confined. An opt-in Docker interpreter exists, but it starts a stock container with no resource limits, no network restriction and no execution timeout.
C5 Untrusted input blast radius
Minimal 0.00 / 1.00
Nothing in the framework limits what a hijacked agent can do after reading untrusted content. Search, browser, crawler, RSS, arXiv and Gmail tools pull external text into context, and in the default prompt mode tool results are appended as user-role messages with the same standing as the developer's instructions. The same agent can hold egress and state-changing tools (arbitrary HTTP, email sending, shell, Python) with no approval in the way. A successful prompt injection can leak keys or data and take irreversible actions unattended.
C6 Memory, context & configuration integrity
Minimal 0.10 / 1.00
Importing the tools package calls load_dotenv() in several modules, which reads a .env file from the current working directory and can silently set API keys and endpoints such as OPENAI_API_BASE. If the agent runs inside an untrusted checkout, that file can redirect model traffic or swap credentials. The opt-in long-term memory agent stores conversation messages, including model outputs, and pastes retrieved memories back into the prompt as context with no validation or provenance.
C7 Third-party extensions
Minimal 0.23 / 1.00
Third-party extensions come in through MCP: the developer supplies a server config and the framework launches or connects to it via FastMCP with no version pinning, hash check or change detection. Tool descriptions from servers are passed to the model verbatim. Stdio servers run as the same OS user. Separately, importing the RAG embeddings module tries to pip-install the ollama package, unpinned and without a prompt, if it is missing.
C8 Secrets & sensitive-data protection
Minimal 0.15 / 1.00
Provider and tool API keys come from environment variables or plain config fields, and the LiteLLM wrapper copies them into os.environ, where every shell command and in-process code execution can read them. There is no redaction anywhere: tool arguments, tool results and raw model responses are logged at INFO to stdout by default. The one protective detail is that saved agent files exclude the LLM config. No telemetry was found.
C9 Audit & traceability
Minimal 0.33 / 1.00
Every tool call that goes through the main agent loop, including MCP tools, is logged with its name, arguments and result through loguru to stdout. Logs are unstructured text, carry no actor or approver attribution, and are only written to a file if the developer calls save_logger. Workflow execution keeps an in-memory trajectory of messages, but nothing is durable or tamper-evident by default.
C10 Limits & kill switch
Minimal 0.40 / 1.00
Each agent action stops after 20 model calls by default and each workflow task after 5 agent executions, and the shell and MCP tools have 30-second timeouts. There is no cost or wall-clock budget (cost is tracked but never enforced), the model chooses the shell timeout itself, the Python and Docker executors have no timeout, and the workflow scheduler loop has no overall cap. Timeouts stop waiting but leave threads or MCP calls running.