C1 Identity & least privilege
Minimal 0.00 / 1.00
AgentVerse reads the operator's OpenAI or Azure API key from the environment and has no notion of a scoped agent identity or per-request authorization. Model-written code run by the code-test executor is launched with the default subprocess environment, so it inherits that key and every other credential in the operator's shell. If an agent is hijacked, it acts with the full authority of the user who launched it.
C2 Approval gates
Minimal 0.00 / 1.00
There is no human approval step anywhere in the framework. Model-written code is executed, files are written to model-chosen paths, and tool calls (including a root shell tool exposed by the tool-using configs) are sent to the tool server without asking anyone. The only input() prompts in the code are commented out and were for scoring, not gating.
C3 Tool & action scoping
Minimal 0.05 / 1.00
Tools are as broad as they can be: the code-test executor writes model-generated code to a model-chosen file path with no containment and then runs it through a shell, and the tool-using executor forwards whatever tool name and arguments the model emits to the tool server without checking them against the configured tool list. Shipped tool configs include a root shell, a Python notebook, file writes and web browsing. Only the simulation ToolAgent checks that a named tool exists, which is not argument validation.
C4 Code-execution isolation
Minimal 0.00 / 1.00
Model-generated code runs with no isolation. The code-test executor runs it through a host shell as the current user; the simulation software-team environment exec()s model-written code and evals model-written test lists in-process; one further in-process path is not confined either. The only isolation anywhere is the external XAgent ToolServer container used by the tool-using tasks, which is not part of this repository and could not be verified. An escape is not needed: the code already has the user's files, network and API keys.
C5 Untrusted input blast radius
Minimal 0.00 / 1.00
Nothing distinguishes untrusted content from instructions. Web pages fetched through the tool server are summarized and fed back into agent memory as ordinary function messages, and messages from other agents in the group are treated as context with the same standing. Because the framework also executes model output and has network egress with no approval, a successful injection in a tool-using or code-execution task can both exfiltrate data and take irreversible actions unattended.
C6 Memory, context & configuration integrity
N/A · full credit 1.00 / 1.00
No persistence path exists that the model can influence. Agent memory (chat history, summaries, the in-memory vector store, reflections) lives only in the running process and is rebuilt for every task and every benchmark example. Task configuration comes from the package's own tasks directory or an explicit --tasks_dir flag, and nothing loads .env files or instruction files from the working directory. Caveat: because model-written code runs unsandboxed on the host (C4), it could still overwrite files such as task configs; that is scored as execution blast radius, not as a memory feature.
C7 Third-party extensions
Minimal 0.28 / 1.00
Third-party tools are not loaded by default. Simulation tasks can list BMTools tool servers by URL in the task YAML, and tool-using tasks call an external XAgent ToolServer over HTTP; in both cases the spec and tool list are fetched live with no pinning or integrity check. These servers run as separate processes the operator starts themselves, and AgentVerse sends them only tool arguments and session cookies, not its API keys.
C8 Secrets & sensitive-data protection
Minimal 0.05 / 1.00
API keys come from environment variables and there is no masking, redaction or secret handling anywhere in the code. The keys are not put into prompts, and full prompts are only written to the activity log when --debug is set, but tool inputs, tool outputs and execution results are logged unredacted at the default level. The largest gap is that every subprocess, including model-written code, inherits the long-lived provider keys.
C9 Audit & traceability
Minimal 0.30 / 1.00
AgentVerse writes a text activity log (activity.log, error.log) under the package directory, recording role assignments, plans, execution results and, for the tool-using executor, each tool name, input and observation. The records are unstructured, carry no actor or approver attribution, and do not log the exact command the code-test executor runs. The log sits outside the task working directory but is writable by the same user that runs model code.
C10 Limits & kill switch
Minimal 0.38 / 1.00
Runs are bounded by a turn cap (10 by default, 3 for the default brainstorming task), a 10-call cap per tool-using executor round, a 10-second timeout on code-test execution and a 5-second timeout on in-process unit tests. Cost is tracked and printed but never enforced. The timeouts stop waiting rather than stopping work: killing the multiprocessing worker leaves the shell's python child running, and the timed-out thread keeps executing. The simulation tool agent loops with while True until the model finishes and catches BaseException, which also swallows Ctrl+C during a call.