BoundBench

Coze Studio

Open-source low-code platform for building, debugging and publishing AI agents, apps and workflows.

github.com/coze-dev/coze-studio · 2026-10-05 · fefb05f

Defense-in-depth score

3.0 / 10

Minimal

Coze Studio agents call plugins, workflows and a raw-SQL database tool with no human approval, and nothing limits what injected content can make them do, including sending data to any URL. Workflow code runs in a Deno and Pyodide sandbox by default, but that sandbox lives inside the server container that holds every deployment secret. Credential storage and logging defaults are weak, and registration is open. Read the README's public-network warning and close registration before exposing it.

Key gaps (3)

  1. Untrusted content can drive unattended exfiltration and database writes or deletes in one session. C5 · Untrusted input blast radius
  2. The code sandbox runs inside the server container and inherits its environment, which holds all deployment secrets. C4 · Code-execution isolation
  3. The server identity holds deployment-wide credentials for all tenants' data. C1 · Identity & least privilege

Criteria

C1 Identity & least privilege

Minimal 0.20 / 1.00

The agent acts with the platform server's own authority and with whatever credentials the plugin author attached. Service-token plugins use one static credential for every end user of a published agent; only OAuth plugins act on behalf of the individual user. Resource authorization is a creator-only check, and the README itself lists horizontal privilege escalation in some APIs as a known risk. The server process holds credentials for the shared database, object store and search index of every tenant, so a failure of the authorization layer exposes the whole deployment.

C2 Approval gates

Minimal 0.00 / 1.00

There is no human approval step for any tool call. The ReAct agent invokes plugins, workflows (which can contain code and HTTP nodes) and a raw-SQL database tool directly, and the run handler persists and streams the calls but never pauses for a decision. Workflow question and input nodes exist, but they are placed by the author for data entry and do not gate actions. Write and delete operations on external APIs and on user databases happen unattended.

C3 Tool & action scoping

Moderate 0.50 / 1.00

Agents only receive the plugins, workflows and databases their author explicitly attaches, which keeps the default tool set empty. Plugin arguments follow typed OpenAPI schemas with required-field checks, and model-written SQL is parsed, rewritten onto the bound table, checked against a denylist and a table-name pattern, and filtered per user in the default read-write mode. Outbound HTTP from plugins and workflow nodes accepts any URL, with no internal-address filtering.

C4 Code-execution isolation

Moderate 0.55 / 1.00

Workflow code nodes run in a sandbox by default: Python executes under Pyodide inside Deno with no network, environment, subprocess or FFI permission, file access limited to a module cache, a 60-second timeout and a 100 MB memory cap, and errors do not fall back to host execution. An operator can switch to a local runner that executes Python directly in the server container, through an environment variable or the admin settings, without a warning. The Deno process runs inside the main server container and inherits its environment, which holds the deployment's secrets, so an escape from the runtime sandbox lands next to every credential.

C5 Untrusted input blast radius

Minimal 0.00 / 1.00

Nothing limits what untrusted content can make the agent do. Published agents take messages from end users through the API and chat SDK, and plugin responses, HTTP results and knowledge documents return to the model as plain tool output with no provenance or taint handling. In the same session the agent can read private databases and knowledge, send data to any URL through plugins, and write or delete records, all without a human. On a multi-user deployment a hijacked agent can also touch data and credentials that other users configured.

C6 Memory, context & configuration integrity

Minimal 0.30 / 1.00

Agents can persist state through user variables and user databases that the model writes with tools, and these are fed back into later conversations and can drive further tool calls. Isolation is per user by default: variables are keyed to the end user and connector, and databases default to a mode that adds a per-user filter to queries. Nothing validates, reviews or expires what the model writes, and there is no provenance or rollback. There are no workspace instruction files or repo configs to auto-load, since this is a hosted platform.

C7 Third-party extensions

Minimal 0.40 / 1.00

Extensions are plugins: OpenAPI descriptions of remote HTTP services, either from the operator-curated official catalog or created by users with their own server URL. No third-party code runs inside the platform; MCP invocation is not implemented and the Coze SaaS plugin source is off by default. User-created plugins point at whatever the remote server currently does, with no pinning or change detection, but a malicious plugin only receives the arguments sent to it and its own credentials.

C8 Secrets & sensitive-data protection

Minimal 0.20 / 1.00

Secrets come from environment variables and the database. Plugin credentials are encrypted at rest, but model provider keys are stored without encryption (a source comment leaves this to the deployer), and credential encryption is not hardened in the default deployment. There is no log redaction, and the shipped environment and Helm values set debug-level logging. Frontend monitoring is a no-op stub, so no telemetry leaves the deployment.

C9 Audit & traceability

Moderate 0.50 / 1.00

Every agent tool call and tool response is stored as a message with run, conversation, agent and user identifiers, and workflow runs store each node's inputs, outputs, status, duration and errors in the database. This gives a structured, per-action record outside anything the model can edit. There is no approval record (there are no approvals), no tamper-evident storage or standard export, and records are written alongside execution rather than before it.

C10 Limits & kill switch

Minimal 0.33 / 1.00

The agent loop inherits the agent framework's default step cap because no limit is set, and sandboxed code has a 60-second timeout. Workflows, which agents can call as tools, have no run timeout and no node-count limit by default, and plugin HTTP calls have no timeout. Cancelling an agent run only updates its database status; workflow cancellation is cooperative, checked every 200 ms. There are no token or cost budgets.