BoundBench

Bifrost

High-performance AI gateway unifying 20+ LLM providers behind an OpenAI-compatible API, with an MCP gateway and agent mode for tool calling.

github.com/maximhq/bifrost · 2026-10-04 · 3b31be0

Defense-in-depth score

4.5 / 10

Minimal

Bifrost's MCP agent mode is cautious where it counts: no tool runs unattended unless the operator lists it, and model-written code-mode scripts run in a capability-limited Starlark interpreter. The dominant risk is the default deployment posture: admin and inference auth are off, the Docker image listens on all interfaces, and every caller shares the gateway's provider keys and MCP credentials. Native plugins load in-process without integrity checks, and secrets are stored in plaintext unless an encryption key is set.

Key gaps (1)

  1. Native .so plugins are dlopen'd into the gateway process (optionally downloaded from a URL) with no hash or signature check, and stdio MCP servers inherit the gateway environment, so a malicious extension gets every provider key. C7 · Third-party extensions

Criteria

C1 Identity & least privilege

Minimal 0.47 / 1.00

Bifrost holds the operator's provider API keys and each MCP server's credential, and by default every caller shares them. Out of the box neither the management API nor the inference API requires authentication, and a request with no virtual key gets no per-caller tool restriction at all. Virtual keys with per-key MCP tool grants, per-user OAuth, per-user headers and RFC 8693 token exchange exist and are good designs, but all are opt-in. Several dangerous registrations (stdio servers, native plugins, env/vault secret references, private-network MCP targets) are refused while auth is off, which limits what an unauthenticated caller can widen.

C2 Approval gates

Moderate 0.53 / 1.00

Bifrost never auto-runs an MCP tool unless the operator lists it in tools_to_auto_execute; that list is empty by default, unknown tools are treated as not auto-executable, and the same rule is re-checked inside code-mode scripts at the exact point each nested tool is called. Anything not auto-approved is returned verbatim to the calling application, which must execute it explicitly. Bifrost does not itself provide a human approval step: whether a person sees the call is up to the application, and there is no argument-level policy. With dashboard auth off by default, any caller who can reach the management API can set the auto-execute list to '*', and there is no undo for tool actions.

C3 Tool & action scoping

Minimal 0.38 / 1.00

Scoping is done at the tool-name level: each MCP client exposes only the tools in tools_to_execute (deny-by-default), per-virtual-key grants and request headers can only narrow that set, and the execute path re-applies the same filters so a hidden tool can't be called by name. Tool arguments are passed to the MCP server without validation against bounds or allowlists. HTTP MCP targets registered without admin auth are pinned to public addresses on every dial.

C4 Code-execution isolation

Moderate 0.65 / 1.00

The only place model-written code runs is opt-in code mode, which executes scripts in an embedded Starlark interpreter (a memory-safe Go runtime) whose only capabilities are the MCP tool functions Bifrost injects; the package imports no OS, process or network modules, and scripts are cancelled when the execution timeout fires. There is no unsandboxed fallback. The interpreter runs inside the gateway process, so a hypothetical escape would land in a process that holds every provider key and full network access. Stdio MCP servers run as host subprocesses but their commands are operator-configured, not model-generated.

C5 Untrusted input blast radius

Minimal 0.45 / 1.00

MCP tool results and upstream servers' instructions go straight into the model's context (server instructions as a system message), with no provenance tagging or taint tracking. What limits a hijacked session is the auto-execute allowlist: the agent loop only runs operator-listed tools unattended and hands everything else back to the application. That holds regardless of what content was read, so it does not stop an auto-listed tool from being used for exfiltration. Shared MCP credentials mean a hijack in one user's session can reach data visible to the shared identity.

C6 Memory, context & configuration integrity

Moderate 0.50 / 1.00

The model has no memory tool and Bifrost auto-loads no workspace files: configuration comes from the operator's app directory or database, and there is no .env loading. The persistence the model can influence is the opt-in semantic cache, which stores responses (including tool calls) and serves them to later requests sharing a caller-supplied cache key until the TTL expires. Cached entries are not validated or provenance-tagged.

C7 Third-party extensions

Minimal 0.17 / 1.00

Two kinds of third-party code can run: native Go .so plugins, which are downloaded from a URL or read from disk and dlopen'd into the gateway process with no hash or signature check, and stdio MCP servers, which run whatever command the operator configures (often an unpinned npx package) as a subprocess that inherits the gateway environment. Neither is enabled by default, and both are refused while dashboard auth is off, so only an authenticated admin (or config.json) can add them. Once added, a malicious plugin has everything the gateway has.

C8 Secrets & sensitive-data protection

Minimal 0.40 / 1.00

Secrets are typed (SecretVar) and redacted when configuration is read back through the API, can reference env or vault values, and internal errors are sanitized before reaching clients. Encryption at rest is only on when an encryption key is set; otherwise provider keys and MCP credentials sit in plaintext in the local database. Prompt and response content logging is on by default. Nothing is sent to a vendor telemetry service, and secrets are never placed in model context, but stdio MCP subprocesses inherit the gateway's environment.

C9 Audit & traceability

Minimal 0.45 / 1.00

Logging is on by default and records every MCP tool execution through the plugin hooks, with the virtual key and user that made the call and a link back to the originating agent request. Records go to a local SQLite (or Postgres) log store via a batched asynchronous writer, outside any workspace. There is no approval record (approval lives in the calling application), no tamper evidence, and with dashboard auth off by default anyone who can reach the API can delete MCP logs.

C10 Limits & kill switch

Moderate 0.50 / 1.00

The agent loop stops after 10 iterations by default, each MCP tool call has a 30-second timeout, and code-mode scripts are cancelled when their timeout fires. Token and cost budgets and rate limits exist through virtual keys but are opt-in. There is no wall-clock cap on the whole agent run, and cancelling a timeout stops waiting but cannot recall an external action already sent.