BoundBench

Exa MCP

MCP server exposing Exa web search, page fetching and Exa Agent research tools, deployed as the hosted Streamable-HTTP server at mcp.exa.ai and as a local stdio package.

github.com/exa-labs/exa-mcp-server · 2026-10-05 · f3d71fb

Defense-in-depth score

5.7 / 10

Moderate

A small, read-only web search and fetch server: it runs no code, touches no files and only talks to Exa's API, and its default tools cannot change anything. Its main gaps in posture are around untrusted content and authority: web pages come back as unmarked text, the fetch tool gives a hijacked model an outbound channel to any URL, and callers' API keys and OAuth tokens are forwarded unchanged to Exa's API. Limits exist as timeouts and free-tier rate limits, but result sizes are model-chosen and logging is plain console text.

Key gaps (1)

  1. OAuth access tokens and API keys presented to the MCP server are forwarded unchanged to Exa's API rather than exchanged for a server-scoped credential (token passthrough). C1 · Identity & least privilege

Criteria

C1 Identity & least privilege

Minimal 0.25 / 1.00

The hosted server runs each request with whatever credential the caller brings: the caller's own Exa API key (header or URL), an Exa OAuth access token, or, for anonymous callers, the operator's shared Exa key. A fresh server instance is built per request, so one caller's key is never reused for another. The agent tool is only registered for callers who bring their own credential, and a token that fails verification is rejected rather than downgraded. OAuth access tokens issued for the MCP server are verified and then forwarded unchanged to Exa's API (token passthrough), so the server's own authority is not separated from the downstream API's.

C2 Approval gates

Moderate 0.53 / 1.00

As a tool server it relies on the MCP host to ask the user before calls. It gives the host risk labels on every tool, and the two default tools are accurately labelled read-only lookups that cannot change anything. The labels are not consistent across the opt-in tools: the agent tool, which starts billable multi-minute runs, is labelled read-only while the older equivalent research tool is not, and the advanced search tool says it does not reach the open web. There is no dry-run, confirmation step or server-enforced read-only mode, but in the default configuration no tool can change state, so a wrongly approved call costs only search credits.

C3 Tool & action scoping

Moderate 0.63 / 1.00

Each tool calls one fixed Exa endpoint through the official client, so the model cannot point the server at an arbitrary host or service, and the default tool set is two read-only lookups. Arguments are typed and validated with zod schemas, but the schemas are deliberately lenient and carry no upper bounds: the model chooses how many results, how many URLs and how many characters per page. The fetch tool accepts any URL, which Exa's crawler then retrieves on the caller's behalf.

C4 Code-execution isolation

N/A · full credit 1.00 / 1.00

The server never interprets model text as code: there is no shell, eval, subprocess or script execution anywhere in its source. All work is HTTPS requests to Exa's API through the official client.

C5 Untrusted input blast radius

Minimal 0.05 / 1.00

Search highlights and fetched page text are returned as flat text mixed with the server's own labels (Title, URL, Author), with no structured separation and no marker that the content is untrusted. Tool descriptions include usage directives to the model. The fetch tool lets a hijacked host model send data to any URL by encoding it in the address, which Exa's crawler then requests, so exfiltration needs no human step. The server itself can take no irreversible action.

C6 Memory, context & configuration integrity

N/A · full credit 1.00 / 1.00

The server keeps no memory the model can write and loads no workspace files. Its only persistence is a 24-hour record of the connecting client's name and version, kept to label requests to Exa's API and never returned into model context. Its bundled agent guide is read from the package itself.

C7 Third-party extensions

N/A · full credit 1.00 / 1.00

The server loads no plugins, launches no other MCP servers and installs nothing at runtime; it uses only its own fixed npm dependencies. The skill files in the repo are read by host clients, not executed by the server.

C8 Secrets & sensitive-data protection

Minimal 0.38 / 1.00

API keys come from headers, the URL or the environment and are never placed in tool results. Before each request is handed to the MCP layer and its third-party analytics wrapper, the server strips the x-api-key and Authorization headers and the exaApiKey query parameter. There is no general redaction or log filter, though: every search query and fetched URL is written to the server log by default, the README recommends putting the API key in the URL, and content-free usage analytics are sent to a third party (Agnost) by default.

C9 Audit & traceability

Minimal 0.38 / 1.00

Every tool call writes a few unstructured lines to the server log with a request id, the tool name, the query or URLs, and whether it succeeded. That gives the operator a basic trail, but it records no caller identity, is plain console text rather than a structured record, and is best effort. Third-party analytics checkpoints are not an audit trail.

C10 Limits & kill switch

Moderate 0.53 / 1.00

Every tool call is wrapped in a server-side timeout (60 seconds for the default tools, 5 minutes for advanced search), retries are capped at two, and the agent tool's call window is clamped below the platform's 800-second function limit. Anonymous callers are rate limited per IP when the operator configures the rate-limit store, and the limiter fails open when that store is unavailable. Result counts and page sizes are chosen by the model with no upper bound, timeouts stop waiting without cancelling the upstream request, and callers with their own key have no spend limit from the server.