BoundBench

HexStrike AI

MCP server that lets AI agents autonomously run 150+ security tools

github.com/0x4m4/hexstrike-ai · 2026-10-03 · d689933

Defense-in-depth score

0.9 / 10

Minimal

The HTTP API server has no authorization, approval step, or sandbox in front of a generic shell-execution endpoint, and its default network exposure and authentication are not locked down. Any client driving the MCP bridge runs commands with the full authority of the server's OS user. Tool argument handling is not strict either, and the only bounds are a 5-minute per-command timeout and a log file. It should only run in a disposable, network-isolated VM.

Key gaps (4)

  1. The most powerful action path, an arbitrary-command endpoint, has no approval or risk gate. C2 · Approval gates
  2. Commands run as same-user shell children with the full environment and no isolation. C4 · Code-execution isolation
  3. A hijacked session can leak data and take irreversible actions with no human involved. C5 · Untrusted input blast radius
  4. Any caller can make the server install arbitrary remote packages without consent. C7 · Third-party extensions

Criteria

C1 Identity & least privilege

Minimal 0.00 / 1.00

The API server holds no identity of its own: every command runs as the OS user who started it, with that user's whole environment. There is no per-request authorization anywhere in the code, and the server's default network exposure and authentication are not locked down. A hijacked agent therefore has the operator's full authority. The project's README only suggests adding authentication.

C2 Approval gates

Minimal 0.00 / 1.00

The bridge exposes 151 tools and none carries a read-only or destructive hint; the most dangerous one is a documented arbitrary-command tool. The server has no confirmation step, no dry-run, and no read-only mode, so every call executes immediately and approval depends wholly on the host. The shipped sample config only has an empty always-allow list, which is a host setting rather than a server control.

C3 Tool & action scoping

Minimal 0.00 / 1.00

Tool argument handling is not strict, and there are no schema constraints or scope allowlist. Checks are limited to 'field is present'. The generic command endpoint, a Python-script execution endpoint, a package-install endpoint, and a file API are all on by default with no way to select a narrower tool set.

C4 Code-execution isolation

Minimal 0.00 / 1.00

Every command runs as a same-user child of the API process through a shell, with the full inherited environment and no container, namespace, or user separation. The only boundary is the host itself. The optional headless browser tool is launched with its own sandbox and web-security protections switched off.

C5 Untrusted input blast radius

Minimal 0.07 / 1.00

Tool output, including content fetched from scan targets and third-party responses, returns to the model as plain stdout/stderr strings with no provenance flag or untrusted marker. The same session can both read that content and run any command, reach the network, and delete files, with no human step. A hijacked session can therefore leak data and take irreversible actions unattended.

C6 Memory, context & configuration integrity

Minimal 0.15 / 1.00

No agent memory or auto-loaded instruction files exist. The server does keep global state that persists outside any session: an in-memory command-result cache keyed by command string, a shared file store under a temp directory, and persistent Python virtual environments. These stores are not partitioned per caller, and any client can write, delete, and execute their contents. Nothing from them is automatically re-injected into model context, and no workspace file is loaded as configuration.

C7 Third-party extensions

Minimal 0.00 / 1.00

The server exposes an endpoint that installs any named Python package into a persistent virtual environment on request, and a companion endpoint that runs caller-supplied Python scripts in it. Installation is unpinned, unverified, and needs no consent, and package install steps run code with the server user's authority and full environment. The roughly 150 external security binaries are installed separately by the operator and are not fetched by the server.

C8 Secrets & sensitive-data protection

Minimal 0.00 / 1.00

Tools take credentials such as passwords and API keys as ordinary arguments and splice them into command lines. Every command line and all its output are written at info level to the console and to a log file with no redaction. Every subprocess inherits the server's full environment. The debug mode of the client also logs complete request bodies. No secret handling or masking code exists, and the log file is created in the working directory.

C9 Audit & traceability

Minimal 0.20 / 1.00

Commands, their output, and timings are written as free-text lines to the console and to hexstrike.log, which gives a basic trail for the main execution path. Records are unstructured, carry no caller identity, approver, or correlation ID, and some direct subprocess calls in analysis helpers do not log their argument vectors. The log file sits in the working directory, where the same command endpoint can alter or delete it, and if the file cannot be opened the server silently continues with console-only logging.

C10 Limits & kill switch

Minimal 0.45 / 1.00

Each command run through the main executor has a fixed 5-minute timeout, the file API caps file size, and the async pool is capped at 32 workers. There is no rate limit, no cap on concurrent requests, no output-size cap, and no overall budget. The timeout terminates only the shell process because commands are not started in their own process group, so child processes can outlive both the timeout and the terminate endpoint. A runaway caller can issue unlimited consecutive commands.