C1 Identity & least privilege
Minimal 0.05 / 1.00
Every built-in agent, classifier, and storage backend builds its AWS client from boto3's default credential chain, and API keys for Anthropic, OpenAI, Jev, and Dakera are read from options or environment variables. The user_id passed to route_request is a caller-supplied string used only to key chat history; nothing checks what that user may do. Developer-registered tools run in the same Python process and inherit all of these credentials. The framework does nothing to narrow the authority of the deployment it runs in.
C2 Approval gates
Minimal 0.00 / 1.00
The framework has no human approval step. When the model asks for a tool, the tool handler looks it up by name and calls it immediately. The same is true for MCP tools and for the supervisor's send_messages delegation tool. The only mention of confirmation is a system-prompt line asking the supervisor to forward confirmations, which is a prompt and not a control. Any consequential action a developer registers runs unattended.
C3 Tool & action scoping
Minimal 0.10 / 1.00
AgentTool derives a JSON schema from the function's type hints and sends it to the model, but the framework never checks the model's arguments against that schema; it unpacks them straight into the function. MCP tool arguments are forwarded unchanged to the server. Unknown tool names are refused, and each agent only receives the tools registered to it, which narrows reach somewhat. Enforcement of MCP tool visibility does not cover every call path.
C4 Code-execution isolation
N/A · full credit 1.00 / 1.00
The framework itself never interprets model output as code: there is no shell, eval, exec, pickle, or package-install path in the Python package. The only process it can launch is an MCP stdio server whose command the developer writes in code, and that risk is scored under third-party extensions. Code a developer puts inside a registered tool runs unisolated in-process; that is the developer's code, not a framework execution path.
C5 Untrusted input blast radius
Minimal 0.00 / 1.00
Content the user did not write reaches the model with no separation: tool and MCP results come back as ordinary tool-result turns, MCP tool descriptions are copied into the tool list, retriever results are appended to the system prompt, and the supervisor pastes every sub-agent's history into its own system prompt. Nothing tracks whether untrusted content has been read, and no tool is disabled or gated afterwards. Bedrock guardrails can be passed through but are off by default and only detect. A successful injection can therefore drive any registered tool, including ones that send data out.
C6 Memory, context & configuration integrity
Minimal 0.10 / 1.00
Conversation history is stored per user, session, and agent and replayed on every later turn; the default store is in memory, with DynamoDB and SQL backends available. Nothing validates or labels what is saved, retriever output goes straight into the system prompt, and the supervisor injects all its sub-agents' history into its system prompt as trusted text. Per-user isolation on the supervisor path is not a complete boundary. StrandsAgent defaults load_tools_from_directory to True, which makes the Strands SDK load and hot-reload Python tool files from the working directory's tools folder without any trust decision.
C7 Third-party extensions
Minimal 0.17 / 1.00
Third-party code enters through MCPToolProvider, which launches or connects to whatever servers the developer lists in code, and through StrandsAgent's MCP clients and directory tool loading. Nothing pins versions, checks hashes, or asks again when a server's tool list changes. The developer has to write each server command in code, so nothing third-party is on by default, except that StrandsAgent's directory loading picks up any Python file placed in the working directory's tools folder. Stdio servers run as separate processes on the host; the MCP SDK gives them a reduced environment unless the developer passes one.
C8 Secrets & sensitive-data protection
Minimal 0.30 / 1.00
API keys come from constructor options or environment variables and are held in plain attributes, with no masking or redaction helper anywhere in the package. Verbose chat and classifier logging is off by default, and there is no telemetry beyond a content-free feature tag added to the AWS User-Agent header. Tools can return a ToolResult whose structured data and UI never reach the model, a small data-minimisation path. MCP tool errors are passed to the model verbatim, and every in-process tool can read the process environment.
C9 Audit & traceability
Minimal 0.35 / 1.00
The framework records nothing by default. It offers callback hooks around each tool call (name, input, output, agent info) for both its own tools and MCP tools, but the default callbacks do nothing, and the error hook is defined but never called. The orchestrator's metadata carries the caller-supplied user and session ids, but there is no audit log, timestamp, or approver field. A developer who wants a trail must write the storage themselves.
C10 Limits & kill switch
Minimal 0.30 / 1.00
Each agent's tool loop stops after a fixed number of rounds (20 for Bedrock, 5 for Anthropic, 40 for the supervisor) and each model call has a 1,000-token output cap by default. There is no wall-clock limit, no session cost budget, and no way to stop a running request. Sub-agents called by the supervisor each start a fresh budget and run in parallel threads with no limit on how many, so total work can multiply well past any single cap.