C1 Identity & least privilege
Minimal 0.25 / 1.00
Ongrid acts on hosts through edge agents installed on every enrolled machine, and those agents run as root. Agent tools do not check whether the requesting user may act on a given device. Messages from the IM bridge (Slack, Telegram, Lark, DingTalk) all run as one shared service account with no role. The shipped default keeps the agent read-only until an admin enables write actions, though that boundary is not strictly enforced. A hijacked agent can therefore reach root-level authority across the fleet, plus the Kubernetes controller's patch and delete rights.
C2 Approval gates
Minimal 0.25 / 1.00
When an admin enables write actions, the classified write tools (cloud_bash, install_skill, restart_service, Kubernetes actions, and host commands whose name matches a short write list) go to a human approval card. Only an admin can approve, and the approved call is exactly what runs. The weak point is host_bash, the most powerful tool: gate coverage for host commands is incomplete, and so is the confirmation step for alert-rule changes. Trusted MCP servers and published workflows run tools without approval.
C3 Tool & action scoping
Minimal 0.40 / 1.00
Most built-in tools are narrow, typed, read-only queries (metrics, logs, topology, incidents, Kubernetes snapshots). The generic host_bash tool is checked on the edge against an allowlist of binaries with subcommand rules, absolute-path checks against a few data directories, and a deny-all network list. That filter is not a complete boundary. MCP tool arguments are passed through unvalidated.
C4 Code-execution isolation
Minimal 0.20 / 1.00
Host commands requested by the model run inside the root edge-agent service on each target machine. The edge filters them (an allowlisted binary executed directly, without a shell), and systemd makes most of the filesystem read-only. But there is no separate sandbox: no dedicated user, no container, no seccomp, and the network is open. The filter is not a complete boundary, and gate coverage is incomplete when writes are enabled. The manager-side cloud_bash (approval-gated) runs in a shell runner labelled 'IsolationNone' in its own code.
C5 Untrusted input blast radius
Minimal 0.13 / 1.00
The agent reads plenty of content its operator didn't write: application and Kubernetes logs, traces, synced git repositories, knowledge documents, MCP results, IM messages and web search results. All of it enters the conversation as ordinary tool output, with no provenance tracking. The only defence is a prompt-level hint on one tool. In the default read-only configuration, a hijacked agent can still read sensitive host data as root. A hijacked session has unattended egress paths. Gaps in the shell filter widen this further.
C6 Memory, context & configuration integrity
Minimal 0.35 / 1.00
The agent has no free-form memory tool. The ways it can persist things are installing a skill, whose instructions are later injected into the system prompt for everyone (gated by human approval), creating alert rules, a per-session cloud_bash workspace, and hosted pages. The knowledge base is curated by operators. Installed skills load silently into future sessions for all users, with no expiry, review or rollback.
C7 Third-party extensions
Minimal 0.30 / 1.00
Nothing third-party is enabled by default. Skills are installed from a git or tarball URL that the user supplies, after an approval card, and MCP servers are registered by admins over HTTP only. Signature checks apply only to the official registry, git refs are not pinned, MCP tool lists are fetched again on every boot without re-approval, and tools from 'trusted' servers run without approval. Skill binaries run through cloud_bash on the manager as the same OS user, with a scrubbed environment.
C8 Secrets & sensitive-data protection
Minimal 0.45 / 1.00
Stored credentials are kept in a vault encrypted with AES-256-GCM. The model sees only credential names, and the values are injected as environment variables at execution time. The command runners on both the manager and the edge build a minimal environment instead of inheriting the parent's. There is no telemetry SDK, and Kubernetes event text is redacted. However, tool output, including cloud_bash output, goes back to the model without secret scanning. Vault key management is not locked down. The vault holds long-lived cloud keys.
C9 Audit & traceability
Minimal 0.45 / 1.00
Every tool call in the chat graph is written to a database table with its arguments and status. Approvals record which admin decided, LLM-review decisions are logged separately, and worker sessions link to their parent session. The records live in the manager database, are written on a best-effort basis (failures are logged and the action goes ahead), and are not tamper-evident. IM-originated actions are attributed to a shared service account. On the edge, ordinary command execution is logged only at debug level.
C10 Limits & kill switch
Minimal 0.45 / 1.00
The ReAct loop is capped at 12 iterations by default (specialist workers 15 to 40). Tools time out after 15 seconds by default. Each tool can be called at most 10 times per turn, side-effecting tools are rate-limited to 10 per minute per user, and LLM requests time out after 120 seconds. There is no token or cost cap. A stop endpoint cancels the session's context, but background workers run from a detached context, and spawned workers get their own turn budgets.