Contain prompt injection: limit the blast radius
Constle does not detect prompt injection. It assumes the agent can be steered by whatever text it reads, and limits what a steered agent can do, from outside it.
On this page
- What you'll get
- The approach
- 1. Ask what a hijacked run could do
- 2. Write the narrowest Agentfile that still does the job
- 3. Check every control is armed
- 4. Run it, and test the refusals
- 5. Verify the log after an incident
- What a hijacked agent can still do under this file
- FAQ
- Does Constle detect or block prompt injections?
- Is a human gate enough on its own?
- Why should the agent's rules not live in its system prompt?
Source release coming soon
The source and installers are not public yet, so there is nothing to install today. This guide documents Constle as it behaves now. More in Project status.
What you'll getLink to this section
- A worked Agentfile for an agent that reads untrusted text (support tickets, here) with each way out narrowed: network, tools, consequential calls, credentials, time.
- A list of what an injected agent can still do under that file, so you can decide what else it needs.
- A signed record of every attempt, allowed or refused, to find out afterwards that something happened.
The approachLink to this section
A rule the model reads in its prompt, or a check that runs in the agent's own process, is exposed to the same injected text that steers the agent. Constle's rules live in a separate host process, at the points the agent's traffic must cross: the egress proxy, the MCP gate, the agent-to-agent gate and the supervisor. The injection can change what the agent tries; it cannot change what those points allow. The blast radius is then whatever the Agentfile allows, so write it for the worst run, not the typical one. Why this has to be outside the agent: Architecture.
1. Ask what a hijacked run could doLink to this section
For each question, the Agentfile field that answers it:
| If the agent is steered, what can it… | Bounded by |
|---|---|
| …connect to? | sandbox.network.allowed_hosts |
| …call, on which tool servers? | mcp.servers[].tools |
| …do without a person agreeing? | human_gates.require_approval_for |
| …spend on priced tools? | spending.max_per_run_usd, max_per_day_usd |
| …read from your environment? | credentials |
| …keep doing, for how long? | limits.max_duration_seconds, sandbox.memory_mb |
2. Write the narrowest Agentfile that still does the jobLink to this section
apiVersion: constle.dev/v1alpha1
kind: AgentManifest
identity:
name: support-triage
version: "1.0.0"
did: did:key:z6Mk...4doK # signs every audit entry
sandbox:
image: support-triage:1.4.2
memory_mb: 1024
network:
allowed_hosts:
- api.groq.com # the model API; nothing else on the internet
capabilities:
- external_api
credentials:
- name: GROQ_API_KEY # the only host variable the agent receives
mcp:
servers:
- id: tickets
url: https://mcp.tickets.internal/mcp
tools: [get_ticket, add_comment, close_ticket] # every other tool is refused
limits:
max_duration_seconds: 600
human_gates:
enabled: true
require_approval_for:
- close_ticket # closing waits for a person
approver_pubkey: did:key:z6Mk...Ap7w
on_timeout: abortThe DIDs are shortened placeholders: constle identity create and constle webhook-keygen print the full ones.
3. Check every control is armedLink to this section
constle validate agent.yamlRead every warning. A gate with enabled left false, a gate entry that names no declared tool, or a spending cap with nothing to meter all look like controls and are not; validate names each one. More in Agentfile in 5 minutes.
4. Run it, and test the refusalsLink to this section
Before relying on the file, make the agent try what an injection would: reach an undeclared host, call a tool that is not listed, call close_ticket. Each should fail outside the agent and leave a line in the audit log:
| Attempt | Refused by | Audit event |
|---|---|---|
Send data to attacker.example |
the egress proxy (403) | network_blocked |
Call delete_ticket |
the MCP gate | mcp_tool_blocked |
Call close_ticket with nobody answering |
the human gate, then on_timeout: abort |
gate_triggered, gate_timeout |
| Keep going past ten minutes | the supervisor | terminated_by_limit |
5. Verify the log after an incidentLink to this section
constle audit verify --agentfile=agent.yaml ~/.constle/logs/support-triage-$(date -u +%F).jsonlWith identity.did set, every entry is signed and hash-chained, so an edited, deleted or reordered line is reported with its position. See Verify what an AI agent did.
What a hijacked agent can still do under this fileLink to this section
Containment is not prevention. Under the Agentfile above, an injected instruction can still:
- Send what it can read to
api.groq.com. Every allowed host is a destination, and the model API receives the prompts the agent writes. - Call
get_ticketandadd_commentfreely. An ungated tool is a channel: a comment can carry another customer's data. - Read
GROQ_API_KEY. Declared credentials are inside the sandbox, and can be sent to an allowed host. - Ask for
close_ticketapproval with a misleading context. The prompt shows the exact call and its arguments; the decision is the person's. - Write anything inside its own sandbox. There is no filesystem policy and no gate on file writes.
- Spend on the model API. Direct LLM traffic is not metered (limitation 3); the time limit bounds it.
Each of these is narrower than an agent with your shell's environment, an open network and every tool its servers expose, which is the point. What is left is yours to size: gate more tools, split the agent in two, or keep sensitive data out of its reach. The complete list of gaps: Known limitations.
FAQLink to this section
Does Constle detect or block prompt injections?Link to this section
No. It does not read prompts or judge model output. It limits what the agent can reach, call and spend, whatever the model has been told. In-agent guardrails that judge content can sit on top; they cover a different layer.
Is a human gate enough on its own?Link to this section
It holds the calls you name, and only MCP tool calls. A steered agent can still use every ungated tool and every allowed host, so gates work alongside a short allowed_hosts list and a tool allowlist, not instead of them.
Why should the agent's rules not live in its system prompt?Link to this section
Because the system prompt is text in the same context as the injected text. A rule there can be argued with; a refusal at the proxy cannot.