constle docs constle docs pre-1.0

Stop an AI agent exfiltrating data: egress allowlists

An agent that reads untrusted text can be told to send what it knows somewhere else. The reliable control is not the model's judgement but the network: one way out, checked against a list the agent cannot change.

On this page

Source release coming soon

The source and installers are not public yet, so there is nothing to install today. This guide documents Constle as it behaves now. More in Project status.

What you'll getLink to this section

  • A sandbox with no default route, IPv4 or IPv6, where the only way out to the internet is a proxy checking each connection against allowed_hosts.
  • Refusals for the usual detours: undeclared names, raw IP addresses, names that resolve to internal or metadata addresses, other ports, DNS.
  • A network_allowed or network_blocked audit event for every attempt, so a refused exfiltration attempt is something you can find afterwards.

The ideaLink to this section

Exfiltration needs a destination. If the only destinations the sandbox can reach are the ones you listed, an injected instruction to send a file to attacker.example fails at the network layer whatever the model decides, and the attempt is recorded. The list lives in the Agentfile, which the runtime reads once before the sandbox starts; nothing the agent sends to the proxy changes it. How the topology is built on each backend: Network isolation.

1. List only the hosts the agent needsLink to this section

agent.yamlyaml
sandbox:
  network:
    allowed_hosts:
      - api.groq.com          # the model API this agent calls

Start from the smallest list that lets the agent do its job: usually its model API and nothing else. An empty or absent list reaches no host at all. Entries are plain lowercase hostnames; a leading . admits a domain and every subdomain under it.

Be careful with a leading . on a shared domain. An entry for a hosting or storage domain where anyone can create a subdomain admits every customer's subdomain, including one an attacker controls.

2. Leave MCP servers and peers out of the listLink to this section

yaml
mcp:
  servers:
    - id: docs-search
      url: https://mcp.search.internal/mcp     # reached through the gate, never directly
      tools: [search]

Declare MCP servers under mcp.servers and agent-to-agent peers under a2a.peers, not in allowed_hosts. They are then reachable only through Constle's gates, where tool allowlists, approvals and metering apply. An Agentfile that lists a declared server's or peer's host in allowed_hosts is rejected at validate time, because that entry would let the agent skip the gate. More in Lock down MCP servers.

3. Validate the listLink to this section

shell
constle validate agent.yaml

A scheme, port, path, wildcard, uppercase letter or non-ASCII name in allowed_hosts is a validation error, as is localhost, 127.0.0.1 or host.docker.internal when mcp or a2a is declared. sandbox.network.egress is not a substitute: it parses and is ignored (limitation 5).

4. Run, and read what was refusedLink to this section

shell
constle run agent.yaml
grep network ~/.constle/logs/<agent>-$(date -u +%F).jsonl

An agent with allowed_hosts: [api.groq.com] reaching for one declared host and one undeclared one:

Output
  ┌─ agent output ──────────────────────────
  │ https://api.groq.com/        CONNECT allowed   TLS tunnel opened, server replied
  │ https://evil.example.com/    CONNECT refused   Tunnel connection failed: 403 Forbidden
  └─────────────────────────────────────────

{"event":"network_allowed","details":{"bytes":5314,"host":"api.groq.com","http_status":200,"method":"CONNECT"}}
{"event":"network_blocked","details":{"bytes":3404,"host":"evil.example.com","http_status":403,"method":"CONNECT"}}

The network events are written when the run ends, from the proxy's own log, so in the file they follow run_finished. With identity.did set they are signed and chained like every other entry (Verify what an AI agent did).

The channels this closesLink to this section

Attempt Result
HTTPS to an undeclared host CONNECT refused with 403 · network_blocked
The raw IP of a declared host, or any IP literal not listed Refused: matching is by name, with reverse lookups off
An address whose PTR record names a declared host Refused, for the same reason
A declared name that resolves to loopback, link-local, a cloud metadata address, a private or CGNAT range Refused on the resolved address, whatever DNS answered
A tunnel to a declared host on port 22, 5432 or anything but 443 Refused: a CONNECT may name 443 only, a plain request 80 or 443
Dialling out directly, with the proxy variables cleared Network is unreachable: there is no route, IPv4 or IPv6
Encoding data in DNS lookups Nothing resolves inside the sandbox, and a refused name is refused before anything resolves it
An MCP server or peer that is not declared, on a host that is not allowed Not reachable: declared ones are reached only through the gates

The channels it leaves openLink to this section

An allowlist decides where data can go, not what is in it. Size these honestly:

  • Every allowed host is a channel. The agent can put anything it can read into a request to a host you allowed. A model API receives the prompts the agent sends; a code host you allowed receives what the agent pushes.
  • Declared MCP tools are channels. A tool's arguments go to its server. Tool allowlists and human approval narrow which calls happen; nothing inspects their contents.
  • DNS is trusted for public addresses. If a declared name resolves to a public address an attacker controls, the proxy lets it through. What DNS cannot do is aim a declared name inside: those destinations are refused on the resolved address.
  • Declared credentials are inside. The variables named under credentials are readable by the agent, and can be sent to any allowed host.
  • Contents are never inspected. Constle does not intercept the agent's TLS, so it cannot apply per-path rules or look for secrets in a request body. That is a deliberate trade-off (What Constle is not).
  • The policy is per sandbox. Every process inside reaches the same hosts.

FAQLink to this section

Can a prompt injection change the allowlist?Link to this section

No. The allowlist is read from the Agentfile before the sandbox starts and enforced by a proxy in the host's control. Nothing the agent can send changes it, and there is no setting or endpoint inside the sandbox to act on.

Does the agent need DNS to reach an allowed host?Link to this section

No. Proxy-aware HTTP clients send the hostname to the proxy, which resolves it. Inside the sandbox nothing resolves, which is also what closes DNS as a channel.

Can I allow a host on a port other than 443?Link to this section

Not through the allowlist. A CONNECT tunnel may name port 443 only, and a plain request port 80 or 443, so an allowed name cannot become a raw TCP path to an SSH or database port.

How do I find out the agent tried something it shouldn't?Link to this section

Every refused connection is a network_blocked event with the host it tried, in the run's audit log. With an identity declared, the log is signed and hash-chained, and constle audit verify names any line that was changed afterwards.