01 / The problem
The problem
Agents call models, tools and other agents directly. They share API keys, have no identity of their own, follow no central policy and leave little audit trail. When something goes wrong, nobody can say which agent did what.
02 / The approach
The approach
I route all agent traffic through one gateway: agentgateway, an open-source project hosted by the Linux Foundation. The gateway authenticates the calling agent, limits which tools it can reach, inspects prompts and answers, and writes one log line per call. I use NIST CSF 2.0 as the structure, so I can ask the same six questions of every project on this blog.
03 / Architecture
Architecture
How to read it
- Controller and proxy are separate. The controller turns Kubernetes resources into configuration; the proxy is what traffic passes through.
- One policy point. Agents never talk to the model or the tool server directly.
- Agents are identities. orchestrator and worker are tokens I sign myself; the gateway trusts only the public key.
- The provider key stays with the gateway. Agents present their own identity and never hold the model provider's key.
quickstart
Run it yourself
Clone the repository and follow the quickstart. It takes me about 15 minutes on a laptop.
git clone https://github.com/fsclyde/agentic-gateway-lab.git04 / The six functions
The six functions
Each function has two parts: what I did in the lab, and how an organisation would apply the same idea as a base practice.
GOVERN
Who owns this and what are the rules?
GOVERN
Who owns this and what are the rules?
What I did
- I keep every rule as a file in Git.
- I pinned the versions of the gateway, the Gateway API and the MCP server.
- Every resource has an owner label.
- Only namespaces I labelled can publish routes on the gateway.
- The token signing key stays out of Git.
Base practices
- Assign ownership: every agent, tool server, model and policy has a named owner and a risk tier.
- Write the rules of use: which data may go to which model, which actions need a human approval.
- Change policy only through review: pull request, second reviewer, CI validation, pipeline deploy.
- Control the supply chain: an approved catalogue of models and MCP servers, pinned and reviewed before upgrade.
- Gate onboarding: a team publishes through the gateway only after registering its agent and tools.
- Keep evidence and report: Git history answers "who allowed this and when".
IDENTIFY
What do I have and what can go wrong?
IDENTIFY
What do I have and what can go wrong?
What I did
- I listed the backends, routes and policies straight from the cluster and wrote a one-page threat model.
- Then I ran my attack script against the unprotected setup to get a baseline: no identity needed, a secret reaches the model, every agent sees every tool, and one tool returns the server's environment with a password I planted.
Base practices
- Keep a live inventory generated from gateway configuration and logs, not a spreadsheet.
- Find shadow usage: teams calling model APIs or MCP servers directly.
- Classify tools (read-only, writing, destructive) and data by sensitivity; flag agents that combine untrusted input, private data and a way to send data out.
- Threat-model every agent path before go-live with one template.
- Measure the baseline and repeat after each significant change.
PROTECT
What stops a bad action?
PROTECT
What stops a bad action?
What I did
- I added JWT authentication on the gateway, so every call carries an agent identity.
- I wrote one tool rule in CEL: anyone can use echo, only the orchestrator can use get-sum, nobody can use get-env.
- I added prompt guards that reject a prompt containing an AWS key or a card number and mask emails and keys in answers.
- The gateway injects the provider key from a Secret, so the agents never hold it.
- Each agent gets 5 LLM requests per minute.
matchExpressions:
- 'mcp.tool.name == "echo" || (jwt.sub == "orchestrator" && mcp.tool.name == "get-sum")'Base practices
- Give every agent its own short-lived identity from the corporate identity provider; no shared keys.
- Least privilege on tools: deny by default, allow per agent, human approval for destructive actions.
- Hold provider keys centrally and rotate them.
- Limit rate and spend per identity.
- Inspect content in both directions; start rules in audit mode, then enforce.
- Close the bypass: the gateway must be the only route to models and tools.
- Train the people who build and review agents.
DETECT
How do I know something happened?
DETECT
How do I know something happened?
What I did
- I read one access log for both model calls and tool calls.
- Each line has the agent identity, the route, the status and a reason; tool calls also have the tool name.
- I also checked the Prometheus metrics for MCP requests, guardrail checks and token usage.
2026-10-04T15:02:58.323097Z error request gateway=agentgateway-system/agentgateway-proxy listener=http route=agentic-lab/llm src.addr=127.0.0.1:53068 http.method=POST http.host=localhost http.path=/v1/chat/completions http.version=HTTP/1.1 http.status=429 jwt.sub=worker protocol=http error="rate limit exceeded" reason=RateLimit duration=0msBase practices
- Centralise gateway logs in the security monitoring platform.
- Define a short list of alerts from a baseline: bursts of denied calls, any guardrail rejection, first use of a tool, repeated rate-limit hits.
- Watch cost and volume per agent.
- Trace end to end with OpenTelemetry.
- Keep an audit trail for decisions an agent contributes to.
- Test the detections on a schedule.
RESPOND
What do I do about it?
RESPOND
What do I do about it?
What I did
- My scenario: the worker agent's token is stolen.
- I apply a deny policy I prepared in advance.
- It blocks that one identity on every route while the orchestrator keeps working.
- Then I rotate the signing key so every old token stops working, and I commit the change under an incident reference.
Base practices
- Write runbooks for agent incidents: stolen credential, malicious tool, prompt injection, runaway agent.
- Prepare and test containment levers in advance: block one agent, hide one tool, close one route.
- Revoke and rotate centrally.
- Scope the incident from the gateway logs.
- Record and communicate under an incident reference.
- Exercise it and measure time to contain.
RECOVER
How do I get back to normal?
RECOVER
How do I get back to normal?
What I did
- I delete the cluster and rebuild it from the repository only, then I run the proof script again.
- The rebuild time is my recovery time for this setup.
- The one thing that is not in Git is the signing key, so I re-issue it.
Base practices
- Make everything rebuildable from Git through a pipeline.
- Roll back by reverting a commit, then run automated control tests before reopening traffic.
- Set recovery objectives and prove them with a timed rebuild.
- Protect what is not in Git: keys and secrets.
- Be able to undo what agents changed.
- Feed lessons back into the rules, the threat model and the policies.
05 / Proof
Proof
I wrote a script that runs eleven checks against the gateway. I run it once before any policy exists and again after.
| # | Check | Before | After |
|---|---|---|---|
| 1 | Caller with no identity reaches the LLM | Allowed (200) | Blocked (401) |
| 2 | Expired token | Allowed (200) | Blocked (401) |
| 3 | Orchestrator with a valid token reaches the LLM | Allowed | Allowed |
| 4 | Prompt containing an AWS key | Reaches the model | Blocked (403) |
| 5 | Model answer containing an email and a key | Returned as is | Masked |
| 6 | Tools each agent can see | Both see all 13 | Orchestrator 2, worker 1 |
| 7 | Worker calls a tool reserved for the orchestrator | Allowed | Blocked |
| 8 | Any agent dumps the tool server's environment | Password exposed | Blocked |
| 9 | Orchestrator calls its own tool | Allowed | Allowed |
| 10 | Caller with no identity opens an MCP session | Allowed (200) | Blocked (401) |
| 11 | Worker loops eight LLM calls | All served | Throttled after 5 (429) |
From my run
These are exemple from running the lab on my laptop. The full output for each step is in the lab guide.
A FAIL means the open behaviour the baseline script expects is gone.
[FAIL] PROTECT 1. Caller with no identity reaches the LLM -> allowed (no policy yet)
HTTP 401
[FAIL] PROTECT 4. Prompt containing an AWS key -> allowed (no policy yet)
HTTP 403 Blocked by gateway: the prompt contains a secret or card number.
[FAIL] PROTECT 5. Model answer containing an email and a key -> returned as is
HTTP 200 content":"Sure, the admin contact is <EMAIL_ADDRESS> and the key is <masked>.","role":"assistant"},"index":0,"
[FAIL] IDENTIFY 6. Tool inventory seen by each agent
orchestrator sees 2: ['echo', 'get-sum'] | worker sees 1: ['echo']
[FAIL] PROTECT 10. Caller with no identity opens an MCP session -> allowed (no policy yet)
HTTP 401
[FAIL] PROTECT 11. Worker loops 8 LLM calls in a row -> all served
200 200 200 200 200 429 429 4292026-10-04T15:02:58.323097Z error request gateway=agentgateway-system/agentgateway-proxy listener=http route=agentic-lab/llm src.addr=127.0.0.1:53068 http.method=POST http.host=localhost http.path=/v1/chat/completions http.version=HTTP/1.1 http.status=429 jwt.sub=worker protocol=http error="rate limit exceeded" reason=RateLimit duration=0ms
2026-10-04T15:02:58.326208Z error request gateway=agentgateway-system/agentgateway-proxy listener=http route=agentic-lab/llm src.addr=127.0.0.1:53074 http.method=POST http.host=localhost http.path=/v1/chat/completions http.version=HTTP/1.1 http.status=429 jwt.sub=worker protocol=http error="rate limit exceeded" reason=RateLimit duration=0ms
2026-10-04T15:02:58.331921Z error request gateway=agentgateway-system/agentgateway-proxy listener=http route=agentic-lab/llm src.addr=127.0.0.1:53076 http.method=POST http.host=localhost http.path=/v1/chat/completions http.version=HTTP/1.1 http.status=429 jwt.sub=worker protocol=http error="rate limit exceeded" reason=RateLimit duration=0ms06 / Known gaps
Known gaps
What this lab does not cover yet, and what I want to fix next.
- Pods can still reach the tool server and the model directly, bypassing the gateway, until a NetworkPolicy forbids it.
- I mint tokens with a script, not a real identity provider such as Keycloak.
- I have not covered agent-to-agent (A2A) traffic yet.
- The prompt guards are regular expressions; they catch known formats only.
- The listener is plain HTTP.
- Logs stay in the pod; I don't ship them to a collector yet.
07 / Other gateways I want to compare
Other gateways I want to compare
I plan to compare agentgateway with other gateways such as Bifrost and Kong AI Gateway and write up the trade-offs in a later note.
feedback
Questions or corrections? Open an issue on the repo.