Skip to content
AgenticCyber

lab

Lab guide: agentgateway on a kind cluster

This is how I deployed it on my laptop, step by step. Each step says why it is there, gives the commands and ends with a check. Run every command from the root of the lab repository.

QUICKSTART

Quickstart: all the commands

If you only want to run it, this is the whole lab. It takes me about 15 minutes. The steps below explain each part.

shell
git clone https://github.com/fsclyde/agentic-gateway-lab.git && cd agentic-gateway-lab
python3 -m venv venv && source venv/bin/activate
pip install PyJWT cryptography

kind create cluster --name agentic

kubectl apply --server-side \
  -f https://github.com/kubernetes-sigs/gateway-api/releases/download/v1.6.2/standard-install.yaml

helm upgrade -i agentgateway-crds oci://cr.agentgateway.dev/charts/agentgateway-crds \
  --create-namespace --namespace agentgateway-system --version v1.6.0

helm upgrade -i agentgateway oci://cr.agentgateway.dev/charts/agentgateway \
  --namespace agentgateway-system --version v1.6.0 --wait --timeout 5m

kubectl apply -f manifests/00-namespace.yaml -f manifests/01-gateway.yaml
kubectl wait --for=condition=Programmed gateway/agentgateway-proxy \
  -n agentgateway-system --timeout 5m

kubectl apply -f manifests/10-mock-llm.yaml -f manifests/11-mcp-server.yaml \
  -f manifests/12-backends-routes.yaml
kubectl rollout status deploy/mock-llm deploy/mcp-everything -n agentic-lab --timeout 5m

In a second terminal, leave this running:

shell
kubectl port-forward -n agentgateway-system service/agentgateway-proxy 8080:80

Back in the first terminal:

shell
python3 scripts/tokens.py init
python3 scripts/check.py --stage before      # 11/11: everything is open

python3 scripts/tokens.py policy > manifests/20-jwt-policy.yaml
kubectl apply -f manifests/20-jwt-policy.yaml      # agent identity (JWT)
kubectl apply -f manifests/21-mcp-tool-rules.yaml  # which agent may use which tool
kubectl apply -f manifests/22-llm-guardrails.yaml  # inspect prompts and answers
kubectl apply -f manifests/23-llm-rate-limit.yaml  # budget per agent

python3 scripts/check.py                     # 11/11: everything is enforced

Clean up:

shell
kind delete cluster --name agentic

step 0 / SETUP

Tools

Why: I use kind because it runs a Kubernetes cluster inside Docker. The whole lab stays on my laptop and I can delete it with one command.

shell
brew install kind kubectl helm

Check: kind version, kubectl version --client and helm version --short each print a version, and Docker is running.

step 1 / GOVERN

Repo and rules before anything runs

Why: I decide the rules before I deploy anything: every rule is a file in Git, versions are pinned, every resource has an owner, and the signing key stays out of the repository.

shell
git clone https://github.com/fsclyde/agentic-gateway-lab.git && cd agentic-gateway-lab
cat .gitignore
grep -rn "owner:\|csf-function:\|@2026\|v1.6" manifests README.md | head -20
shell
# Python environment for the token and proof scripts
python3 -m venv venv && source venv/bin/activate
pip install PyJWT cryptography

Run source venv/bin/activate again in every new terminal you use for the Python scripts.

Check: manifests/01-gateway.yaml only accepts routes from namespaces labelled agentgateway-access: "true".

step 2 / SETUP

Cluster, controller, Gateway

Why: The Gateway API adds the standard Gateway and HTTPRoute resource types. The first Helm chart adds agentgateway's own types (backends and policies) and the second installs the controller. When I apply a Gateway resource, the controller creates the proxy for me.

shell
kind create cluster --name agentic

kubectl apply --server-side \
  -f https://github.com/kubernetes-sigs/gateway-api/releases/download/v1.6.2/standard-install.yaml

helm upgrade -i agentgateway-crds oci://cr.agentgateway.dev/charts/agentgateway-crds \
  --create-namespace --namespace agentgateway-system --version v1.6.0

helm upgrade -i agentgateway oci://cr.agentgateway.dev/charts/agentgateway \
  --namespace agentgateway-system --version v1.6.0 --wait --timeout 5m

kubectl apply -f manifests/00-namespace.yaml -f manifests/01-gateway.yaml
kubectl wait --for=condition=Programmed gateway/agentgateway-proxy -n agentgateway-system --timeout 5m
kubectl get pods,svc -n agentgateway-system

Check: pods agentgateway (controller) and agentgateway-proxy are Running.

step 3 / IDENTIFY

Deploy, take inventory, attack the open setup

Why: I want to know what I have and what can go wrong before I choose controls. The "before" result is my baseline. Without it the "after" result means nothing.

shell
kubectl apply -f manifests/10-mock-llm.yaml -f manifests/11-mcp-server.yaml -f manifests/12-backends-routes.yaml
kubectl rollout status deploy/mock-llm deploy/mcp-everything -n agentic-lab --timeout 5m

# second terminal, leave running
kubectl port-forward -n agentgateway-system service/agentgateway-proxy 8080:80
shell
kubectl get gateway,httproute,agentgatewaybackend,agentgatewaypolicy -A
cat docs/threat-model.md

source venv/bin/activate   # if not already active
python3 scripts/tokens.py init
python3 scripts/check.py --stage before

Check: 11/11 pass in the "before" sense: no identity needed, the AWS key reaches the model, every agent sees all tools, and get-env returns the planted DEMO_DB_PASSWORD.

What I got

output · my run, 4 Oct 2026 · kind on macOS
Gateway: http://localhost:8080   Expecting: open state

[PASS] PROTECT  1. Caller with no identity reaches the LLM -> allowed (no policy yet)
         HTTP 200
[PASS] PROTECT  2. Expired token -> allowed (no policy yet)
         HTTP 200
[PASS] PROTECT  3. Orchestrator with a valid token reaches the LLM -> allowed
         HTTP 200
[PASS] PROTECT  4. Prompt containing an AWS key -> allowed (no policy yet)
         HTTP 200 {"model":"mock-model","usage":{"prompt_tokens":12,"completion_tokens":9,"total_t
[PASS] PROTECT  5. Model answer containing an email and a key -> returned as is
         HTTP 200 content":"Sure, the admin contact is admin@example.com and the key is AKIAIOSFODNN7EXAMPLE.","role":"assistant
[PASS] IDENTIFY 6. Tool inventory seen by each agent
         orchestrator sees 13: ['echo', 'get-annotated-message', 'get-env', 'get-resource-links']... | worker sees 13: ['echo', 'get-annotated-message', 'get-env', 'get-resource-l
[PASS] PROTECT  7. Worker calls a tool reserved for the orchestrator (get-sum) -> allowed (no policy yet)
         {"jsonrpc": "2.0", "id": 3, "result": {"content": [{"type": "text", "text": "The sum of 2 and 3 is 5."}]}}
[PASS] PROTECT  8. Any agent dumps the tool server's environment (get-env) -> allowed (no policy yet)
         environment returned, including DEMO_DB_PASSWORD
[PASS] PROTECT  9. Orchestrator calls get-sum -> allowed
         {"jsonrpc": "2.0", "id": 4, "result": {"content": [{"type": "text", "text": "The sum of 2 and 3 is 5."}]}}
[PASS] PROTECT  10. Caller with no identity opens an MCP session -> allowed (no policy yet)
         HTTP 200
[PASS] PROTECT  11. Worker loops 8 LLM calls in a row -> all served
         200 200 200 200 200 200 200 200

step 4 / PROTECT

Apply one control at a time

Why: I apply one control at a time and re-run the proof script after each one, so I see exactly which control stops which attack.

shell
# 4a. Identity: every request needs a signed token      (flips checks 1, 2, 10)
python3 scripts/tokens.py policy > manifests/20-jwt-policy.yaml
kubectl apply -f manifests/20-jwt-policy.yaml

# 4b. Authorization: which agent may use which tool      (flips checks 6, 7, 8)
kubectl apply -f manifests/21-mcp-tool-rules.yaml

# 4c. Data protection: inspect prompts and answers        (flips checks 4, 5)
kubectl apply -f manifests/22-llm-guardrails.yaml

# 4d. Limits: a budget per agent identity                 (flips check 11)
kubectl apply -f manifests/23-llm-rate-limit.yaml

python3 scripts/check.py
shell
# 4e. See the provider key swap
ORCH=$(python3 scripts/tokens.py mint orchestrator)
curl -s localhost:8080/v1/chat/completions -H "authorization: Bearer $ORCH" \
  -H 'content-type: application/json' \
  -d '{"model":"mock-model","messages":[{"role":"user","content":"hello"}]}'
kubectl logs -n agentic-lab deploy/mock-llm --tail=3

Check: 11/11 pass in the secured sense. The mock LLM's log shows the gateway's provider key, not your agent token. If you run the script twice within a minute a check can fail with 429; wait 60 seconds.

What I got

To watch the controls take effect, I re-ran the baseline script (--stage before) after applying them. A FAIL here is the result I want: the open behaviour it expects is gone.

output · my run, 4 Oct 2026 · kind on macOS
[FAIL] PROTECT  1. Caller with no identity reaches the LLM -> allowed (no policy yet)
         HTTP 401
[FAIL] PROTECT  4. Prompt containing an AWS key -> allowed (no policy yet)
         HTTP 403 Blocked by gateway: the prompt contains a secret or card number.
[FAIL] PROTECT  5. Model answer containing an email and a key -> returned as is
         HTTP 200 content":"Sure, the admin contact is <EMAIL_ADDRESS> and the key is <masked>.","role":"assistant"},"index":0,"
[FAIL] IDENTIFY 6. Tool inventory seen by each agent
         orchestrator sees 2: ['echo', 'get-sum'] | worker sees 1: ['echo']
[FAIL] PROTECT  10. Caller with no identity opens an MCP session -> allowed (no policy yet)
         HTTP 401
[FAIL] PROTECT  11. Worker loops 8 LLM calls in a row -> all served
         200 200 200 200 200 429 429 429

step 5 / DETECT

Read the logs and metrics

Why: A control that blocks silently is only half the job. Because every call goes through the proxy, I get one log stream for model calls and tool calls.

shell
kubectl logs -n agentgateway-system deploy/agentgateway-proxy --tail=200 | grep -E '401|403|429'
kubectl logs -n agentgateway-system deploy/agentgateway-proxy --tail=200 | grep 'tools/call'

# third terminal
kubectl port-forward -n agentgateway-system deploy/agentgateway-proxy 15020:15020
curl -s localhost:15020/metrics | grep -E 'agentgateway_(mcp_requests|guardrail_checks|gen_ai_client_token_usage_count)'

Check: each refused call shows the agent (jwt.sub), the route, the status and a reason.

What I got

gateway access log · my run, 4 Oct 2026
2026-10-04T15:02:58.323097Z	error	request gateway=agentgateway-system/agentgateway-proxy listener=http route=agentic-lab/llm src.addr=127.0.0.1:53068 http.method=POST http.host=localhost http.path=/v1/chat/completions http.version=HTTP/1.1 http.status=429 jwt.sub=worker protocol=http error="rate limit exceeded" reason=RateLimit duration=0ms
2026-10-04T15:02:58.326208Z	error	request gateway=agentgateway-system/agentgateway-proxy listener=http route=agentic-lab/llm src.addr=127.0.0.1:53074 http.method=POST http.host=localhost http.path=/v1/chat/completions http.version=HTTP/1.1 http.status=429 jwt.sub=worker protocol=http error="rate limit exceeded" reason=RateLimit duration=0ms
2026-10-04T15:02:58.331921Z	error	request gateway=agentgateway-system/agentgateway-proxy listener=http route=agentic-lab/llm src.addr=127.0.0.1:53076 http.method=POST http.host=localhost http.path=/v1/chat/completions http.version=HTTP/1.1 http.status=429 jwt.sub=worker protocol=http error="rate limit exceeded" reason=RateLimit duration=0ms

step 6 / RESPOND

The worker agent is compromised

Why: I want to stop one agent fast without stopping the others, then remove the attacker's access for good.

shell
# contain: block that one identity on every route
kubectl apply -f manifests/30-respond-block-worker.yaml

# eradicate: rotate the signing key so every old token stops working
python3 scripts/tokens.py init
python3 scripts/tokens.py policy > manifests/20-jwt-policy.yaml
kubectl apply -f manifests/20-jwt-policy.yaml

# close the incident
kubectl delete -f manifests/30-respond-block-worker.yaml
git add -A && git commit -m "LAB-001: rotate signing key after worker compromise"

Check: while the block is applied the worker gets 403 and the orchestrator still gets 200; after rotation old tokens get 401.

What I got

output · my run, 4 Oct 2026 · kind on macOS
$ curl -s -o /dev/null -w "%{http_code}\n" localhost:8080/v1/chat/completions -H "authorization: Bearer $WORK" ...
200
$ kubectl apply -f manifests/30-respond-block-worker.yaml
agentgatewaypolicy.agentgateway.dev/incident-block-worker created
$ curl -s -o /dev/null -w "%{http_code}\n" localhost:8080/v1/chat/completions -H "authorization: Bearer $WORK" ...
403

step 7 / RECOVER

Rebuild from Git

Why: Recover means I can get back to a known-good state. I destroy the cluster, rebuild it from the repository only and prove the controls are back.

shell
kind delete cluster --name agentic
# repeat Step 2, then:
kubectl apply -f manifests/10-mock-llm.yaml -f manifests/11-mcp-server.yaml -f manifests/12-backends-routes.yaml
kubectl apply -f manifests/20-jwt-policy.yaml -f manifests/21-mcp-tool-rules.yaml \
  -f manifests/22-llm-guardrails.yaml -f manifests/23-llm-rate-limit.yaml
# restart the port-forward, then:
python3 scripts/check.py

Check: 11/11 again. The rebuild time is your recovery time.

reading

Further reading

← Back to the project