💪Exercise💪 — reload trap with an AI agent: did my deploy take effect?
Same two broken configs as in the manual variant. One is silently ignored, the other crash-loops the pod, and both times
helm upgradesaysdeployed. You still deploy by hand. After each upgrade you ask an AI agent, which sees your release only through the lab’s Grafana MCP server, one question: “is the config I just deployed running?” The config metrics answer “yes” both times. You predict where the agent will fall for them, review its evidence, and settle it withkubectl.
Goal
- For each broken config you have three answers: your prediction, the agent’s verdict with evidence, and what
kubectlshows. - You name the signals that report success for a rejected config, and the ones that do not.
- You find the link in the liveness-probe chain that is invisible from Grafana.
- You write a log-based alert expression for a rejected reload and check that it would have fired.
Agent: Claude Code (any MCP client with Streamable HTTP works), Grafana MCP only, no shell.
Prerequisites
- The
ex-alloy-$MErelease andex-alloy.values.yamlfrom First pipeline, working (fixed after its step 4).MEandNSexported as there. - A named Grafana account on the lab (
user1…userN) and the MCP server token, both from the lab handout. - Two terminals: A for
kubectl/helm, B for the agent. - Do one variant of this exercise.
| Step | Who | Why |
|---|---|---|
values files, helm upgrade |
you, terminal A | MCP has no path to the cluster |
| “is it running?” | agent | Prometheus: Alloy self-metrics via ServiceMonitor, kube-state-metrics; Loki: pod logs with namespace and container labels |
ground truth: pod status, events, /-/healthy |
you, terminal A | Kubernetes events are not shipped to Loki in this lab; the agent cannot see probe failures |
No shell for the agent. With Bash it would run kubectl get pods and the exercise collapses into reading pod status. The agent also inherits your kubeconfig, and it could “recover” the release itself. Its answer has to come from telemetry, the same view an on-call engineer has without cluster access.
💪Exercise💪 — steps
1. Agent access
Do steps 1–2 of Clustering split with an AI agent: a Viewer token for mcp-ro-<login>, the grafana MCP server in ~/collectors-ai, the agent started with claude --disallowedTools "Bash" "Edit" "Write". Already done and the token is younger than 8 h: start a new session (no --continue). Otherwise run claude mcp remove grafana and repeat both steps (token name mcp-ro-08b-…).
2. Liveness probe, ServiceMonitor, two broken copies (terminal A)
Run step 1 of the manual variant (liveness probe on /-/healthy, check that it reached the pod). Before making the two copies, add a ServiceMonitor so the release’s metrics reach the shared Prometheus:
grep -q '^serviceMonitor:' ex-alloy.values.yaml || printf 'serviceMonitor:\n enabled: true\n' >> ex-alloy.values.yaml
helm upgrade ex-alloy-$ME grafana/alloy --version 1.8.1 -n $NS -f ex-alloy.values.yaml
sed 's/otelcol.exporter.debug.default.input, /otelcol.exporter.debug.defualt.input, /' ex-alloy.values.yaml > bad-reference.values.yaml
sed 's/verbosity = "detailed"/verbosity = "loud"/' ex-alloy.values.yaml > bad-argument.values.yaml
3. Baseline: which signals does the agent trust? (terminal B)
Wait two minutes for the first scrapes.
Via Grafana MCP, read only. Prometheus uid "prometheus", Loki uid "loki".
My Alloy release ex-alloy-<ME>, one replica, namespace "collectors-<ME>". Its /metrics is scraped
every 30 s. Its pod logs are in Loki with labels namespace and container (containers "alloy" and
"config-reloader"). Kube-state-metrics are in Prometheus.
After every deploy I will ask you one question: is the config I just deployed running?
Now establish the baseline: list the signals you would use to answer it (metric names you have
verified exist, log selectors), their current values, and what each one actually measures.
| Check | Correct | Typical agent deviation |
|---|---|---|
| metric names | verified with list_prometheus_metric_names before use: alloy_config_last_load_successful, alloy_config_load_failures_total, alloy_config_hash, alloy_component_controller_running_components, kube_pod_container_status_restarts_total |
an invented name (alloy_config_reload_success); the empty result reads like “no failures” |
| log selector | {namespace="collectors-<ME>", container="alloy"} and container="config-reloader" |
service_name or app guessed; empty result reported as “no errors” |
| meaning | says what each signal measures | “last_load_successful = the running config is the deployed one” |
Loki has no pod label in this lab. If other Alloy releases run in your namespace, their lines are mixed in. Note whether the agent says so.
Before step 4, write down your prediction: after a broken-reference upgrade, which of these signals change and which stay the same? Use the tables in Config lifecycle.
4. Broken reference (terminal A, then B)
helm upgrade ex-alloy-$ME grafana/alloy --version 1.8.1 -n $NS -f bad-reference.values.yaml
date -u +%H:%M
Wait 90 s. Do not tell the agent what you changed:
I deployed a new config at <HH:MM> UTC. Is it running? Answer yes, no or unknown.
Back it with every query and its result. Say what would prove you wrong.
| Signal | What it really shows | Typical agent reading |
|---|---|---|
alloy_config_last_load_successful = 1, alloy_config_load_failures_total = 0 |
parsing, not loading: success for a rejected config | “yes, loaded successfully, zero failures” |
alloy_config_hash changed |
hash of the rejected file | “the hash changed, so the new config is live” |
running_components{health_type="unhealthy"} = 0, no restarts |
the old graph is healthy | “healthy, so the deploy worked” |
failed to reload config (container alloy), 400 Bad Request every 5 s (container config-reloader) |
the only evidence that the reload was rejected | not queried; rate() on the failures counter instead |
Correct verdict: no. The reload was rejected and the old config keeps running, backed by the log lines. Deviation: “yes” from the metrics, or “unknown” without ever querying Loki.
Verify (terminal A): step 2 of the manual variant (pod status, the log grep, /metrics and /-/healthy through a port-forward). Answer yourself: if you had to alert on this, which signal would you use, and why not alloy_config_last_load_successful?
5. Broken argument (terminal A, then B)
Write your prediction first: what will the agent see in the 2 minutes after this upgrade, and what will it not be able to see?
helm upgrade ex-alloy-$ME grafana/alloy --version 1.8.1 -n $NS -f bad-argument.values.yaml
date -u +%H:%M
for i in $(seq 1 8); do sleep 20; kubectl -n $NS get pods -l app.kubernetes.io/instance=ex-alloy-$ME --no-headers; done
Once the pod is in CrashLoopBackOff, ask the agent the same question as in step 4 with the new time.
| Correct | Typical agent deviation |
|---|---|
no: unhealthy components at 1 right after the reload, if a 30 s scrape caught the ~30 s before the kill |
reports only the latest samples and misses that window |
kube_pod_container_status_restarts_total{container="alloy"} climbing |
not checked |
| self-metrics stop: no new samples since the crash. Missing data is a finding, not “nothing wrong” | the last sample before the gap (last_load_successful 1) presented as current |
| cause backed by a log line it quotes (the reload error or the startup error) | a confident cause with no line behind it: “OOMKilled”, “probe misconfigured”, or the right argument guessed |
probe failure not visible: says the restart reason needs kubectl events |
“the liveness probe failed” stated as observed |
Verify (terminal A): the events command from step 3 of the manual variant. Answer yourself: which link of the chain /-/healthy 500 → kill → restart → exit 1 → CrashLoopBackOff did the agent observe, which did it infer, and which can only be seen with kubectl?
6. Alert on the reload outcome (you)
Write the LogQL yourself. It should fire for step 4’s rejected reload, without depending on the parse gauges. Run it in Grafana → Explore → Loki over the last 30 minutes. Compare with this reference only after yours returns data:
sum by (namespace) (count_over_time({namespace="collectors-ab", container="alloy"} |= "failed to reload config" [5m]))
Then ask the agent what your expression misses and what to pair it with. Expected: after the restart in step 5 the broken config fails at startup and logs a different line, so pair it with running_components{health_type!="healthy"} and restarts. Accept only the parts it backs with a query.
7. Recover (terminal A)
Step 4 of the manual variant. Ask the agent the step 4 question once more. Correct: “yes”, with the restart counter left over as evidence of the earlier crash loop, and self-metrics arriving again after the gap.
Success criteria
- For both upgrades you have your prediction, the agent’s verdict with its evidence, and the
kubectlresult. Where the agent was wrong, you can say which signal misled it. - You can explain why
alloy_config_last_load_successfulstayed1andalloy_config_hashchanged for a config that never ran. - You can name the link in the probe chain the agent could not observe, and why.
- Your LogQL alert returns data for the step 4 window.
⭐Stretch: move the probe to /-/ready
Run the stretch of the manual variant (probe on /-/ready, repeat the broken argument). Ask the step 4 question again. Correct: “partially”. The pod is running with no restarts and one component unhealthy, so the new graph is applied in part. Deviation: “yes, no restarts”. This is the state a probe on /-/ready leaves visible in metrics instead of fatal, and the answer you want the agent, or a dashboard, to give.
Restore the good values file when done. Keep the agent setup if you continue with Silent drop hunt with an AI agent. Otherwise clean up as at the end of Clustering split with an AI agent.
Related lessons
- Reload trap — the same exercise without an agent
- Config lifecycle
- Self-monitoring