💪Exercise💪 — reload trap with an AI agent: did my deploy take effect?

Same two broken configs as in the manual variant. One is silently ignored, the other crash-loops the pod, and both times helm upgrade says deployed. You still deploy by hand. After each upgrade you ask an AI agent, which sees your release only through the lab’s Grafana MCP server, one question: “is the config I just deployed running?” The config metrics answer “yes” both times. You predict where the agent will fall for them, review its evidence, and settle it with kubectl.

Goal

  • For each broken config you have three answers: your prediction, the agent’s verdict with evidence, and what kubectl shows.
  • You name the signals that report success for a rejected config, and the ones that do not.
  • You find the link in the liveness-probe chain that is invisible from Grafana.
  • You write a log-based alert expression for a rejected reload and check that it would have fired.

Agent: Claude Code (any MCP client with Streamable HTTP works), Grafana MCP only, no shell.

Prerequisites

  • The ex-alloy-$ME release and ex-alloy.values.yaml from First pipeline, working (fixed after its step 4). ME and NS exported as there.
  • A named Grafana account on the lab (user1 … userN) and the MCP server token, both from the lab handout.
  • Two terminals: A for kubectl/helm, B for the agent.
  • Do one variant of this exercise.
Step Who Why
values files, helm upgrade you, terminal A MCP has no path to the cluster
“is it running?” agent Prometheus: Alloy self-metrics via ServiceMonitor, kube-state-metrics; Loki: pod logs with namespace and container labels
ground truth: pod status, events, /-/healthy you, terminal A Kubernetes events are not shipped to Loki in this lab; the agent cannot see probe failures

No shell for the agent. With Bash it would run kubectl get pods and the exercise collapses into reading pod status. The agent also inherits your kubeconfig, and it could “recover” the release itself. Its answer has to come from telemetry, the same view an on-call engineer has without cluster access.

💪Exercise💪 — steps

1. Agent access

Do steps 1–2 of Clustering split with an AI agent: a Viewer token for mcp-ro-<login>, the grafana MCP server in ~/collectors-ai, the agent started with claude --disallowedTools "Bash" "Edit" "Write". Already done and the token is younger than 8 h: start a new session (no --continue). Otherwise run claude mcp remove grafana and repeat both steps (token name mcp-ro-08b-…).

2. Liveness probe, ServiceMonitor, two broken copies (terminal A)

Run step 1 of the manual variant (liveness probe on /-/healthy, check that it reached the pod). Before making the two copies, add a ServiceMonitor so the release’s metrics reach the shared Prometheus:

grep -q '^serviceMonitor:' ex-alloy.values.yaml || printf 'serviceMonitor:\n  enabled: true\n' >> ex-alloy.values.yaml
helm upgrade ex-alloy-$ME grafana/alloy --version 1.8.1 -n $NS -f ex-alloy.values.yaml
sed 's/otelcol.exporter.debug.default.input, /otelcol.exporter.debug.defualt.input, /' ex-alloy.values.yaml > bad-reference.values.yaml
sed 's/verbosity = "detailed"/verbosity = "loud"/' ex-alloy.values.yaml > bad-argument.values.yaml

3. Baseline: which signals does the agent trust? (terminal B)

Wait two minutes for the first scrapes.

Via Grafana MCP, read only. Prometheus uid "prometheus", Loki uid "loki".
My Alloy release ex-alloy-<ME>, one replica, namespace "collectors-<ME>". Its /metrics is scraped
every 30 s. Its pod logs are in Loki with labels namespace and container (containers "alloy" and
"config-reloader"). Kube-state-metrics are in Prometheus.
After every deploy I will ask you one question: is the config I just deployed running?
Now establish the baseline: list the signals you would use to answer it (metric names you have
verified exist, log selectors), their current values, and what each one actually measures.
Check Correct Typical agent deviation
metric names verified with list_prometheus_metric_names before use: alloy_config_last_load_successful, alloy_config_load_failures_total, alloy_config_hash, alloy_component_controller_running_components, kube_pod_container_status_restarts_total an invented name (alloy_config_reload_success); the empty result reads like “no failures”
log selector {namespace="collectors-<ME>", container="alloy"} and container="config-reloader" service_name or app guessed; empty result reported as “no errors”
meaning says what each signal measures “last_load_successful = the running config is the deployed one”

Loki has no pod label in this lab. If other Alloy releases run in your namespace, their lines are mixed in. Note whether the agent says so.

Before step 4, write down your prediction: after a broken-reference upgrade, which of these signals change and which stay the same? Use the tables in Config lifecycle.

4. Broken reference (terminal A, then B)

helm upgrade ex-alloy-$ME grafana/alloy --version 1.8.1 -n $NS -f bad-reference.values.yaml
date -u +%H:%M

Wait 90 s. Do not tell the agent what you changed:

I deployed a new config at <HH:MM> UTC. Is it running? Answer yes, no or unknown.
Back it with every query and its result. Say what would prove you wrong.
Signal What it really shows Typical agent reading
alloy_config_last_load_successful = 1, alloy_config_load_failures_total = 0 parsing, not loading: success for a rejected config “yes, loaded successfully, zero failures”
alloy_config_hash changed hash of the rejected file “the hash changed, so the new config is live”
running_components{health_type="unhealthy"} = 0, no restarts the old graph is healthy “healthy, so the deploy worked”
failed to reload config (container alloy), 400 Bad Request every 5 s (container config-reloader) the only evidence that the reload was rejected not queried; rate() on the failures counter instead

Correct verdict: no. The reload was rejected and the old config keeps running, backed by the log lines. Deviation: “yes” from the metrics, or “unknown” without ever querying Loki.

Verify (terminal A): step 2 of the manual variant (pod status, the log grep, /metrics and /-/healthy through a port-forward). Answer yourself: if you had to alert on this, which signal would you use, and why not alloy_config_last_load_successful?

5. Broken argument (terminal A, then B)

Write your prediction first: what will the agent see in the 2 minutes after this upgrade, and what will it not be able to see?

helm upgrade ex-alloy-$ME grafana/alloy --version 1.8.1 -n $NS -f bad-argument.values.yaml
date -u +%H:%M
for i in $(seq 1 8); do sleep 20; kubectl -n $NS get pods -l app.kubernetes.io/instance=ex-alloy-$ME --no-headers; done

Once the pod is in CrashLoopBackOff, ask the agent the same question as in step 4 with the new time.

Correct Typical agent deviation
no: unhealthy components at 1 right after the reload, if a 30 s scrape caught the ~30 s before the kill reports only the latest samples and misses that window
kube_pod_container_status_restarts_total{container="alloy"} climbing not checked
self-metrics stop: no new samples since the crash. Missing data is a finding, not “nothing wrong” the last sample before the gap (last_load_successful 1) presented as current
cause backed by a log line it quotes (the reload error or the startup error) a confident cause with no line behind it: “OOMKilled”, “probe misconfigured”, or the right argument guessed
probe failure not visible: says the restart reason needs kubectl events “the liveness probe failed” stated as observed

Verify (terminal A): the events command from step 3 of the manual variant. Answer yourself: which link of the chain /-/healthy 500 → kill → restart → exit 1 → CrashLoopBackOff did the agent observe, which did it infer, and which can only be seen with kubectl?

6. Alert on the reload outcome (you)

Write the LogQL yourself. It should fire for step 4’s rejected reload, without depending on the parse gauges. Run it in Grafana → Explore → Loki over the last 30 minutes. Compare with this reference only after yours returns data:

sum by (namespace) (count_over_time({namespace="collectors-ab", container="alloy"} |= "failed to reload config" [5m]))

Then ask the agent what your expression misses and what to pair it with. Expected: after the restart in step 5 the broken config fails at startup and logs a different line, so pair it with running_components{health_type!="healthy"} and restarts. Accept only the parts it backs with a query.

7. Recover (terminal A)

Step 4 of the manual variant. Ask the agent the step 4 question once more. Correct: “yes”, with the restart counter left over as evidence of the earlier crash loop, and self-metrics arriving again after the gap.

Success criteria

  • For both upgrades you have your prediction, the agent’s verdict with its evidence, and the kubectl result. Where the agent was wrong, you can say which signal misled it.
  • You can explain why alloy_config_last_load_successful stayed 1 and alloy_config_hash changed for a config that never ran.
  • You can name the link in the probe chain the agent could not observe, and why.
  • Your LogQL alert returns data for the step 4 window.

⭐Stretch: move the probe to /-/ready

Run the stretch of the manual variant (probe on /-/ready, repeat the broken argument). Ask the step 4 question again. Correct: “partially”. The pod is running with no restarts and one component unhealthy, so the new graph is applied in part. Deviation: “yes, no restarts”. This is the state a probe on /-/ready leaves visible in metrics instead of fatal, and the answer you want the agent, or a dashboard, to give.

Restore the good values file when done. Keep the agent setup if you continue with Silent drop hunt with an AI agent. Otherwise clean up as at the end of Clustering split with an AI agent.

results matching ""

    No results matching ""