💪Exercise💪 — clustering split with an AI agent: who scrapes what
Same end state as the manual variant: duplicate scraping without the component block, a split with it, targets moving when a third replica joins. You still change the release by hand. Counting and interpreting move to an AI agent that reads your replicas’ metrics from the shared Prometheus through the lab’s Grafana MCP server. You predict each number before the agent reports it, check its arithmetic, and catch where it states a guess as a measurement.
Goal
- The agent counts targets per replica (
query_prometheus) after each change. You compare its numbers with your prediction and with one direct count. - You decide which number is the real target count, and why the sum across replicas is not.
- You separate what the metrics measured about target movement from what the agent inferred from how clustering works.
Agent: Claude Code (any MCP client with Streamable HTTP works), Grafana MCP only, no shell.
Prerequisites
- Your namespace and variables from First pipeline (
ME,NS, Grafana Helm repo). - A named Grafana account on the lab (
user1…userN, from the lab handout). It has the Admin role, so it can create service accounts. The handout also carries the MCP server token. curl,jq, Claude Code.- Two terminals: A for
kubectl/helm(yours), B for the agent. - Do one variant of this exercise.
| Step | Who | Why |
|---|---|---|
values file, helm upgrade, kubectl scale |
you, terminal A | MCP has no path to the cluster |
| targets per replica, peers | agent | query_prometheus on the shared Prometheus; your release is scraped through its ServiceMonitor |
| ground truth | you | count-targets.sh from the manual variant, Alloy UI Clustering page |
No shell for the agent — an explicit risk decision. Terminal B inherits your kubeconfig. An agent with Bash could run kubectl scale or edit the values file to “fix” the duplication. Then nothing is left to analyse, and a mistake lands on a shared cluster. Start it with Bash, Edit and Write disallowed.
💪Exercise💪 — steps
1. A read-only token for the agent
The agent only reads, so it gets its own service account with the Viewer role. Terminal B:
GRAFANA_URL=https://grafana.workshop2.indexoutofrange.com
LOGIN=<your Grafana login>
read -rsp "Grafana password: " P && echo
H=(-u "$LOGIN:$P" -H "Content-Type: application/json")
SA=$(curl -sS "${H[@]}" "$GRAFANA_URL/api/serviceaccounts/search?query=mcp-ro-$LOGIN" \
| jq -j --arg n "mcp-ro-$LOGIN" '.serviceAccounts[] | select(.name==$n) | .id')
[ -n "$SA" ] || SA=$(curl -sS "${H[@]}" -X POST "$GRAFANA_URL/api/serviceaccounts" \
-d "{\"name\":\"mcp-ro-$LOGIN\",\"role\":\"Viewer\"}" | jq -j .id)
SA_TOKEN=$(curl -sS "${H[@]}" -X POST "$GRAFANA_URL/api/serviceaccounts/$SA/tokens" \
-d "{\"name\":\"mcp-ro-05b-$(date +%s)\",\"secondsToLive\":28800}" | jq -j .key)
curl -sS -o /dev/null -w '%{http_code}\n' -H "Authorization: Bearer $SA_TOKEN" -H 'Content-Type: application/json' \
-X POST "$GRAFANA_URL/api/dashboards/db" -d '{"dashboard":{"title":"mcp-ro probe","panels":[]}}'
Nothing is exported. The agent you start from this shell inherits exported variables, and your password belongs to an Admin.
Check: the probe prints 403 and $SA_TOKEN starts with glsa_.
2. Connect Grafana MCP and start the agent
mkdir -p ~/collectors-ai && cd ~/collectors-ai
read -rsp "MCP server token: " MCP_SERVER_TOKEN && echo
claude mcp add --transport http grafana https://mcp.workshop2.indexoutofrange.com/mcp \
--header "Authorization: Bearer $MCP_SERVER_TOKEN" \
--header "X-Grafana-Service-Account-Token: $SA_TOKEN"
claude mcp list
claude --disallowedTools "Bash" "Edit" "Write"
The default local scope keeps both tokens in ~/.claude.json, bound to this empty directory. The agent cannot see your values files from here. Read test:
Via Grafana MCP, read only: which Grafana identity and role are you using?
In Prometheus (uid "prometheus") list the values of the label "namespace" on prometheus_scrape_targets_gauge.
Check: the identity is the mcp-ro-<login> service account with role Viewer. Your collectors-<ME> namespace is not in the list yet.
3. Deploy two replicas, scraped by the gateway (terminal A)
Write ex-cluster.values.yaml exactly as in step 1 of the manual variant (component block commented out), then add a ServiceMonitor. The shared gateway’s prometheus.operator.servicemonitors picks it up from every namespace and scrapes every 30 s.
printf 'serviceMonitor:\n enabled: true\n' >> ex-cluster.values.yaml
helm upgrade --install ex-alloy-cluster-$ME grafana/alloy --version 1.8.1 -n $NS -f ex-cluster.values.yaml
kubectl -n $NS rollout status deploy/ex-alloy-cluster-$ME
Write count-targets.sh from step 2 of the manual variant too. You run it once per step as ground truth.
Before asking the agent, fill in your prediction column (about 60 pods expose http-metrics in monitoring):
| After | Your prediction (per pod, sum) | Agent (per pod, sum) | count-targets.sh |
|---|---|---|---|
| step 3 — no component block | |||
| step 5 — component block | |||
| step 6 — third replica |
4. The agent counts (terminal B)
Wait about two minutes for the first scrapes. The prompt describes the situation and the evidence you need, not the query:
Via Grafana MCP, read only. Prometheus uid "prometheus".
My Alloy release ex-alloy-cluster-<ME> runs in namespace "collectors-<ME>": a Deployment with
several replicas, each scraped every 30 s. Each replica exposes prometheus_scrape_targets_gauge.
1. List the label names on that metric in my namespace before you query values.
2. Current value per replica (one row per pod) and the sum. Give every PromQL, the query type
(instant or range) and the evaluation time.
3. Do the replicas split the targets, or does each one scrape all of them? What in the numbers
tells you?
4. From the cluster_node_* metrics: how many peers does each replica see?
You cannot see my config. If you name a cause, name the metric that supports it.
Approve tool calls one at a time and read the arguments. Review the answer:
| What | Correct | Typical agent deviation |
|---|---|---|
| query | instant, per pod: sum by (pod) (prometheus_scrape_targets_gauge{namespace="collectors-<ME>"}) |
rate() or increase() on a gauge; sum() without by (pod): one number, duplication invisible |
| series | one per running pod | a range query or max_over_time that pulls in pods from earlier rollouts |
| real target count | one pod’s number (both equal, e.g. 59) |
the sum (e.g. 118) reported as “number of targets” |
| interpretation | every target scraped twice | “load is evenly balanced”, read from two equal numbers |
| peers | 2 per replica (gossip works) | not checked, or a metric name the agent never verified with list_prometheus_metric_names |
| cause | peers 2 + full count on each replica → the scrape component does not use clustering | “process clustering is off”, contradicting its own peer count; any statement about the config presented as fact |
Answer yourself before going on:
- Which number is the real target count here, and why is the sum wrong?
- Two equal numbers can also come from a working split that happens to be even. What check tells the two cases apart? (Hint: compare the sum with the real total.)
Check: bash count-targets.sh in terminal A matches the agent’s per-pod numbers. Fill in the row. If the agent was off, write down why: wrong query, stale series, or a sample from before the rollout.
5. Add the component block (terminal A), then ask again
Run step 3 of the manual variant (sed + helm upgrade). The reloader applies it in place in about a minute, and the next scrape is up to 30 s later. Fill in your prediction first. Then:
I changed the release. Same questions as before. Also give the timestamp of the newest sample per
pod, so I can tell values from before and after the change apart. Compare with your previous answer.
| Correct | Typical agent deviation |
|---|---|
two numbers that add up to the step 4 per-pod number, not necessarily equal (e.g. 21 + 38) |
“targets lost” because one pod has not reloaded yet, or its newest sample predates the change |
| lumpy split explained by few targets on a hash ring | uneven split reported as a bug or “rebalancing in progress” |
Open the UI of one pod (kubectl -n $NS port-forward <pod> 12345:12345, then http://localhost:12345/clustering) and confirm both peers.
6. Scale out (terminal A), then ask what moved
kubectl -n $NS scale deploy/ex-alloy-cluster-$ME --replicas=3
kubectl -n $NS rollout status deploy/ex-alloy-cluster-$ME
Prediction first, then:
A third replica joined. Per-replica counts and the sum. How many targets moved, and from which
replicas? Split your answer into what you measured and what you infer from how clustering works.
Correct:
- Three numbers, same total (e.g.
16+28+15). - Measured: both old replicas lost targets, and at least the new pod’s count moved.
- Inferred: with consistent hashing, targets move only to the new node, about 1/N of them. The gauge carries no per-target data, so it cannot show whether targets also swapped between the old pods.
Typical deviation: “exactly a third moved” stated as measured. Or “all targets were redistributed”. Or the new pod reported as “not participating” because its first scrape has not arrived.
Then run the same helm upgrade as in step 5 and ask once more for the replica count. Correct: 2 replicas, kubectl scale was drift. Deviation: the deleted pod’s last value counted as a current replica.
Success criteria
- Your table has three rows with prediction, agent and
count-targets.shvalues, and the agent’s numbers match the direct count, or you know why they did not. - You can explain why the step 4 sum is not the target count, and which check separates “duplicated” from “evenly split”.
- You can say roughly what share of targets moved, which part of that statement is measured, and which part comes from the hashing model.
- You found at least one place where the agent’s query or conclusion needed correcting, or you can show the check that confirmed it.
⭐Stretch: the block without the process flag
Run the stretch of the manual variant (alloy.clustering.enabled: false, component block kept). This changes the pod args, so the pods roll. Then:
Every replica reports the full target count again. Using metrics only: is the clustering setting
missing at the component level or at the process level? Say what each case would look like in
the metrics, and which of the two the data shows.
Correct: the agent uses the peer metrics, or their absence, and says how far they go. Without the process flag each replica is a cluster of one. Whether the component block is also missing cannot be read from the gauge, because both cases produce the same counts. Deviation: “the component block is missing”, repeated from step 4 without new evidence. Check the peer claim on the UI Clustering page of one pod.
Clean up: helm -n $NS uninstall ex-alloy-cluster-$ME. If you continue with Reload trap with an AI agent or Silent drop hunt with an AI agent, keep the agent setup. Otherwise, in terminal B: claude mcp remove grafana (in ~/collectors-ai) and curl -sS "${H[@]}" -X DELETE "$GRAFANA_URL/api/serviceaccounts/$SA".
Related lessons
- Clustering split — the same exercise without an agent
- Topologies on Kubernetes
- Grafana Alloy — clustering in our stack