💪Exercise💪 — silent drop hunt with an AI agent: half the logs are gone and everything is green
Same bug as in the manual variant: a colleague’s “drop DEBUG noise” rule loses more than DEBUG, while every component is healthy and the receiver answers 200. You deploy and send the records by hand. An AI agent hunts the loss through the lab’s Grafana MCP server, from collector metrics alone, and cannot read the config. You predict the counts, review every query for the traps of fresh counters, and review the fix it proposes before you apply it.
Goal
- The agent localises the loss from metrics (
query_prometheus) and counts records in Loki (query_loki_logs), without seeing the config. - You check its numbers against your prediction and against Explore. You catch the
increase()trap and the empty “dropped” metric. - You give the agent only the filter block and review its fix for boundary and scope errors before you apply it.
- After the fix, 15 of 20 records reach Loki.
Agent: Claude Code (any MCP client with Streamable HTTP works), Grafana MCP only, no shell.
Prerequisites
MEandNSfrom First pipeline.- A named Grafana account on the lab (
user1…userN) and the MCP server token, both from the lab handout. - Two terminals: A for
kubectl/helm/curl, B for the agent. - Do one variant of this exercise.
| Step | Who | Why |
|---|---|---|
| deploy, send 20 records, apply the fix | you, terminal A | MCP has no path to the cluster |
| count in Loki, localise the loss | agent | query_loki_logs, query_prometheus; the release is scraped through its ServiceMonitor |
| propose the fix | agent, from the block you paste | it never sees the rest of the config |
| verify | you | Explore, the release’s own /metrics |
No shell, no files for the agent. It runs in ~/collectors-ai, not in the directory with ex-drop.values.yaml, and without Bash, Edit and Write. The manual variant’s rule is “find the loss without reading the config first”. This setup enforces that rule. It also keeps an agent with your kubeconfig from “fixing” the release itself.
💪Exercise💪 — steps
1. Agent access
Do steps 1–2 of Clustering split with an AI agent: a Viewer token for mcp-ro-<login>, the grafana MCP server in ~/collectors-ai, the agent started with claude --disallowedTools "Bash" "Edit" "Write". Already done and the token is younger than 8 h: start a new session (no --continue). Otherwise run claude mcp remove grafana and repeat both steps (token name mcp-ro-10b-…).
2. Deploy and send (terminal A)
Run steps 1 and 2 of the manual variant: deploy ex-alloy-drop-$ME with its ServiceMonitor, send 20 records (5 each of DEBUG, INFO, WARN, ERROR), and get twenty 200s. Note the time.
Fill in your prediction for the intended rule (drop DEBUG) and the actual column later:
| Stage | Intended | Agent | Explore / /metrics |
|---|---|---|---|
| receiver accepted | 20 | ||
filter noise in → out |
20 → 15 | ||
attributes loki_hints in → out |
15 → 15 | ||
| sent to Loki | 15 | ||
| in Loki | 15 |
3. Symptom only (terminal B)
Wait a minute for the first scrape. The prompt gives the facts a reporter would have, and no queries:
Via Grafana MCP, read only. Loki uid "loki", Prometheus uid "prometheus".
At <HH:MM> UTC a test client sent 20 log records (5 each DEBUG, INFO, WARN, ERROR) with
service.name "drop-<ME>" to the Alloy release ex-alloy-drop-<ME> in namespace "collectors-<ME>".
All 20 got HTTP 200. The pipeline is supposed to drop only DEBUG. The release is fresh.
Other Alloy releases may run in the same namespace.
1. How many records reached Loki? By severity, if you can derive it; say how. Give the LogQL.
2. From the collector's own metrics only: where were records lost? Chain: receiver accepted →
each processor in/out → sent to Loki. Verify that each metric and label exists before querying.
Separate this release from others by a label. Give every PromQL and the raw value.
3. Which component lost them, and what kind of rule produces exactly this loss?
You cannot see the config. Do not propose a fix yet.
Approve tool calls one at a time. Review:
| Where | Correct | Typical agent deviation |
|---|---|---|
| Loki count | sum(count_over_time({service_name="drop-<ME>"}[15m])) = 10, a window that covers the send |
counting lines returned by query_loki_logs (capped at 100 per call); a window that misses the send |
| release filter | namespace plus a label that identifies the release, found with list_prometheus_label_names |
namespace only, mixing in counters of another release |
| counter values | raw instant values: accepted 20, filter 20 → 10, attributes 10 → 10, loki_write_sent_entries_total 10 |
increase() / rate(): everything ≈ 0. Fresh counters start at 20 on the first scrape, so the jump from “no series” is never counted |
conclusion from increase() |
— | “no loss in the collector, Loki must have dropped them”: a confident wrong cause |
otelcol_processor_dropped_log_records_total |
empty, because only memory_limiter emits it |
empty read as “0 dropped”; a metric name the agent never verified |
| attribution | filter noise lost 10: DEBUG and INFO. The boundary is too high, everything below WARN |
“the filter dropped DEBUG as intended” without doing the arithmetic; blame on the attributes processor or Loki |
If the agent finds a filter-specific counter through list_prometheus_metric_names, check its value against incoming − outgoing. The difference works for every processor; a filter counter covers only this one.
Answer yourself:
- Which single number proves the loss happened in the collector and not in Loki?
- Why is
increase()wrong here, and why is it the right tool on a collector that has run for days?
Verify: the three PromQL queries and the LogQL from steps 3–4 of the manual variant, in Explore. Fill in the actual column. If the agent used increase(), give it your raw values and see whether it revises the conclusion or defends it.
4. Review the fix before you apply it
Paste only the filter block from ex-drop.values.yaml:
This is the filter block from the pipeline. The rest of the config is out of scope.
<paste the otelcol.processor.filter "noise" block>
The comment says "drop DEBUG noise". Propose the minimal change to the statement so it matches
the comment, using the OTel severity number ranges. Change nothing else. Tell me the counts I should
see after I resend the 20 records: filter in/out, and how to count the new batch in Loki.
| Check | Correct | Typical agent deviation |
|---|---|---|
| boundary | severity_number < SEVERITY_NUMBER_INFO (drops 1–8: TRACE and DEBUG) |
<= SEVERITY_NUMBER_DEBUG keeps DEBUG2–4 (6–8); severity_text == "DEBUG" misses debug and records without text |
| records without severity | says that < SEVERITY_NUMBER_INFO also drops severity_number 0 (unset). Offers severity_number >= SEVERITY_NUMBER_DEBUG and severity_number < SEVERITY_NUMBER_INFO if those must stay |
not mentioned |
| scope | only the statement changes | error_mode changed, the block rewritten as otelcol.processor.transform, the attributes processor “cleaned up” |
| predicted counts | filter 20 → 15, because the reload rebuilds the component and its counters restart; count the new batch in Loki over a window that starts after the resend |
40 → 25 cumulative; a 15-minute window that adds the first batch (10 + 15 = 25) |
Decide yourself which of the two statements you apply, and write down why: does this pipeline receive records without a severity?
5. Apply, resend, verify (terminal A, then B)
Edit the statement, upgrade, wait about a minute for the reload, and send the 20 records again (step 5 of the manual variant). Then:
I applied a fix and resent the 20 records at <HH:MM> UTC. Repeat the chain from step 3 for the new
batch only, and say how you kept the first batch out of the numbers.
Check: filter 20 → 15. In Loki, 15 records for the new batch, and no DEBUG among them. Confirm the filter counters with curl -s localhost:12346/metrics | grep otelcol_processor_ through the port-forward from the manual variant. Restart the port-forward if the pod changed.
Success criteria
- You found the loss from metrics: accepted 20, filter out 10, sent 10. The agent never saw the config, and your table matches Explore.
- You can explain why
increase()reads about0on fresh counters and whyotelcol_processor_dropped_*is empty. - You rejected or corrected at least one point of the agent’s analysis or fix, or you can show the check that confirmed it.
- You chose the fix statement yourself and can say what it does to records without a severity.
- After the fix, 15 of 20 records of the new batch reach Loki.
⭐Stretch: an alert the agent drafts, you break
Ask the agent for an alert expression that would have caught this in production without firing for every legitimate filter: the per-processor drop ratio (incoming − outgoing) / incoming against its own value offset 1d. Review it yourself:
- What does it return when
incomingis 0? Division by zero givesNaN, and the alert silently never fires. - What does it return on the first day of a new processor, when
offset 1dhas no data? Would it have caught this colleague’s rule, deployed with the bug from day one? - Is it grouped by release and processor, or does one sum hide a single bad filter?
Test each of your answers with query_prometheus through the agent or in Explore, not by argument.
Clean up when done: helm -n $NS uninstall ex-alloy-drop-$ME. Remove the agent setup as at the end of Clustering split with an AI agent.
Related lessons
- Silent drop hunt — the same exercise without an agent
- Self-monitoring and silent drops
- Pipeline model — processor order
- LogQL