đŸ’ȘExerciseđŸ’Ș — PromQL with an AI agent: judge every query it writes

Same questions and the same end state as sections 2–4 of the manual variant: node memory, memory-hungry jobs, an SLO ranking, percentiles, cardinality, a disk forecast, an “absent” expression, native histograms. An AI agent writes and runs the PromQL through the lab’s Grafana MCP server. You ask in plain language, predict the result, read every query and verify it yourself. Section 1 (selectors, range vectors, rate vs increase, intervals, histogram buckets), 3.1 (irate vs rate on a chart) and 3.4 ($__rate_interval) stay manual: they are about what you see on a chart, and the MCP query tools do not expand Grafana variables.

Goal

  • You ask questions about the lab’s metrics in plain language. The agent finds metrics and labels, writes PromQL and runs it (list_prometheus_metric_names, list_prometheus_label_names, list_prometheus_label_values, list_prometheus_metric_metadata, query_prometheus, query_prometheus_histogram).
  • Before each answer you predict the metric, the aggregation and the shape of the result.
  • You review each query: metric choice and what it actually measures, label matchers, by / without, the rate() window against the sample interval, le in histograms, units, and whether the number answers the question.
  • You verify in Grafana Explore, the Prometheus UI and the Prometheus API, and work out what the agent assumed when the question was ambiguous.

Environment: workshop cluster, Grafana Explore, Prometheus UI, a terminal, an MCP-capable AI client.

Prerequisites

  • Section 1 of the manual variant done, plus 3.1 and 3.4. You need your own feel for rate() and windows to review the agent.
  • Your lab login (user<N>), its Grafana password and the MCP server token, all from the lab handout.
  • An AI client that supports remote MCP servers over Streamable HTTP with custom headers. The examples use Claude Code.
  • curl, jq, Git Bash or WSL.
  • An empty working directory for the agent. In the course repo the agent reads CLAUDE.md, which describes the lab and its faults. Do not start it there.

Who does what:

Step Who Why
selectors, range vectors, rate vs increase on a chart, Min step, buckets (manual 1), irate (3.1), $__rate_interval (3.4) you, in Explore and the Prometheus UI about charts and Grafana macros; query_prometheus sends the query straight to Prometheus without expanding $__rate_interval
metric and label discovery, PromQL for every question, the numbers agent list_prometheus_*, query_prometheus, query_prometheus_histogram
predictions, review, verdicts you the agent’s query is a claim until you have read it
charts, legends, cross-checks you, in Explore and the Prometheus API MCP returns numbers, not charts

đŸ’ȘExerciseđŸ’Ș — steps

Step 1 — a read-only token for the agent

The agent only reads, so its service account is Viewer with no folder permissions. The script reuses mcp-<login>-ro if another exercise created it and issues a new token valid for 8 h.

export GRAFANA_URL=https://grafana.workshop2.indexoutofrange.com
LOGIN=<login>
read -rsp "Grafana password: " P && GRAFANA_AUTH="$LOGIN:$P" && echo
H=(-u "$GRAFANA_AUTH" -H "Content-Type: application/json")
SA_NAME="mcp-$LOGIN-ro"

SA=$(curl -sS "${H[@]}" "$GRAFANA_URL/api/serviceaccounts/search?query=$SA_NAME" \
  | jq -j --arg n "$SA_NAME" '.serviceAccounts[] | select(.name==$n) | .id')
if [ -z "$SA" ]; then
  SA=$(curl -sS "${H[@]}" -X POST "$GRAFANA_URL/api/serviceaccounts" \
    -d "{\"name\":\"$SA_NAME\",\"role\":\"Viewer\"}" | jq -j .id)
fi
export GRAFANA_SERVICE_ACCOUNT_TOKEN=$(curl -sS "${H[@]}" -X POST "$GRAFANA_URL/api/serviceaccounts/$SA/tokens" \
  -d "{\"name\":\"mcp-prom11-$(date +%s)\",\"secondsToLive\":28800}" | jq -j .key)

The password stays in a shell variable, without export. claude started from this terminal inherits exported variables, and an agent with a Bash tool could then call the Grafana API as your Admin account, bypassing MCP.

Check the token’s boundary. A dashboard write must fail:

curl -sS -o /dev/null -w '%{http_code}\n' -H "Authorization: Bearer $GRAFANA_SERVICE_ACCOUNT_TOKEN" -H 'Content-Type: application/json' \
  -X POST "$GRAFANA_URL/api/dashboards/db" -d "{\"dashboard\":{\"title\":\"$SA_NAME probe\",\"panels\":[]}}"

Check that: the response is 403 and $GRAFANA_SERVICE_ACCOUNT_TOKEN starts with glsa_.

Step 2 — connect Grafana MCP in an empty directory

mkdir -p ~/promql-11 && cd ~/promql-11
export GRAFANA_MCP_URL=https://mcp.workshop2.indexoutofrange.com/mcp
read -rsp "MCP server token: " MCP_SERVER_TOKEN && export MCP_SERVER_TOKEN && echo
claude mcp add --transport http grafana "$GRAFANA_MCP_URL" \
  --header "Authorization: Bearer $MCP_SERVER_TOKEN" \
  --header "X-Grafana-Service-Account-Token: $GRAFANA_SERVICE_ACCOUNT_TOKEN"
claude mcp list
  • Claude Code’s default local scope binds the server to this directory. Do not use --scope project: it writes the tokens to .mcp.json.
  • Other clients need the same URL and the same two headers. Without X-Grafana-Service-Account-Token the server falls back to a shared Viewer account, and your token is not the one being tested.

Start claude and run a smoke test:

Via Grafana MCP, Prometheus datasource uid "prometheus": list the values of the label server_address
for the metric http_client_request_duration_seconds_count, grouped by job (one instant query is fine),
and the label names of traces_spanmetrics_latency_bucket. Analyse nothing.

Check that: the agent called list_prometheus_label_names and query_prometheus (or list_prometheus_label_values with matches). For opentelemetry-demo/payment and opentelemetry-demo/frontend the only server_address is the Pyroscope service; checkout calls shipping and email. The span metrics labels include service_name, span_kind, status_code and source.

Answer in your notes before going on: http_client_request_duration_seconds is a client histogram. For payment, what does it measure in this lab? Is that the response time of payment?

Step 3 — predict before you ask

Fill the second and third columns now, from the Overview and Metric Cardinality lessons and section 1. You fill the rest in steps 4–6.

# Question (plain language) Your query sketch Predicted result shape (series, labels, unit, magnitude) Agent’s query Verified Verdict
P1 Memory used per Kubernetes node, in GB, labelled by node name          
P2 Top 5 most memory-hungry processes, grouped by job          
P3 Three services furthest from “90% of requests under 10 ms”          
P4 p90 response time of checkout          
P5 Top 20 metrics by number of series, and the labels behind the top 3          
P6 Free disk per node in 4 h, from the last hour’s trend, in GB          
P7 An expression that returns 1 when payment stops sending app_payment_transactions_total          
P8 p50, p95, p99 per service; the largest p50–p99 gap          
P9 Share of Prometheus /api/v1/write requests under 100 ms, and their p99, from the native histogram          

Step 4 — the spec

The prompt states what to find and what to report. It contains no PromQL.

Via Grafana MCP, read-only. Prometheus datasource uid "prometheus", last 1 hour unless stated.
Before writing a query, check metric names, metadata and label values. For every question report:
the PromQL, the tool and its parameters (queryType, startTime/endTime, stepSeconds), the number of
series, the result with its unit, and every assumption you made (which metric and what it measures,
which labels, the rate() window and why it is long enough for this metric's sample interval).
If a question can be read two ways, say so and pick one explicitly. NaN or an empty result is a
finding to explain, not a value to report.

P1. Memory used on each Kubernetes node (not pod), in GB, with the node name as the series label.
P2. The 5 most memory-hungry processes, grouped by job.
P3. The 3 services with the biggest problem meeting the SLO "90% of requests complete in under 10 ms".
    The SLO is per service, not per instance. The latency must be the service's own response time.
P4. The 90th percentile response time of the checkout service.
P5. The 20 metrics with the most series, then the labels that drive the series count of the top 3.
P6. Free disk space on each node 4 hours from now, extrapolated from the last hour, in GB.
P7. An expression that returns 1 when app_payment_transactions_total for job opentelemetry-demo/payment
    stops existing. What does it return now?
P8. p50, p95 and p99 of the response time for every OTel Demo service. Which service has the largest
    gap between p50 and p99?
P9. From the native histogram prometheus_http_request_duration_seconds: the share of requests to
    /api/v1/write faster than 100 ms, and their p99, over the last 5 minutes.

Approve tool calls one at a time and read the arguments. Copy each PromQL into the “Agent’s query” column before you read the agent’s answer.

Check that: every answer has a query, parameters, a series count, a result with a unit and a list of assumptions.

Step 5 — review the agent

A deviation is not wrong by itself if the agent said what it assumed and the assumption is defensible. A silent one is.

P Correct Typical agent deviation
P1 (node_memory_MemTotal_bytes - node_memory_MemAvailable_bytes) / 1e9 * on (instance) group_left (nodename) node_uname_info; one series per node with nodename; the agent says GB (1e9) or GiB (2^30) MemTotal - MemFree: counts the page cache as used, several times larger; legend instance (an IP and port) presented as the node name; container_memory_working_set_bytes: not collected in this lab, empty result
P2 topk(5, sum by (job) (process_resident_memory_bytes)), and the agent says that a job with several replicas is summed topk(5, process_resident_memory_bytes): pods, not jobs; max by (job) without saying so; the claim that a chart of this query shows exactly 5 jobs (in a range query topk picks per step)
P3 server-side latency: traces_spanmetrics_latency_bucket{span_kind="SPAN_KIND_SERVER", source="tempo"} by service_name, and the agent says the bucket boundaries are 0.008 and 0.016, so “10 ms” is approximated, and which side it took http_client_request_duration_seconds_bucket{le="0.01"} by job: ranks outgoing calls, and for payment and frontend that means profile uploads to Pyroscope; le="0.01" on span metrics: matches only a few llm series from Beyla (source="beyla"), a silently partial result; rate(...[1m]) on the OTel Demo app metrics: they arrive every 60 s, so the result is empty
P4 histogram_quantile(0.9, sum by (le) (rate(traces_spanmetrics_latency_bucket{service_name="checkout", span_kind="SPAN_KIND_SERVER", source="tempo"}[5m]))); the agent says that checkout’s only server span is oteldemo.CheckoutService/PlaceOrder the client histogram of checkout: the latency of its calls to shipping and email; query_prometheus_histogram with percentile = 0.9: the tool expects 90, so the result is the 0.9th percentile
P5 topk(20, count by (__name__) ({__name__=~".+"})), then count by (<label>) (<metric>) or count(count by (<label>) (<metric>)) per candidate label list_prometheus_metric_names with its default limit (10) presented as “all metrics”; prometheus_tsdb_head_series presented as the per-metric answer; the “problem label” named from training data, not from a query
P6 predict_linear(node_filesystem_avail_bytes{mountpoint="/"}[1h], 4*3600) / 1e9 (or another explicit mount), joined to nodename no mountpoint / fstype matcher: seven filesystems per node including tmpfs and /boot/efi, read as “the node’s disk”; predict_linear on a counter
P7 absent(app_payment_transactions_total{job="opentelemetry-demo/payment"}); now: empty result; the agent notes it can take up to the 5 min lookback to turn 1 absent(rate(...)) or ... == 0: different questions; “returns 0” (it returns no series)
P8 histogram_quantile(Q, sum by (le, service_name) (rate(traces_spanmetrics_latency_bucket{span_kind="SPAN_KIND_SERVER", source="tempo"}[5m]))) for 0.5, 0.95, 0.99 no le in by: no result; source not filtered: llm mixes two bucket layouts; NaN (no requests in the window) reported as a value; a p99 equal to the top finite bucket (4.1) read as a measurement, not as “above 4.1 s”
P9 histogram_fraction(0, 0.1, rate(prometheus_http_request_duration_seconds{job="prometheus-native-histograms", handler="/api/v1/write"}[5m])) and histogram_quantile(0.99, rate(
same
[5m])) no job matcher with classic _bucket / _count: the same Prometheus is also scraped as a classic histogram by another job, counts double; sum by (le) on a native histogram; no rate(): the share since Prometheus started

Answer in your notes:

  1. P1. How large is the difference between MemFree and MemAvailable per node? Which one answers “memory used”, and why?
  2. P3. Name the three jobs the client-metric query ranks worst and what each of them is actually calling (server_address). Does that answer the SLO question?
  3. P3 and P8. What is the sample interval of traces_spanmetrics_* and of http_client_request_duration_seconds_*? Which rate() window is the smallest safe one for each?
  4. P8. Why can histogram_quantile never return a value between 4.1 and +Inf for span metrics?

Step 6 — verify independently

  1. In Explore → Prometheus, run the agent’s P1, P3 and P8 queries yourself. For P1 set Legend to ``. For P8 put p50 and p99 on one chart (queries A and B).
  2. The raw API, with the agent’s token (Viewer is enough for the datasource proxy):
pq() { curl -sS -G -H "Authorization: Bearer $GRAFANA_SERVICE_ACCOUNT_TOKEN" \
  "$GRAFANA_URL/api/datasources/proxy/uid/prometheus/api/v1/query" --data-urlencode "query=$1" \
  | jq -c '.data.result[] | [.metric, .value[1]]'; }

pq 'count by (job, server_address) (http_client_request_duration_seconds_count)'
pq 'count by (le) (traces_spanmetrics_latency_bucket{source="tempo"})'
pq 'count by (service_name) (traces_spanmetrics_latency_bucket{le="0.01"})'
pq 'sum by (instance) (node_memory_MemAvailable_bytes - node_memory_MemFree_bytes) / 1e9'
  1. Sample intervals in the Prometheus UI (http://prometheus.workshop2.indexoutofrange.com/), Table tab: traces_spanmetrics_calls_total{service_name="payment"}[2m] and app_payment_transactions_total[5m]. Compare the timestamps.

Check that:

  • payment and frontend client series point only at Pyroscope, so the client-metric SLO ranking is not about those services,
  • span metrics from Tempo have no 0.01 bucket, and le="0.01" returns only llm,
  • span metrics arrive every 15 s and the OTel Demo app metrics every 60 s, so rate(app_payment_transactions_total[1m]) is empty,
  • the agent’s P1 values match yours and the legend shows node names, not IPs.

Fill the “Verified” and “Verdict” columns. A verdict is correct, correct with stated assumption, silently wrong or answers another question.

Step 7 — an ambiguous question on purpose

Start a new agent session (/clear) and type only:

How much memory does the cluster use?

When it answers, ask:

List every assumption you made: nodes or pods, which metric and why it exists in this Prometheus,
MemFree or MemAvailable, usage or requests, sum or per node, GB or GiB, instant or a range.
Decision The agent’s choice Alternatives What each gives (Explore or pq)
level   nodes (node_memory_*), pods  
metric exists here?   node_memory_*, kube_pod_container_resource_requests, container_memory_* (not collected)  
“used”   MemTotal - MemAvailable, MemTotal - MemFree  
usage or reservation   actual usage, kube_pod_container_resource_requests{resource="memory"}  
unit   GB, GiB  

Then write a one-sentence question that removes every ambiguity, ask it in the same session and compare with your pq result.

Check that: you can name the choice the agent made silently. Typical: a container_memory_* query that returns nothing, explained as “no data”; requests reported as usage; MemFree without a word about the cache.

Step 8 — planted bugs: does the agent catch them?

Write your own verdict for each query first. Then paste them to the agent:

Each query below is supposed to answer the question next to it. Without running it, say whether it
answers the question and predict the result. Then run it via query_prometheus and compare.

a) "Payments per second, by currency":
   sum by (app_payment_currency) (rate(app_payment_transactions_total[1m]))
b) "p95 latency per service":
   histogram_quantile(0.95, sum by (service_name) (rate(traces_spanmetrics_latency_bucket{span_kind="SPAN_KIND_SERVER"}[5m])))
c) "Error ratio of payment's server spans":
   sum(rate(traces_spanmetrics_calls_total{service_name="payment", status_code="STATUS_CODE_ERROR"}[5m]))
   / sum(rate(traces_spanmetrics_calls_total{service_name="payment"}[5m]))
Query Bug Agent caught it before running? After running?
a 1 min window on a metric that arrives every 60 s: empty    
b le missing from by: no result    
c no span_kind="SPAN_KIND_SERVER": payment’s internal and client spans never carry the error and dilute the denominator    

Check that: you found all bugs yourself before the agent answered, and you recorded which ones the agent only noticed after it had seen the result, or not at all.

Step 9 — cleanup

If you continue with another agent variant (LogQL, TraceQL), skip this step: the token is valid for 8 h. Otherwise, in the same terminal as step 1:

cd ~/promql-11 && claude mcp remove grafana
curl -sS "${H[@]}" -X DELETE "$GRAFANA_URL/api/serviceaccounts/$SA" | jq -c .
unset P GRAFANA_AUTH H

Success criteria

  • The prediction table is complete for P1–P9, and every row has a verdict backed by your own Explore, Prometheus UI or API result.
  • You can explain, without the agent, why the client-histogram SLO ranking does not measure the services’ response time in this lab, and which metric and bucket you would use instead.
  • You know the sample interval of span metrics and of the OTel Demo app metrics, and the smallest safe rate() window for each.
  • P1 in Explore shows used memory per node with `` legends, and you can say how much of MemTotal - MemFree is page cache.
  • For the ambiguous question in step 7 you listed the agent’s silent assumptions and wrote a spec sentence that removes them.
  • You found the bugs in step 8 before the agent did.

⭐Stretch: a Grafana macro sent to Prometheus (failure mode)

Ask the agent: Run the SLO query from P3 with [$__rate_interval] instead of a fixed window and give me the result. query_prometheus sends the expression to Prometheus as it is, and Prometheus does not know $__rate_interval.

  • Correct: the agent says the MCP query tools do not expand Grafana macros, runs the query with an explicit window it names (at least 4× the sample interval), and says that a panel would use $__rate_interval.
  • Deviation: it reports the parse error as “no data”, or swaps in a window without saying so and presents the result as the $__rate_interval one.

Compare its window with the one Grafana uses for the same query in Explore: Query inspector → Query → Refresh, executedQueryString in the response.

results matching ""

    No results matching ""