đȘExerciseđȘ â PromQL with an AI agent: judge every query it writes
Same questions and the same end state as sections 2â4 of the manual variant: node memory, memory-hungry jobs, an SLO ranking, percentiles, cardinality, a disk forecast, an âabsentâ expression, native histograms. An AI agent writes and runs the PromQL through the labâs Grafana MCP server. You ask in plain language, predict the result, read every query and verify it yourself. Section 1 (selectors, range vectors,
ratevsincrease, intervals, histogram buckets), 3.1 (iratevsrateon a chart) and 3.4 ($__rate_interval) stay manual: they are about what you see on a chart, and the MCP query tools do not expand Grafana variables.
Goal
- You ask questions about the labâs metrics in plain language. The agent finds metrics and labels, writes PromQL and runs it (
list_prometheus_metric_names,list_prometheus_label_names,list_prometheus_label_values,list_prometheus_metric_metadata,query_prometheus,query_prometheus_histogram). - Before each answer you predict the metric, the aggregation and the shape of the result.
- You review each query: metric choice and what it actually measures, label matchers,
by/without, therate()window against the sample interval,lein histograms, units, and whether the number answers the question. - You verify in Grafana Explore, the Prometheus UI and the Prometheus API, and work out what the agent assumed when the question was ambiguous.
Environment: workshop cluster, Grafana Explore, Prometheus UI, a terminal, an MCP-capable AI client.
Prerequisites
- Section 1 of the manual variant done, plus 3.1 and 3.4. You need your own feel for
rate()and windows to review the agent. - Your lab login (
user<N>), its Grafana password and the MCP server token, all from the lab handout. - An AI client that supports remote MCP servers over Streamable HTTP with custom headers. The examples use Claude Code.
curl,jq, Git Bash or WSL.- An empty working directory for the agent. In the course repo the agent reads
CLAUDE.md, which describes the lab and its faults. Do not start it there.
Who does what:
| Step | Who | Why |
|---|---|---|
selectors, range vectors, rate vs increase on a chart, Min step, buckets (manual 1), irate (3.1), $__rate_interval (3.4) |
you, in Explore and the Prometheus UI | about charts and Grafana macros; query_prometheus sends the query straight to Prometheus without expanding $__rate_interval |
| metric and label discovery, PromQL for every question, the numbers | agent | list_prometheus_*, query_prometheus, query_prometheus_histogram |
| predictions, review, verdicts | you | the agentâs query is a claim until you have read it |
| charts, legends, cross-checks | you, in Explore and the Prometheus API | MCP returns numbers, not charts |
đȘExerciseđȘ â steps
Step 1 â a read-only token for the agent
The agent only reads, so its service account is Viewer with no folder permissions. The script reuses mcp-<login>-ro if another exercise created it and issues a new token valid for 8 h.
export GRAFANA_URL=https://grafana.workshop2.indexoutofrange.com
LOGIN=<login>
read -rsp "Grafana password: " P && GRAFANA_AUTH="$LOGIN:$P" && echo
H=(-u "$GRAFANA_AUTH" -H "Content-Type: application/json")
SA_NAME="mcp-$LOGIN-ro"
SA=$(curl -sS "${H[@]}" "$GRAFANA_URL/api/serviceaccounts/search?query=$SA_NAME" \
| jq -j --arg n "$SA_NAME" '.serviceAccounts[] | select(.name==$n) | .id')
if [ -z "$SA" ]; then
SA=$(curl -sS "${H[@]}" -X POST "$GRAFANA_URL/api/serviceaccounts" \
-d "{\"name\":\"$SA_NAME\",\"role\":\"Viewer\"}" | jq -j .id)
fi
export GRAFANA_SERVICE_ACCOUNT_TOKEN=$(curl -sS "${H[@]}" -X POST "$GRAFANA_URL/api/serviceaccounts/$SA/tokens" \
-d "{\"name\":\"mcp-prom11-$(date +%s)\",\"secondsToLive\":28800}" | jq -j .key)
The password stays in a shell variable, without export. claude started from this terminal inherits exported variables, and an agent with a Bash tool could then call the Grafana API as your Admin account, bypassing MCP.
Check the tokenâs boundary. A dashboard write must fail:
curl -sS -o /dev/null -w '%{http_code}\n' -H "Authorization: Bearer $GRAFANA_SERVICE_ACCOUNT_TOKEN" -H 'Content-Type: application/json' \
-X POST "$GRAFANA_URL/api/dashboards/db" -d "{\"dashboard\":{\"title\":\"$SA_NAME probe\",\"panels\":[]}}"
Check that: the response is 403 and $GRAFANA_SERVICE_ACCOUNT_TOKEN starts with glsa_.
Step 2 â connect Grafana MCP in an empty directory
mkdir -p ~/promql-11 && cd ~/promql-11
export GRAFANA_MCP_URL=https://mcp.workshop2.indexoutofrange.com/mcp
read -rsp "MCP server token: " MCP_SERVER_TOKEN && export MCP_SERVER_TOKEN && echo
claude mcp add --transport http grafana "$GRAFANA_MCP_URL" \
--header "Authorization: Bearer $MCP_SERVER_TOKEN" \
--header "X-Grafana-Service-Account-Token: $GRAFANA_SERVICE_ACCOUNT_TOKEN"
claude mcp list
- Claude Codeâs default
localscope binds the server to this directory. Do not use--scope project: it writes the tokens to.mcp.json. - Other clients need the same URL and the same two headers. Without
X-Grafana-Service-Account-Tokenthe server falls back to a shared Viewer account, and your token is not the one being tested.
Start claude and run a smoke test:
Via Grafana MCP, Prometheus datasource uid "prometheus": list the values of the label server_address
for the metric http_client_request_duration_seconds_count, grouped by job (one instant query is fine),
and the label names of traces_spanmetrics_latency_bucket. Analyse nothing.
Check that: the agent called list_prometheus_label_names and query_prometheus (or list_prometheus_label_values with matches). For opentelemetry-demo/payment and opentelemetry-demo/frontend the only server_address is the Pyroscope service; checkout calls shipping and email. The span metrics labels include service_name, span_kind, status_code and source.
Answer in your notes before going on: http_client_request_duration_seconds is a client histogram. For payment, what does it measure in this lab? Is that the response time of payment?
Step 3 â predict before you ask
Fill the second and third columns now, from the Overview and Metric Cardinality lessons and section 1. You fill the rest in steps 4â6.
| # | Question (plain language) | Your query sketch | Predicted result shape (series, labels, unit, magnitude) | Agentâs query | Verified | Verdict |
|---|---|---|---|---|---|---|
| P1 | Memory used per Kubernetes node, in GB, labelled by node name | Â | Â | Â | Â | Â |
| P2 | Top 5 most memory-hungry processes, grouped by job |
 |  |  |  |  |
| P3 | Three services furthest from â90% of requests under 10 msâ | Â | Â | Â | Â | Â |
| P4 | p90 response time of checkout |
 |  |  |  |  |
| P5 | Top 20 metrics by number of series, and the labels behind the top 3 | Â | Â | Â | Â | Â |
| P6 | Free disk per node in 4 h, from the last hourâs trend, in GB | Â | Â | Â | Â | Â |
| P7 | An expression that returns 1 when payment stops sending app_payment_transactions_total |
 |  |  |  |  |
| P8 | p50, p95, p99 per service; the largest p50âp99 gap | Â | Â | Â | Â | Â |
| P9 | Share of Prometheus /api/v1/write requests under 100 ms, and their p99, from the native histogram |
 |  |  |  |  |
Step 4 â the spec
The prompt states what to find and what to report. It contains no PromQL.
Via Grafana MCP, read-only. Prometheus datasource uid "prometheus", last 1 hour unless stated.
Before writing a query, check metric names, metadata and label values. For every question report:
the PromQL, the tool and its parameters (queryType, startTime/endTime, stepSeconds), the number of
series, the result with its unit, and every assumption you made (which metric and what it measures,
which labels, the rate() window and why it is long enough for this metric's sample interval).
If a question can be read two ways, say so and pick one explicitly. NaN or an empty result is a
finding to explain, not a value to report.
P1. Memory used on each Kubernetes node (not pod), in GB, with the node name as the series label.
P2. The 5 most memory-hungry processes, grouped by job.
P3. The 3 services with the biggest problem meeting the SLO "90% of requests complete in under 10 ms".
The SLO is per service, not per instance. The latency must be the service's own response time.
P4. The 90th percentile response time of the checkout service.
P5. The 20 metrics with the most series, then the labels that drive the series count of the top 3.
P6. Free disk space on each node 4 hours from now, extrapolated from the last hour, in GB.
P7. An expression that returns 1 when app_payment_transactions_total for job opentelemetry-demo/payment
stops existing. What does it return now?
P8. p50, p95 and p99 of the response time for every OTel Demo service. Which service has the largest
gap between p50 and p99?
P9. From the native histogram prometheus_http_request_duration_seconds: the share of requests to
/api/v1/write faster than 100 ms, and their p99, over the last 5 minutes.
Approve tool calls one at a time and read the arguments. Copy each PromQL into the âAgentâs queryâ column before you read the agentâs answer.
Check that: every answer has a query, parameters, a series count, a result with a unit and a list of assumptions.
Step 5 â review the agent
A deviation is not wrong by itself if the agent said what it assumed and the assumption is defensible. A silent one is.
| P | Correct | Typical agent deviation |
|---|---|---|
| P1 | (node_memory_MemTotal_bytes - node_memory_MemAvailable_bytes) / 1e9 * on (instance) group_left (nodename) node_uname_info; one series per node with nodename; the agent says GB (1e9) or GiB (2^30) |
MemTotal - MemFree: counts the page cache as used, several times larger; legend instance (an IP and port) presented as the node name; container_memory_working_set_bytes: not collected in this lab, empty result |
| P2 | topk(5, sum by (job) (process_resident_memory_bytes)), and the agent says that a job with several replicas is summed |
topk(5, process_resident_memory_bytes): pods, not jobs; max by (job) without saying so; the claim that a chart of this query shows exactly 5 jobs (in a range query topk picks per step) |
| P3 | server-side latency: traces_spanmetrics_latency_bucket{span_kind="SPAN_KIND_SERVER", source="tempo"} by service_name, and the agent says the bucket boundaries are 0.008 and 0.016, so â10 msâ is approximated, and which side it took |
http_client_request_duration_seconds_bucket{le="0.01"} by job: ranks outgoing calls, and for payment and frontend that means profile uploads to Pyroscope; le="0.01" on span metrics: matches only a few llm series from Beyla (source="beyla"), a silently partial result; rate(...[1m]) on the OTel Demo app metrics: they arrive every 60 s, so the result is empty |
| P4 | histogram_quantile(0.9, sum by (le) (rate(traces_spanmetrics_latency_bucket{service_name="checkout", span_kind="SPAN_KIND_SERVER", source="tempo"}[5m]))); the agent says that checkoutâs only server span is oteldemo.CheckoutService/PlaceOrder |
the client histogram of checkout: the latency of its calls to shipping and email; query_prometheus_histogram with percentile = 0.9: the tool expects 90, so the result is the 0.9th percentile |
| P5 | topk(20, count by (__name__) ({__name__=~".+"})), then count by (<label>) (<metric>) or count(count by (<label>) (<metric>)) per candidate label |
list_prometheus_metric_names with its default limit (10) presented as âall metricsâ; prometheus_tsdb_head_series presented as the per-metric answer; the âproblem labelâ named from training data, not from a query |
| P6 | predict_linear(node_filesystem_avail_bytes{mountpoint="/"}[1h], 4*3600) / 1e9 (or another explicit mount), joined to nodename |
no mountpoint / fstype matcher: seven filesystems per node including tmpfs and /boot/efi, read as âthe nodeâs diskâ; predict_linear on a counter |
| P7 | absent(app_payment_transactions_total{job="opentelemetry-demo/payment"}); now: empty result; the agent notes it can take up to the 5 min lookback to turn 1 |
absent(rate(...)) or ... == 0: different questions; âreturns 0â (it returns no series) |
| P8 | histogram_quantile(Q, sum by (le, service_name) (rate(traces_spanmetrics_latency_bucket{span_kind="SPAN_KIND_SERVER", source="tempo"}[5m]))) for 0.5, 0.95, 0.99 |
no le in by: no result; source not filtered: llm mixes two bucket layouts; NaN (no requests in the window) reported as a value; a p99 equal to the top finite bucket (4.1) read as a measurement, not as âabove 4.1 sâ |
| P9 | histogram_fraction(0, 0.1, rate(prometheus_http_request_duration_seconds{job="prometheus-native-histograms", handler="/api/v1/write"}[5m])) and histogram_quantile(0.99, rate(âŠsameâŠ[5m])) |
no job matcher with classic _bucket / _count: the same Prometheus is also scraped as a classic histogram by another job, counts double; sum by (le) on a native histogram; no rate(): the share since Prometheus started |
Answer in your notes:
- P1. How large is the difference between
MemFreeandMemAvailableper node? Which one answers âmemory usedâ, and why? - P3. Name the three jobs the client-metric query ranks worst and what each of them is actually calling (
server_address). Does that answer the SLO question? - P3 and P8. What is the sample interval of
traces_spanmetrics_*and ofhttp_client_request_duration_seconds_*? Whichrate()window is the smallest safe one for each? - P8. Why can
histogram_quantilenever return a value between4.1and+Inffor span metrics?
Step 6 â verify independently
- In Explore â Prometheus, run the agentâs P1, P3 and P8 queries yourself. For P1 set Legend to ``. For P8 put p50 and p99 on one chart (queries A and B).
- The raw API, with the agentâs token (Viewer is enough for the datasource proxy):
pq() { curl -sS -G -H "Authorization: Bearer $GRAFANA_SERVICE_ACCOUNT_TOKEN" \
"$GRAFANA_URL/api/datasources/proxy/uid/prometheus/api/v1/query" --data-urlencode "query=$1" \
| jq -c '.data.result[] | [.metric, .value[1]]'; }
pq 'count by (job, server_address) (http_client_request_duration_seconds_count)'
pq 'count by (le) (traces_spanmetrics_latency_bucket{source="tempo"})'
pq 'count by (service_name) (traces_spanmetrics_latency_bucket{le="0.01"})'
pq 'sum by (instance) (node_memory_MemAvailable_bytes - node_memory_MemFree_bytes) / 1e9'
- Sample intervals in the Prometheus UI (
http://prometheus.workshop2.indexoutofrange.com/), Table tab:traces_spanmetrics_calls_total{service_name="payment"}[2m]andapp_payment_transactions_total[5m]. Compare the timestamps.
Check that:
paymentandfrontendclient series point only at Pyroscope, so the client-metric SLO ranking is not about those services,- span metrics from Tempo have no
0.01bucket, andle="0.01"returns onlyllm, - span metrics arrive every 15 s and the OTel Demo app metrics every 60 s, so
rate(app_payment_transactions_total[1m])is empty, - the agentâs P1 values match yours and the legend shows node names, not IPs.
Fill the âVerifiedâ and âVerdictâ columns. A verdict is correct, correct with stated assumption, silently wrong or answers another question.
Step 7 â an ambiguous question on purpose
Start a new agent session (/clear) and type only:
How much memory does the cluster use?
When it answers, ask:
List every assumption you made: nodes or pods, which metric and why it exists in this Prometheus,
MemFree or MemAvailable, usage or requests, sum or per node, GB or GiB, instant or a range.
| Decision | The agentâs choice | Alternatives | What each gives (Explore or pq) |
|---|---|---|---|
| level | Â | nodes (node_memory_*), pods |
 |
| metric exists here? | Â | node_memory_*, kube_pod_container_resource_requests, container_memory_* (not collected) |
 |
| âusedâ | Â | MemTotal - MemAvailable, MemTotal - MemFree |
 |
| usage or reservation | Â | actual usage, kube_pod_container_resource_requests{resource="memory"} |
 |
| unit | Â | GB, GiB | Â |
Then write a one-sentence question that removes every ambiguity, ask it in the same session and compare with your pq result.
Check that: you can name the choice the agent made silently. Typical: a container_memory_* query that returns nothing, explained as âno dataâ; requests reported as usage; MemFree without a word about the cache.
Step 8 â planted bugs: does the agent catch them?
Write your own verdict for each query first. Then paste them to the agent:
Each query below is supposed to answer the question next to it. Without running it, say whether it
answers the question and predict the result. Then run it via query_prometheus and compare.
a) "Payments per second, by currency":
sum by (app_payment_currency) (rate(app_payment_transactions_total[1m]))
b) "p95 latency per service":
histogram_quantile(0.95, sum by (service_name) (rate(traces_spanmetrics_latency_bucket{span_kind="SPAN_KIND_SERVER"}[5m])))
c) "Error ratio of payment's server spans":
sum(rate(traces_spanmetrics_calls_total{service_name="payment", status_code="STATUS_CODE_ERROR"}[5m]))
/ sum(rate(traces_spanmetrics_calls_total{service_name="payment"}[5m]))
| Query | Bug | Agent caught it before running? | After running? |
|---|---|---|---|
| a | 1 min window on a metric that arrives every 60 s: empty | Â | Â |
| b | le missing from by: no result |
 |  |
| c | no span_kind="SPAN_KIND_SERVER": paymentâs internal and client spans never carry the error and dilute the denominator |
 |  |
Check that: you found all bugs yourself before the agent answered, and you recorded which ones the agent only noticed after it had seen the result, or not at all.
Step 9 â cleanup
If you continue with another agent variant (LogQL, TraceQL), skip this step: the token is valid for 8 h. Otherwise, in the same terminal as step 1:
cd ~/promql-11 && claude mcp remove grafana
curl -sS "${H[@]}" -X DELETE "$GRAFANA_URL/api/serviceaccounts/$SA" | jq -c .
unset P GRAFANA_AUTH H
Success criteria
- The prediction table is complete for P1âP9, and every row has a verdict backed by your own Explore, Prometheus UI or API result.
- You can explain, without the agent, why the client-histogram SLO ranking does not measure the servicesâ response time in this lab, and which metric and bucket you would use instead.
- You know the sample interval of span metrics and of the OTel Demo app metrics, and the smallest safe
rate()window for each. - P1 in Explore shows used memory per node with `` legends, and you can say how much of
MemTotal - MemFreeis page cache. - For the ambiguous question in step 7 you listed the agentâs silent assumptions and wrote a spec sentence that removes them.
- You found the bugs in step 8 before the agent did.
âStretch: a Grafana macro sent to Prometheus (failure mode)
Ask the agent: Run the SLO query from P3 with [$__rate_interval] instead of a fixed window and give me the result. query_prometheus sends the expression to Prometheus as it is, and Prometheus does not know $__rate_interval.
- Correct: the agent says the MCP query tools do not expand Grafana macros, runs the query with an explicit window it names (at least 4Ă the sample interval), and says that a panel would use
$__rate_interval. - Deviation: it reports the parse error as âno dataâ, or swaps in a window without saying so and presents the result as the
$__rate_intervalone.
Compare its window with the one Grafana uses for the same query in Explore: Query inspector â Query â Refresh, executedQueryString in the response.
Related lessons
- đȘExerciseđȘ â PromQL â the same exercise without the agent
- Overview
- Metric Cardinality
- LogQL with an AI agent â the same review discipline for logs
- TraceQL with an AI agent â the same review discipline for traces