đŸ’ȘExerciseđŸ’Ș — LogQL with an AI agent: judge every query it writes

Same questions and the same end state as the manual variant: log volume per service, payment failures per second, cart latency, payment totals per currency, relative errors with an offset, a “service is dead” expression. An AI agent writes and runs the LogQL through the lab’s Grafana MCP server. You ask in plain language, predict the result, read every query and verify the numbers yourself. This variant replaces sections 3.3–5 of the manual exercise. Sections 1–3.2 stay manual.

Goal

  • You ask questions about the lab’s logs in plain language. The agent finds the labels, writes the LogQL and runs it (list_loki_label_names, list_loki_label_values, query_loki_logs).
  • Before each answer you predict the selector, the aggregation and the shape of the result.
  • You review each query: stream selection, line filter and its case, parser, unwrap, grouping, the range vector compared with the question’s time range, and whether the number answers the question at all.
  • You verify in Explore and with exact counts through the Loki API, and work out what the agent assumed when the question was ambiguous.

Environment: workshop cluster, Grafana Explore, a terminal, an MCP-capable AI client.

Prerequisites

  • Sections 1–3.2 of the manual variant done: selectors, line filters, detected fields, range vectors on a Points chart. They train the eye you need for the review.
  • Your lab login (user<N>), its Grafana password and the MCP server token, all from the lab handout.
  • An AI client that supports remote MCP servers over Streamable HTTP with custom headers. The examples use Claude Code.
  • curl, jq, Git Bash or WSL.
  • An empty working directory for the agent. In the course repo the agent reads CLAUDE.md, which describes the lab and its faults. Do not start it there.

Who does what:

Step Who Why
selectors, filters, range vectors on a Points chart (manual 1–3.2) you, in Explore the point is your own eye on the syntax
label discovery, LogQL for every question, the numbers agent list_loki_label_names, list_loki_label_values, query_loki_logs
predictions, review, verdicts you the agent’s query is a claim until you have read it
charts, legends (-offset), exact counts you, in Explore and the Loki API MCP returns numbers, not charts; the legend is a query option in Grafana, not LogQL

đŸ’ȘExerciseđŸ’Ș — steps

Step 1 — a read-only token for the agent

The agent only reads, so its service account is Viewer with no folder permissions. The script reuses mcp-<login>-ro if another exercise created it and issues a new token valid for 8 h.

export GRAFANA_URL=https://grafana.workshop2.indexoutofrange.com
LOGIN=<login>
read -rsp "Grafana password: " P && GRAFANA_AUTH="$LOGIN:$P" && echo
H=(-u "$GRAFANA_AUTH" -H "Content-Type: application/json")
SA_NAME="mcp-$LOGIN-ro"

SA=$(curl -sS "${H[@]}" "$GRAFANA_URL/api/serviceaccounts/search?query=$SA_NAME" \
  | jq -j --arg n "$SA_NAME" '.serviceAccounts[] | select(.name==$n) | .id')
if [ -z "$SA" ]; then
  SA=$(curl -sS "${H[@]}" -X POST "$GRAFANA_URL/api/serviceaccounts" \
    -d "{\"name\":\"$SA_NAME\",\"role\":\"Viewer\"}" | jq -j .id)
fi
export GRAFANA_SERVICE_ACCOUNT_TOKEN=$(curl -sS "${H[@]}" -X POST "$GRAFANA_URL/api/serviceaccounts/$SA/tokens" \
  -d "{\"name\":\"mcp-loki06-$(date +%s)\",\"secondsToLive\":28800}" | jq -j .key)

The password stays in a shell variable, without export. claude started from this terminal inherits exported variables, and an agent with a Bash tool could then call the Grafana API as your Admin account, bypassing MCP.

Check the token’s boundary. A dashboard write must fail:

curl -sS -o /dev/null -w '%{http_code}\n' -H "Authorization: Bearer $GRAFANA_SERVICE_ACCOUNT_TOKEN" -H 'Content-Type: application/json' \
  -X POST "$GRAFANA_URL/api/dashboards/db" -d "{\"dashboard\":{\"title\":\"$SA_NAME probe\",\"panels\":[]}}"

Check that: the response is 403 and $GRAFANA_SERVICE_ACCOUNT_TOKEN starts with glsa_.

Step 2 — connect Grafana MCP in an empty directory

mkdir -p ~/logql-06 && cd ~/logql-06
export GRAFANA_MCP_URL=https://mcp.workshop2.indexoutofrange.com/mcp
read -rsp "MCP server token: " MCP_SERVER_TOKEN && export MCP_SERVER_TOKEN && echo
claude mcp add --transport http grafana "$GRAFANA_MCP_URL" \
  --header "Authorization: Bearer $MCP_SERVER_TOKEN" \
  --header "X-Grafana-Service-Account-Token: $GRAFANA_SERVICE_ACCOUNT_TOKEN"
claude mcp list
  • Claude Code’s default local scope binds the server to this directory. Do not use --scope project: it writes the tokens to .mcp.json.
  • Other clients need the same URL and the same two headers. Without X-Grafana-Service-Account-Token the server falls back to a shared Viewer account, and your token is not the one being tested.

Start claude and run a smoke test:

Via Grafana MCP, Loki datasource uid "loki", selector {namespace="otel"}: list the label names,
then the values of the labels stream, exporter and level. Run no other query.

Check that: the agent called list_loki_label_names and list_loki_label_values with a matcher, the labels include app, stream, exporter and level, and stream has the values otel and stdout.

Answer in your notes before going on:

  1. stream="stdout" is the pod log file, stream="otel" (exporter="OTLP") is the same application sending logs over OTLP. Which of the two can carry level as a label, and why do stdout streams have only detected_level?
  2. A service that writes to both: what happens to a count of its log lines if the selector does not pick a stream?

Step 3 — predict before you ask

Fill the second and third columns now, from the LogQL lesson and sections 1–3.2. You fill the rest in steps 4–6.

# Question (plain language) Your query sketch Predicted result shape (series, labels, magnitude) Agent’s query Verified Verdict
Q1 Which five services in otel produced the most log lines in the last 10 min?          
Q2 How many failed payments per second does payment produce (last 15 min)?          
Q3 Average and p95 request duration of cart, per endpoint (last 15 min)?          
Q4 Total amount of completed payments per currency in the last hour?          
Q5 Share of error lines per service, top 3, now and 5 min earlier?          
Q6 An expression that returns 1 when currency has logged nothing for 5 min?          

Step 4 — the spec

The prompt states what to find and what to report. It contains no LogQL.

Via Grafana MCP, read-only. Loki datasource uid "loki", OTel Demo in namespace "otel".
Before the first query, check label names and values (list_loki_label_names, list_loki_label_values).
For every question report: the LogQL, the query_loki_logs parameters (queryType, start/end,
stepSeconds, limit), the number of series returned, the result, and every assumption you made
(which streams, which filter text and its case, which field). If a question can be read two ways,
say so and pick one explicitly. Count lines with count_over_time, never by the number of lines returned.

Q1. Which five services in namespace otel produced the most log lines in the last 10 minutes?
Q2. How many failed payments per second does the payment service produce, averaged over the last 15 minutes?
Q3. Average and 95th percentile request duration of cart, per endpoint, over the last 15 minutes.
    The duration is in cart's OTLP request logs.
Q4. Total amount of completed payments per currency in the last hour.
Q5. Share of error lines among all lines per service in otel, top 3, and the same value 5 minutes earlier.
Q6. A query that returns 1 when currency has logged nothing for 5 minutes. What does it return now?

Approve tool calls one at a time and read the arguments. Copy each LogQL into the “Agent’s query” column before you read the agent’s answer.

Check that: the agent called query_loki_logs at least six times, and every answer has a query, parameters, a series count, a result and a list of assumptions.

Step 5 — review the agent

Compare every query with the table. A deviation is not wrong by itself if the agent said what it assumed and the assumption is defensible. A silent one is.

Q Correct Typical agent deviation
Q1 topk(5, sum by (app) (count_over_time({namespace="otel"}[10m]))), instant, and the agent says that services writing to both stdout and otel are counted twice, or restricts the stream on purpose no sum by (app): one series per stream, the same app appears twice in the top 5; a range query whose per-step values are added up; both streams counted without a word
Q2 sum(rate({namespace="otel", app="payment", stream="otel"} \|= "Payment request failed" [15m])) or the same selector with level="WARN" \|= "error": no match, the line says Error: and the JSON key is err, and the agent reports “0 failures”; {level="ERROR"}: no match, failed payments are logged as WARN; no stream: each failure is also a multi-line stack trace in stdout, several matching lines per failure
Q3 avg_over_time(
 \| json \| unwrap attributes_ElapsedMilliseconds [15m]) by (attributes_Path) and quantile_over_time(0.95, 
 [15m]) by (attributes_Path) on {namespace="otel", app="cart", stream="otel"} \|= "ElapsedMilliseconds"; three endpoints no by (
): json turns traceid and spanid into labels, so one series per request; avg(quantile_over_time(
)) presented as “the p95 of cart”; a [5m] window reported as “last 15 min”
Q4 sum by (attributes_amount_currencyCode) (sum_over_time({
, app="payment", stream="otel"} \|= "Transaction complete" \| json \| unwrap attributes_amount_units_low [1h])); one series per currency, and the agent notes that units_low drops nanos one total across currencies; \|= "Charge request received": attempts including failed ones, and the field there is attributes_request_amount_units_low; no outer sum by: one series per transaction, json makes attributes_transactionId a label
Q5 topk(3, sum by (app) (count_over_time({namespace="otel"} \|~ "(?i)error" [5m])) / sum by (app) (count_over_time({namespace="otel"}[5m]))) plus the same with offset 5m on both range vectors; the agent says whether it matched case-insensitively \|= "error" only, Error missed; offset on one side of the division only; the claim that a chart of this query shows exactly 3 services: in a range query topk picks the top 3 at each step
Q6 absent_over_time({namespace="otel", app="currency"}[5m]); result now: empty, the service logs “it returns 0” (it returns no series); an error filter inside the selector, so “no errors” reads as “dead”; a negative-only selector, which the server’s Loki guardrail rejects

Answer in your notes:

  1. Q1. Which services were counted twice? For “how chatty is this service”, is counting both streams right or wrong? Write the one sentence you would add to the spec to settle it.
  2. Q2. How many matching stdout lines does one failed payment produce? Is the agent’s per-second number per failure or per line?
  3. Q3. Which labels does json add to a cart line? Why does that turn a missing by into hundreds of series?
  4. Q4. Why is a single total across currencies a wrong answer even when every number in it is correct?
  5. Q5. In a range query over 1 h, how many distinct services can the topk(3, 
) chart show? Why?

Step 6 — verify independently

  1. In Explore → Loki, run the agent’s Q1, Q2 and Q4 queries yourself. Then run each once more with the opposite stream choice and note how the number changes.
  2. Exact counts through the Loki API, with the agent’s token (Viewer is enough for the datasource proxy):
lq() { curl -sS -G -H "Authorization: Bearer $GRAFANA_SERVICE_ACCOUNT_TOKEN" \
  "$GRAFANA_URL/api/datasources/proxy/uid/loki/loki/api/v1/query" --data-urlencode "query=$1" \
  | jq -c '.data.result[] | [.metric, .value[1]]'; }

lq 'sum by (stream) (count_over_time({namespace="otel", app="payment"} |= "Payment request failed" [15m]))'
lq 'sum by (stream) (count_over_time({namespace="otel", app="payment"} |= "error" [15m]))'
lq 'sum by (level) (count_over_time({namespace="otel", app="payment", stream="otel"} [15m]))'
  1. Q5 in Explore: queries A (now) and B (offset 5m), Options → Legend: `` for A, -offset for B. This is the manual part of exercise 4.5.

Check that:

  • your Q2 count for stream="otel", divided by 900 s, matches the agent’s rate to within the drift of the time window,
  • the stdout count for the same failures is several times the otel count (one failure spans several lines), and |= "error" returns nothing,
  • the level split for payment has no ERROR, and failed payments are in WARN,
  • the Q5 chart in Explore shows more than 3 services over the hour, each A series with a matching -offset series.

Fill the “Verified” and “Verdict” columns. A verdict is correct, correct with stated assumption, silently wrong or answers another question.

Step 7 — an ambiguous question on purpose

Start a new agent session (/clear) and type only:

How many errors did payment have in the last hour?

When it answers, ask:

List every assumption you made: which streams, which filter text and case, level label or text,
log lines or failed payments, whether your number is a count or a sample of returned lines,
and the exact time range.
Decision The agent’s choice Alternatives What each gives (Explore)
streams   otel, stdout, both  
“error” means   text error, text Error, level="ERROR", level="WARN", “Payment request failed”  
unit   log lines, failed payments  
method   count_over_time, number of returned lines (capped at 100 by the lab server)  
range   last 1 h, last 5 min of a [5m] window  

Then write a one-sentence question that removes every ambiguity, ask it in the same session and compare the result with your count from step 6 (scaled to 1 h).

Check that: you can name the choice the agent made silently. Typical: “100 errors” (the log-line cap, not a count), “0 errors” from |= "error", or a confident number that adds stdout and otel.

Step 8 — planted bugs: does the agent catch them?

Write your own verdict for each query first. Then paste them to the agent:

Each query below is supposed to answer the question next to it. Without running it, say whether it
answers the question, and predict the result shape. Then run it via query_loki_logs and compare.

a) "Log lines per service in otel, last 10 min":
   count_over_time({namespace="otel"}[10m])
b) "Payment failures per second":
   sum(count_over_time({namespace="otel", app="payment"} |= "failed" [5m]))
c) "p95 of cart request duration per endpoint over the last hour" (instant query):
   quantile_over_time(0.95, {namespace="otel", app="cart", stream="otel"} |= "ElapsedMilliseconds" | json | unwrap attributes_ElapsedMilliseconds [5m]) by (attributes_Path)
Query Bug Agent caught it before running? After running?
a one series per stream, not per service    
b a count per 5 min, not per second; both streams    
c covers the last 5 min, not the hour    

Check that: you found all bugs yourself before the agent answered, and you recorded which ones the agent only noticed after it had seen the result, or not at all.

Step 9 — cleanup

If you continue with another agent variant (TraceQL, PromQL), skip this step: the token is valid for 8 h. Otherwise, in the same terminal as step 1:

cd ~/logql-06 && claude mcp remove grafana
curl -sS "${H[@]}" -X DELETE "$GRAFANA_URL/api/serviceaccounts/$SA" | jq -c .
unset P GRAFANA_AUTH H

Success criteria

  • The prediction table is complete for Q1–Q6, and every row has a verdict backed by your own Explore or API result.
  • You can name, without the agent, the stream that double-counts payment failures, the level failed payments are logged at, and why |= "error" finds none of them.
  • You explained why a missing by after json + unwrap produces one series per request or transaction.
  • The Q5 chart in Explore shows A and B series with `` and -offset legends, and you can say why it shows more than 3 services.
  • For the ambiguous question in step 7 you listed the agent’s silent assumptions and wrote a spec sentence that removes them.
  • You found the bugs in step 8 before the agent did.

⭐Stretch: limits the agent runs into (failure mode)

The lab server runs query_loki_logs with a 100-line cap and a Loki guardrail in enforce mode: a time range above 6 h (range vector included) and a selector with no selective positive matcher are rejected before they reach Loki.

  1. Range cap. Ask: How many lines did cart log in the last 24 hours? Correct: the agent reports the rejection and either asks or splits the range into windows of at most 6 h and says so. Deviation: it silently answers for a shorter range as if it were 24 h.
  2. Negative-only selector. Ask: Count log lines of every app except cart in the last 10 minutes, using only {app!="cart"}. Correct: the agent reports the guardrail rejection and adds a positive matcher such as namespace="otel", saying the scope changed. Deviation: it reports the rejection as “no data”.
  3. Line cap. Ask: Show me all payment log lines from the last 15 minutes and tell me how many there are. Correct: the agent says it got at most 100 lines and counts with count_over_time. Deviation: the number of returned lines given as the total.

results matching ""

    No results matching ""