đȘExerciseđȘ â LogQL with an AI agent: judge every query it writes
Same questions and the same end state as the manual variant: log volume per service, payment failures per second,
cartlatency, payment totals per currency, relative errors with an offset, a âservice is deadâ expression. An AI agent writes and runs the LogQL through the labâs Grafana MCP server. You ask in plain language, predict the result, read every query and verify the numbers yourself. This variant replaces sections 3.3â5 of the manual exercise. Sections 1â3.2 stay manual.
Goal
- You ask questions about the labâs logs in plain language. The agent finds the labels, writes the LogQL and runs it (
list_loki_label_names,list_loki_label_values,query_loki_logs). - Before each answer you predict the selector, the aggregation and the shape of the result.
- You review each query: stream selection, line filter and its case, parser,
unwrap, grouping, the range vector compared with the questionâs time range, and whether the number answers the question at all. - You verify in Explore and with exact counts through the Loki API, and work out what the agent assumed when the question was ambiguous.
Environment: workshop cluster, Grafana Explore, a terminal, an MCP-capable AI client.
Prerequisites
- Sections 1â3.2 of the manual variant done: selectors, line filters, detected fields, range vectors on a Points chart. They train the eye you need for the review.
- Your lab login (
user<N>), its Grafana password and the MCP server token, all from the lab handout. - An AI client that supports remote MCP servers over Streamable HTTP with custom headers. The examples use Claude Code.
curl,jq, Git Bash or WSL.- An empty working directory for the agent. In the course repo the agent reads
CLAUDE.md, which describes the lab and its faults. Do not start it there.
Who does what:
| Step | Who | Why |
|---|---|---|
| selectors, filters, range vectors on a Points chart (manual 1â3.2) | you, in Explore | the point is your own eye on the syntax |
| label discovery, LogQL for every question, the numbers | agent | list_loki_label_names, list_loki_label_values, query_loki_logs |
| predictions, review, verdicts | you | the agentâs query is a claim until you have read it |
charts, legends (-offset), exact counts |
you, in Explore and the Loki API | MCP returns numbers, not charts; the legend is a query option in Grafana, not LogQL |
đȘExerciseđȘ â steps
Step 1 â a read-only token for the agent
The agent only reads, so its service account is Viewer with no folder permissions. The script reuses mcp-<login>-ro if another exercise created it and issues a new token valid for 8 h.
export GRAFANA_URL=https://grafana.workshop2.indexoutofrange.com
LOGIN=<login>
read -rsp "Grafana password: " P && GRAFANA_AUTH="$LOGIN:$P" && echo
H=(-u "$GRAFANA_AUTH" -H "Content-Type: application/json")
SA_NAME="mcp-$LOGIN-ro"
SA=$(curl -sS "${H[@]}" "$GRAFANA_URL/api/serviceaccounts/search?query=$SA_NAME" \
| jq -j --arg n "$SA_NAME" '.serviceAccounts[] | select(.name==$n) | .id')
if [ -z "$SA" ]; then
SA=$(curl -sS "${H[@]}" -X POST "$GRAFANA_URL/api/serviceaccounts" \
-d "{\"name\":\"$SA_NAME\",\"role\":\"Viewer\"}" | jq -j .id)
fi
export GRAFANA_SERVICE_ACCOUNT_TOKEN=$(curl -sS "${H[@]}" -X POST "$GRAFANA_URL/api/serviceaccounts/$SA/tokens" \
-d "{\"name\":\"mcp-loki06-$(date +%s)\",\"secondsToLive\":28800}" | jq -j .key)
The password stays in a shell variable, without export. claude started from this terminal inherits exported variables, and an agent with a Bash tool could then call the Grafana API as your Admin account, bypassing MCP.
Check the tokenâs boundary. A dashboard write must fail:
curl -sS -o /dev/null -w '%{http_code}\n' -H "Authorization: Bearer $GRAFANA_SERVICE_ACCOUNT_TOKEN" -H 'Content-Type: application/json' \
-X POST "$GRAFANA_URL/api/dashboards/db" -d "{\"dashboard\":{\"title\":\"$SA_NAME probe\",\"panels\":[]}}"
Check that: the response is 403 and $GRAFANA_SERVICE_ACCOUNT_TOKEN starts with glsa_.
Step 2 â connect Grafana MCP in an empty directory
mkdir -p ~/logql-06 && cd ~/logql-06
export GRAFANA_MCP_URL=https://mcp.workshop2.indexoutofrange.com/mcp
read -rsp "MCP server token: " MCP_SERVER_TOKEN && export MCP_SERVER_TOKEN && echo
claude mcp add --transport http grafana "$GRAFANA_MCP_URL" \
--header "Authorization: Bearer $MCP_SERVER_TOKEN" \
--header "X-Grafana-Service-Account-Token: $GRAFANA_SERVICE_ACCOUNT_TOKEN"
claude mcp list
- Claude Codeâs default
localscope binds the server to this directory. Do not use--scope project: it writes the tokens to.mcp.json. - Other clients need the same URL and the same two headers. Without
X-Grafana-Service-Account-Tokenthe server falls back to a shared Viewer account, and your token is not the one being tested.
Start claude and run a smoke test:
Via Grafana MCP, Loki datasource uid "loki", selector {namespace="otel"}: list the label names,
then the values of the labels stream, exporter and level. Run no other query.
Check that: the agent called list_loki_label_names and list_loki_label_values with a matcher, the labels include app, stream, exporter and level, and stream has the values otel and stdout.
Answer in your notes before going on:
stream="stdout"is the pod log file,stream="otel"(exporter="OTLP") is the same application sending logs over OTLP. Which of the two can carrylevelas a label, and why dostdoutstreams have onlydetected_level?- A service that writes to both: what happens to a count of its log lines if the selector does not pick a stream?
Step 3 â predict before you ask
Fill the second and third columns now, from the LogQL lesson and sections 1â3.2. You fill the rest in steps 4â6.
| # | Question (plain language) | Your query sketch | Predicted result shape (series, labels, magnitude) | Agentâs query | Verified | Verdict |
|---|---|---|---|---|---|---|
| Q1 | Which five services in otel produced the most log lines in the last 10 min? |
 |  |  |  |  |
| Q2 | How many failed payments per second does payment produce (last 15 min)? |
 |  |  |  |  |
| Q3 | Average and p95 request duration of cart, per endpoint (last 15 min)? |
 |  |  |  |  |
| Q4 | Total amount of completed payments per currency in the last hour? | Â | Â | Â | Â | Â |
| Q5 | Share of error lines per service, top 3, now and 5 min earlier? | Â | Â | Â | Â | Â |
| Q6 | An expression that returns 1 when currency has logged nothing for 5 min? |
 |  |  |  |  |
Step 4 â the spec
The prompt states what to find and what to report. It contains no LogQL.
Via Grafana MCP, read-only. Loki datasource uid "loki", OTel Demo in namespace "otel".
Before the first query, check label names and values (list_loki_label_names, list_loki_label_values).
For every question report: the LogQL, the query_loki_logs parameters (queryType, start/end,
stepSeconds, limit), the number of series returned, the result, and every assumption you made
(which streams, which filter text and its case, which field). If a question can be read two ways,
say so and pick one explicitly. Count lines with count_over_time, never by the number of lines returned.
Q1. Which five services in namespace otel produced the most log lines in the last 10 minutes?
Q2. How many failed payments per second does the payment service produce, averaged over the last 15 minutes?
Q3. Average and 95th percentile request duration of cart, per endpoint, over the last 15 minutes.
The duration is in cart's OTLP request logs.
Q4. Total amount of completed payments per currency in the last hour.
Q5. Share of error lines among all lines per service in otel, top 3, and the same value 5 minutes earlier.
Q6. A query that returns 1 when currency has logged nothing for 5 minutes. What does it return now?
Approve tool calls one at a time and read the arguments. Copy each LogQL into the âAgentâs queryâ column before you read the agentâs answer.
Check that: the agent called query_loki_logs at least six times, and every answer has a query, parameters, a series count, a result and a list of assumptions.
Step 5 â review the agent
Compare every query with the table. A deviation is not wrong by itself if the agent said what it assumed and the assumption is defensible. A silent one is.
| Q | Correct | Typical agent deviation |
|---|---|---|
| Q1 | topk(5, sum by (app) (count_over_time({namespace="otel"}[10m]))), instant, and the agent says that services writing to both stdout and otel are counted twice, or restricts the stream on purpose |
no sum by (app): one series per stream, the same app appears twice in the top 5; a range query whose per-step values are added up; both streams counted without a word |
| Q2 | sum(rate({namespace="otel", app="payment", stream="otel"} \|= "Payment request failed" [15m])) or the same selector with level="WARN" |
\|= "error": no match, the line says Error: and the JSON key is err, and the agent reports â0 failuresâ; {level="ERROR"}: no match, failed payments are logged as WARN; no stream: each failure is also a multi-line stack trace in stdout, several matching lines per failure |
| Q3 | avg_over_time(⊠\| json \| unwrap attributes_ElapsedMilliseconds [15m]) by (attributes_Path) and quantile_over_time(0.95, ⊠[15m]) by (attributes_Path) on {namespace="otel", app="cart", stream="otel"} \|= "ElapsedMilliseconds"; three endpoints |
no by (âŠ): json turns traceid and spanid into labels, so one series per request; avg(quantile_over_time(âŠ)) presented as âthe p95 of cartâ; a [5m] window reported as âlast 15 minâ |
| Q4 | sum by (attributes_amount_currencyCode) (sum_over_time({âŠ, app="payment", stream="otel"} \|= "Transaction complete" \| json \| unwrap attributes_amount_units_low [1h])); one series per currency, and the agent notes that units_low drops nanos |
one total across currencies; \|= "Charge request received": attempts including failed ones, and the field there is attributes_request_amount_units_low; no outer sum by: one series per transaction, json makes attributes_transactionId a label |
| Q5 | topk(3, sum by (app) (count_over_time({namespace="otel"} \|~ "(?i)error" [5m])) / sum by (app) (count_over_time({namespace="otel"}[5m]))) plus the same with offset 5m on both range vectors; the agent says whether it matched case-insensitively |
\|= "error" only, Error missed; offset on one side of the division only; the claim that a chart of this query shows exactly 3 services: in a range query topk picks the top 3 at each step |
| Q6 | absent_over_time({namespace="otel", app="currency"}[5m]); result now: empty, the service logs |
âit returns 0â (it returns no series); an error filter inside the selector, so âno errorsâ reads as âdeadâ; a negative-only selector, which the serverâs Loki guardrail rejects |
Answer in your notes:
- Q1. Which services were counted twice? For âhow chatty is this serviceâ, is counting both streams right or wrong? Write the one sentence you would add to the spec to settle it.
- Q2. How many matching
stdoutlines does one failed payment produce? Is the agentâs per-second number per failure or per line? - Q3. Which labels does
jsonadd to acartline? Why does that turn a missingbyinto hundreds of series? - Q4. Why is a single total across currencies a wrong answer even when every number in it is correct?
- Q5. In a range query over 1 h, how many distinct services can the
topk(3, âŠ)chart show? Why?
Step 6 â verify independently
- In Explore â Loki, run the agentâs Q1, Q2 and Q4 queries yourself. Then run each once more with the opposite stream choice and note how the number changes.
- Exact counts through the Loki API, with the agentâs token (Viewer is enough for the datasource proxy):
lq() { curl -sS -G -H "Authorization: Bearer $GRAFANA_SERVICE_ACCOUNT_TOKEN" \
"$GRAFANA_URL/api/datasources/proxy/uid/loki/loki/api/v1/query" --data-urlencode "query=$1" \
| jq -c '.data.result[] | [.metric, .value[1]]'; }
lq 'sum by (stream) (count_over_time({namespace="otel", app="payment"} |= "Payment request failed" [15m]))'
lq 'sum by (stream) (count_over_time({namespace="otel", app="payment"} |= "error" [15m]))'
lq 'sum by (level) (count_over_time({namespace="otel", app="payment", stream="otel"} [15m]))'
- Q5 in Explore: queries A (now) and B (
offset 5m), Options â Legend: `` for A,-offsetfor B. This is the manual part of exercise 4.5.
Check that:
- your Q2 count for
stream="otel", divided by 900 s, matches the agentâs rate to within the drift of the time window, - the
stdoutcount for the same failures is several times theotelcount (one failure spans several lines), and|= "error"returns nothing, - the
levelsplit forpaymenthas noERROR, and failed payments are inWARN, - the Q5 chart in Explore shows more than 3 services over the hour, each A series with a matching
-offsetseries.
Fill the âVerifiedâ and âVerdictâ columns. A verdict is correct, correct with stated assumption, silently wrong or answers another question.
Step 7 â an ambiguous question on purpose
Start a new agent session (/clear) and type only:
How many errors did payment have in the last hour?
When it answers, ask:
List every assumption you made: which streams, which filter text and case, level label or text,
log lines or failed payments, whether your number is a count or a sample of returned lines,
and the exact time range.
| Decision | The agentâs choice | Alternatives | What each gives (Explore) |
|---|---|---|---|
| streams | Â | otel, stdout, both |
 |
| âerrorâ means | Â | text error, text Error, level="ERROR", level="WARN", âPayment request failedâ |
 |
| unit | Â | log lines, failed payments | Â |
| method | Â | count_over_time, number of returned lines (capped at 100 by the lab server) |
 |
| range | Â | last 1 h, last 5 min of a [5m] window |
 |
Then write a one-sentence question that removes every ambiguity, ask it in the same session and compare the result with your count from step 6 (scaled to 1 h).
Check that: you can name the choice the agent made silently. Typical: â100 errorsâ (the log-line cap, not a count), â0 errorsâ from |= "error", or a confident number that adds stdout and otel.
Step 8 â planted bugs: does the agent catch them?
Write your own verdict for each query first. Then paste them to the agent:
Each query below is supposed to answer the question next to it. Without running it, say whether it
answers the question, and predict the result shape. Then run it via query_loki_logs and compare.
a) "Log lines per service in otel, last 10 min":
count_over_time({namespace="otel"}[10m])
b) "Payment failures per second":
sum(count_over_time({namespace="otel", app="payment"} |= "failed" [5m]))
c) "p95 of cart request duration per endpoint over the last hour" (instant query):
quantile_over_time(0.95, {namespace="otel", app="cart", stream="otel"} |= "ElapsedMilliseconds" | json | unwrap attributes_ElapsedMilliseconds [5m]) by (attributes_Path)
| Query | Bug | Agent caught it before running? | After running? |
|---|---|---|---|
| a | one series per stream, not per service | Â | Â |
| b | a count per 5 min, not per second; both streams | Â | Â |
| c | covers the last 5 min, not the hour | Â | Â |
Check that: you found all bugs yourself before the agent answered, and you recorded which ones the agent only noticed after it had seen the result, or not at all.
Step 9 â cleanup
If you continue with another agent variant (TraceQL, PromQL), skip this step: the token is valid for 8 h. Otherwise, in the same terminal as step 1:
cd ~/logql-06 && claude mcp remove grafana
curl -sS "${H[@]}" -X DELETE "$GRAFANA_URL/api/serviceaccounts/$SA" | jq -c .
unset P GRAFANA_AUTH H
Success criteria
- The prediction table is complete for Q1âQ6, and every row has a verdict backed by your own Explore or API result.
- You can name, without the agent, the stream that double-counts payment failures, the level failed payments are logged at, and why
|= "error"finds none of them. - You explained why a missing
byafterjson+unwrapproduces one series per request or transaction. - The Q5 chart in Explore shows A and B series with `` and
-offsetlegends, and you can say why it shows more than 3 services. - For the ambiguous question in step 7 you listed the agentâs silent assumptions and wrote a spec sentence that removes them.
- You found the bugs in step 8 before the agent did.
âStretch: limits the agent runs into (failure mode)
The lab server runs query_loki_logs with a 100-line cap and a Loki guardrail in enforce mode: a time range above 6 h (range vector included) and a selector with no selective positive matcher are rejected before they reach Loki.
- Range cap. Ask:
How many lines did cart log in the last 24 hours?Correct: the agent reports the rejection and either asks or splits the range into windows of at most 6 h and says so. Deviation: it silently answers for a shorter range as if it were 24 h. - Negative-only selector. Ask:
Count log lines of every app except cart in the last 10 minutes, using only {app!="cart"}.Correct: the agent reports the guardrail rejection and adds a positive matcher such asnamespace="otel", saying the scope changed. Deviation: it reports the rejection as âno dataâ. - Line cap. Ask:
Show me all payment log lines from the last 15 minutes and tell me how many there are.Correct: the agent says it got at most 100 lines and counts withcount_over_time. Deviation: the number of returned lines given as the total.
Related lessons
- đȘExerciseđȘ â LogQL â the same exercise without the agent
- LogQL
- Labels
- TraceQL with an AI agent â the same review discipline for traces
- PromQL with an AI agent â the same review discipline for metrics