💪Exercise💪 — TraceQL with an AI agent: judge every query it writes
Same questions and the same end state as section 2 of the manual variant: HTTP 5xx anywhere in a trace, database spans, structural queries, pipelines,
select(). An AI agent writes and runs the TraceQL through the lab’s Grafana MCP server. You ask in plain language, predict the result, read every query and verify it in Explore. Section 1 of the manual exercise (syntax, types, scopes, structural operators, the pipeline question) stays manual: it is what you read the agent’s queries with.
Goal
- You ask questions about the lab’s traces in plain language. The agent finds the attributes, writes TraceQL and runs it (
list_tempo_attribute_names,list_tempo_attribute_values,search_tempo_traces,query_tempo_metrics,get_tempo_trace). - Before each answer you predict which spans match and in which services.
- You review each query: scope (
span.,resource.,.), attribute name and type, span vs trace duration, operand order of structural operators, what a pipeline aggregates over, and whether the result is a sample or a count. - You verify in Explore and through the Tempo API, and work out what the agent assumed when the question was ambiguous.
Environment: workshop cluster, Grafana Explore, a terminal, an MCP-capable AI client.
Prerequisites
- Section 1 of the manual variant done: basic syntax, data types, scopes, structural operators, the
frontend-proxypipeline query. - Your lab login (
user<N>), its Grafana password and the MCP server token, all from the lab handout. - An AI client that supports remote MCP servers over Streamable HTTP with custom headers. The examples use Claude Code.
curl,jq, Git Bash or WSL.- An empty working directory for the agent. In the course repo the agent reads
CLAUDE.md, which describes the lab and its faults. Do not start it there. - For the ⭐Stretch only:
kubectlaccess that allows a port-forward in namespacemonitoring.
Who does what:
| Step | Who | Why |
|---|---|---|
| syntax, types, scopes, structural operators (manual section 1) | you, in Explore | the point is your own reading of TraceQL |
| attribute discovery, TraceQL for every question, trace lookups | agent | list_tempo_attribute_names, list_tempo_attribute_values, search_tempo_traces, query_tempo_metrics, get_tempo_trace |
| predictions, review, verdicts | you | the agent’s query is a claim until you have read it |
| waterfalls, span details, exact counts | you, in Explore and the Tempo API | you check the trace tree with your own eyes |
💪Exercise💪 — steps
Step 1 — a read-only token for the agent
The agent only reads, so its service account is Viewer with no folder permissions. The script reuses mcp-<login>-ro if another exercise created it and issues a new token valid for 8 h.
export GRAFANA_URL=https://grafana.workshop2.indexoutofrange.com
LOGIN=<login>
read -rsp "Grafana password: " P && GRAFANA_AUTH="$LOGIN:$P" && echo
H=(-u "$GRAFANA_AUTH" -H "Content-Type: application/json")
SA_NAME="mcp-$LOGIN-ro"
SA=$(curl -sS "${H[@]}" "$GRAFANA_URL/api/serviceaccounts/search?query=$SA_NAME" \
| jq -j --arg n "$SA_NAME" '.serviceAccounts[] | select(.name==$n) | .id')
if [ -z "$SA" ]; then
SA=$(curl -sS "${H[@]}" -X POST "$GRAFANA_URL/api/serviceaccounts" \
-d "{\"name\":\"$SA_NAME\",\"role\":\"Viewer\"}" | jq -j .id)
fi
export GRAFANA_SERVICE_ACCOUNT_TOKEN=$(curl -sS "${H[@]}" -X POST "$GRAFANA_URL/api/serviceaccounts/$SA/tokens" \
-d "{\"name\":\"mcp-tempo02-$(date +%s)\",\"secondsToLive\":28800}" | jq -j .key)
The password stays in a shell variable, without export. claude started from this terminal inherits exported variables, and an agent with a Bash tool could then call the Grafana API as your Admin account, bypassing MCP.
Check the token’s boundary. A dashboard write must fail:
curl -sS -o /dev/null -w '%{http_code}\n' -H "Authorization: Bearer $GRAFANA_SERVICE_ACCOUNT_TOKEN" -H 'Content-Type: application/json' \
-X POST "$GRAFANA_URL/api/dashboards/db" -d "{\"dashboard\":{\"title\":\"$SA_NAME probe\",\"panels\":[]}}"
Check that: the response is 403 and $GRAFANA_SERVICE_ACCOUNT_TOKEN starts with glsa_.
Step 2 — connect Grafana MCP in an empty directory
mkdir -p ~/traceql-02 && cd ~/traceql-02
export GRAFANA_MCP_URL=https://mcp.workshop2.indexoutofrange.com/mcp
read -rsp "MCP server token: " MCP_SERVER_TOKEN && export MCP_SERVER_TOKEN && echo
claude mcp add --transport http grafana "$GRAFANA_MCP_URL" \
--header "Authorization: Bearer $MCP_SERVER_TOKEN" \
--header "X-Grafana-Service-Account-Token: $GRAFANA_SERVICE_ACCOUNT_TOKEN"
claude mcp list
- Claude Code’s default
localscope binds the server to this directory. Do not use--scope project: it writes the tokens to.mcp.json. - Other clients need the same URL and the same two headers. Without
X-Grafana-Service-Account-Tokenthe server falls back to a shared Viewer account, and your token is not the one being tested.
Start claude and run a smoke test:
Via Grafana MCP, Tempo datasource uid "tempo": list the span-scope attribute names that start with
"http." or "db.", and the values of resource.service.name. Run no TraceQL search.
Check that: the agent called list_tempo_attribute_names with scope = span and list_tempo_attribute_values. The span attributes include both http.status_code and http.response.status_code, plus db.system and db.statement. The services include frontend, frontend-proxy, checkout, payment, cart and product-catalog.
Answer in your notes: two names for the HTTP status code coexist in this lab. What does that mean for a query that uses only one of them?
Step 3 — predict before you ask
Fill the second and third columns now, from the Integrations & TraceQL lesson and section 1. You fill the rest in steps 4–6.
| # | Question (plain language) | Your query sketch | Predicted: matching services / spans | Agent’s query | Verified | Verdict |
|---|---|---|---|---|---|---|
| T1 | Traces that contain an HTTP request that ended with a 5xx status, anywhere | |||||
| T2 | Spans that carry the text of a PostgreSQL query | |||||
| T3 | Traces with more than one call to Redis | |||||
| T4 | payment work longer than 0.1 ms that ended with an error |
|||||
| T5 | Traces where frontend calls payment, directly or not, and that call fails |
|||||
| T6 | checkout calls payment and email as siblings under one parent |
|||||
| T7 | Services whose average span duration exceeds 20 ms | |||||
| T8 | Errors in product-catalog, showing service, status and duration |
Step 4 — the spec
The prompt states what to find and what to report. It contains no TraceQL.
Via Grafana MCP, read-only. Tempo datasource uid "tempo", OTel Demo, last 1 hour.
Before writing a query, check attribute names, scopes and value types (list_tempo_attribute_names
with a scope, list_tempo_attribute_values). For every question report: the TraceQL, the tool and its
parameters, how many traces or series came back, two example trace IDs, and every assumption you made
(scope, attribute name and type, span or trace duration). If a question can be read two ways, say so
and pick one explicitly. search_tempo_traces returns a sample, not a total: say so whenever you give
a number from it, and use query_tempo_metrics when the question needs a count or an average.
T1. Traces that contain an HTTP request that ended with a 5xx status code, anywhere in the trace.
T2. Spans that contain the text of a PostgreSQL query.
T3. Traces that contain more than one call to Redis.
T4. Work in the payment service that lasted longer than 0.1 ms and ended with an error.
T5. Traces where frontend calls payment, directly or indirectly, and that payment call ends with an error.
T6. Traces where checkout, within one operation, calls both payment and email as siblings.
T7. Services whose average span duration exceeds 20 ms.
T8. Errors in product-catalog; show the service name, status and span duration in the result.
Approve tool calls one at a time and read the arguments. Copy each TraceQL into the “Agent’s query” column before you read the agent’s answer.
Check that: every answer has a query, a tool name, a number with “sample” or “count” next to it, trace IDs and a list of assumptions.
Step 5 — review the agent
A deviation is not wrong by itself if the agent said what it assumed and the assumption is defensible. A silent one is.
| T | Correct | Typical agent deviation |
|---|---|---|
| T1 | { span.http.status_code >= 500 \|\| span.http.status_code =~ "5.." \|\| span.http.response.status_code >= 500 }, and the agent says why: frontend stores http.status_code as an int, frontend-proxy (Envoy) as a string |
{ span.http.status_code >= 500 } only: frontend-proxy 5xx missed, because a numeric comparison never matches a string; http.response.status_code only: no 5xx in this lab; status = error: span status, not HTTP status |
| T2 | { span.db.system = "postgresql" && span.db.statement != nil }; spans from product-reviews |
{ span.db.statement != nil }: also returns cart’s Redis commands; span.db.query.text (newer semantic conventions): not present in this lab, empty result reported as “no Postgres queries” |
| T3 | { span.db.system = "redis" } \| count() > 1; the Redis spans are client spans in cart |
{ resource.service.name = "valkey-cart" }: the Redis server sends no spans; “20 traces have more than one Redis call”: 20 is the search limit, not a total |
| T4 | { resource.service.name = "payment" && duration > 100us && status = error }, and the agent says it read “lasted” as span duration (or uses trace:duration and says so); the error spans are the server spans grpc.oteldemo.PaymentService/Charge |
{ resource.service.name = "payment" } && { status = error }: the error can sit in another service of the same trace; duration > 100ms (unit slip); the claim that the internal charge span carries the error, it stays unset |
| T5 | { resource.service.name = "frontend" } >> { resource.service.name = "payment" && status = error } |
> instead of >>: empty, checkout sits between frontend and payment; operands swapped: the returned spans are frontend spans, or nothing; && between the spansets: no relationship at all |
| T6 | { resource.service.name = "checkout" && name = "oteldemo.PaymentService/Charge" } ~ { resource.service.name = "checkout" && span.server.address = "email" } |
name = "HTTP POST" without the address: also matches the calls to shipping; { resource.service.name = "email" } on the right: empty, the email server span is a child of checkout’s client span, not its sibling |
| T7 | two readings, stated: per service over the hour, { } \| avg_over_time(duration) by (resource.service.name) via query_tempo_metrics, values in seconds, filtered at 0.02; or per trace, { } \| by(resource.service.name) \| avg(duration) > 20ms via search |
the per-trace result presented as service averages; seconds read as milliseconds; long-lived spans (streams, consumers) taken as “this service is slow” without checking kind; { } \| avg(duration) > 20ms \| by(…): averages the whole trace first, then groups |
| T8 | { resource.service.name = "product-catalog" && status = error } \| select(resource.service.name, status, duration) |
the claim that select() filters (it only adds attributes to the result); when there are no errors in the window, a relaxed query, such as dropping status = error, reported as the answer |
Answer in your notes:
- T1. Which service did the int-only query miss, and why does TraceQL not convert
"503"to503? - T3 and T7. Which numbers in the agent’s answers came from
search_tempo_traces(a sample of at most 20 traces, 3 spans per spanset by default) and which fromquery_tempo_metrics(a computed count or average)? Mark each one. - T5. Draw the path
frontend→ … →paymentfrom one trace the agent returned. Which span sits between them? - T7. Which of the two readings answers “which services are slow”? Which
kindwould you filter on, and why?
Step 6 — verify independently
- In Explore → Tempo → TraceQL, run the agent’s T1, T5 and T6 queries yourself. Open one trace of each and find the matching spans in the waterfall.
- Counts and averages through the Tempo API, with the agent’s token (Viewer is enough for the datasource proxy):
tq() { curl -sS -G -H "Authorization: Bearer $GRAFANA_SERVICE_ACCOUNT_TOKEN" \
"$GRAFANA_URL/api/datasources/proxy/uid/tempo/api/metrics/query" --data-urlencode "q=$1" \
--data-urlencode "start=$(( $(date +%s) - 3600 ))" --data-urlencode "end=$(date +%s)" \
| jq -c '.series[]? | [([.labels[] | "\(.key)=\(.value | .[keys[0]])"] | join(",")), .value]'; }
tq '{ span.http.status_code >= 500 } | count_over_time() by (resource.service.name)'
tq '{ span.http.status_code =~ "5.." } | count_over_time() by (resource.service.name)'
tq '{ resource.service.name = "checkout" && name = "HTTP POST" } | count_over_time() by (span.server.address)'
tq '{ } | avg_over_time(duration) by (resource.service.name, kind)'
Check that:
- the int query and the string query return different services, and the agent’s T1 covers both,
HTTP POSTfromcheckoutgoes toemailandshipping, so T6 withoutspan.server.addressis too broad,- in the per-
kindaverages, at least one service has a client or consumer average of minutes (a long-lived stream), far above its server spans. An average over all kinds flags it as slow.
Fill the “Verified” and “Verdict” columns. A verdict is correct, correct with stated assumption, silently wrong or answers another question.
Step 7 — an ambiguous question on purpose
Start a new agent session (/clear) and type only:
Which services are slow?
When it answers, ask:
List every assumption you made: span or trace duration, which span kinds, average or percentile,
the threshold, the time range, and whether each number is a count from query_tempo_metrics or a
sample from search_tempo_traces.
| Decision | The agent’s choice | Alternatives | What each gives (Explore or tq) |
|---|---|---|---|
| duration | span duration, trace:duration |
||
| span kinds | all, kind = server |
||
| statistic | average, p90 (quantile_over_time(duration, .9)) |
||
| threshold | none, 20 ms, the service’s own baseline | ||
| source | query_tempo_metrics, a search sample |
Then write a one-sentence question that removes every ambiguity, ask it in the same session and compare the answer with your tq result.
Check that: you can name the choice the agent made silently. Typical: a service flagged as slow because of long-lived stream or consumer spans, or “slow” decided from the durations of 20 searched traces.
Step 8 — planted bugs: does the agent catch them?
Write your own verdict for each query first. Then paste them to the agent:
Each query below is supposed to answer the question next to it. Without running it, say whether it
answers the question and predict what comes back. Then run it and compare.
a) "Traces where checkout calls payment and the payment call fails":
{ resource.service.name = "checkout" } > { resource.service.name = "payment" } && { status = error }
b) "Spans in cart slower than 50 ms":
{ .service.name = "cart" && duration > 50 }
c) "HTTP 5xx anywhere, for an alert":
{ .http.status_code >= 500 }
| Query | Bug | Agent caught it before running? | After running? |
|---|---|---|---|
| a | && { status = error } is a separate spanset: the error can be in any span of the trace |
||
| b | .service.name searches span and resource attributes (slower, can match a span attribute); duration > 50 has no unit: compare its result with duration > 50ms |
||
| c | misses 5xx stored as strings (frontend-proxy) and under http.response.status_code |
Check that: you found all bugs yourself before the agent answered, and you recorded which ones the agent only noticed after it had seen the result, or not at all.
Step 9 — cleanup
If you continue with another agent variant (LogQL, PromQL), skip this step: the token is valid for 8 h. Otherwise, in the same terminal as step 1:
cd ~/traceql-02 && claude mcp remove grafana
curl -sS "${H[@]}" -X DELETE "$GRAFANA_URL/api/serviceaccounts/$SA" | jq -c .
unset P GRAFANA_AUTH H
Success criteria
- The prediction table is complete for T1–T8, and every row has a verdict backed by your own Explore or API result.
- You can explain, without the agent, why one HTTP 5xx query is not enough in this lab: two attribute names and two value types.
- For every number the agent gave you know whether it was a search sample or a computed count, and you showed one place where the difference changes the answer.
- You drew the
frontend→paymentpath from a real trace and know why>returns nothing there. - For the ambiguous question in step 7 you listed the agent’s silent assumptions and wrote a spec sentence that removes them.
- You found the bugs in step 8 before the agent did.
⭐Stretch: Grafana MCP vs Tempo’s own MCP server (failure mode)
Tempo’s query-frontend has its own MCP server (/api/mcp, enabled in the lab, ClusterIP only). Its tools are traceql-search, traceql-metrics-instant, traceql-metrics-range, get-trace, get-attribute-names, get-attribute-values and docs-traceql.
kubectl port-forward -n monitoring svc/tempo-query-frontend 3200:3200
# second terminal, same directory
claude mcp add --transport http tempo http://localhost:3200/api/mcp
- Ask T1 and T7 again and tell the agent to use only the
temposerver. Compare queries and results with step 5. - Answer in your notes: who authorises a call through the native server? What does Grafana RBAC on your Viewer token do there? What happens to the access boundary if
/api/mcpis exposed through an ingress? - Remove it afterwards:
claude mcp remove tempo, then stop the port-forward.
Correct: the same TraceQL gives the same traces, and you can state that in this lab the native server checks no credentials (Tempo runs without multi-tenancy or auth in front of the query-frontend), so anyone who reaches the port reads every trace. Deviation: the agent mixes tools from both servers in one answer without saying which one each result came from.
Related lessons
- 💪Exercise💪 — TraceQL — the same exercise without the agent
- Integrations & TraceQL — TraceQL and Tempo’s MCP server
- Overview
- LogQL with an AI agent — the same review discipline for logs
- PromQL with an AI agent — the same review discipline for metrics