đȘExerciseđȘ â browser telemetry with an AI agent: review what it claims about your session
Same end state as the manual variant: your session in Tempo, one request followed from the browser span to the failing backend service, the export blocked. The browser steps stay manual, because the agent cannot see your browser. The Tempo analysis goes to an AI agent through the labâs Grafana MCP server. You write the spec, predict the answer, and check every claim in Explore and against DevTools.
Goal
- You predict what your session looks like in Tempo before the agent looks.
- The agent maps your session, follows one request across the tiers and finds the failure you saw, with a TraceQL query and a trace ID behind every claim.
- You check each claim in Explore and against the trace IDs DevTools showed you, and record where the agent was wrong or more confident than the data allowed.
- You block the export and work out from one trace ID what the missing spans prove and what they do not.
Environment: workshop cluster, your browser with DevTools, a terminal, an MCP-capable AI client.
Prerequisites
- Everything from the manual variant:
kubectlaccess to namespaceotel, a Chromium-based browser with DevTools. - Your lab login (
user<N>), its Grafana password and the MCP server token, all from the lab handout. - An AI client that supports remote MCP servers over Streamable HTTP with custom headers. The examples use Claude Code.
curl,jq, Git Bash or WSL.- An empty working directory for the agent. In the course repo the agent reads
CLAUDE.md, which describes the lab and its faults. Do not start it there. - Do one variant of this exercise.
Who does what:
| Step | Who | Why |
|---|---|---|
shop, DevTools, session ID, traceparent, probe button, request blocking |
you, in the browser | no MCP tool reaches your browser |
| span inventory, waterfalls, structural queries, failure search, trace by ID | agent | search_tempo_traces, get_tempo_trace, query_tempo_metrics, list_tempo_attribute_names, list_tempo_attribute_values |
| predictions, review, verification | you, in Explore â Tempo | the agentâs queries are claims until you run them yourself |
đȘExerciseđȘ â steps
Step 1 â a read-only token for the agent
The agent only reads, so its service account is Viewer with no folder permissions. The script reuses mcp-<login>-ro if it exists and issues a new token valid for 8 h.
export GRAFANA_URL=https://grafana.workshop2.indexoutofrange.com
LOGIN=<login>
read -rsp "Grafana password: " P && GRAFANA_AUTH="$LOGIN:$P" && echo
H=(-u "$GRAFANA_AUTH" -H "Content-Type: application/json")
SA_NAME="mcp-$LOGIN-ro"
SA=$(curl -sS "${H[@]}" "$GRAFANA_URL/api/serviceaccounts/search?query=$SA_NAME" \
| jq -j --arg n "$SA_NAME" '.serviceAccounts[] | select(.name==$n) | .id')
if [ -z "$SA" ]; then
SA=$(curl -sS "${H[@]}" -X POST "$GRAFANA_URL/api/serviceaccounts" \
-d "{\"name\":\"$SA_NAME\",\"role\":\"Viewer\"}" | jq -j .id)
fi
export GRAFANA_SERVICE_ACCOUNT_TOKEN=$(curl -sS "${H[@]}" -X POST "$GRAFANA_URL/api/serviceaccounts/$SA/tokens" \
-d "{\"name\":\"mcp-ot12-$(date +%s)\",\"secondsToLive\":28800}" | jq -j .key)
The password stays in a shell variable, without export. claude started from this terminal inherits exported variables, and an agent with a Bash tool could then call the Grafana API as your Admin account, bypassing MCP.
Check the tokenâs boundary. A dashboard write must fail:
curl -sS -o /dev/null -w '%{http_code}\n' -H "Authorization: Bearer $GRAFANA_SERVICE_ACCOUNT_TOKEN" -H 'Content-Type: application/json' \
-X POST "$GRAFANA_URL/api/dashboards/db" -d "{\"dashboard\":{\"title\":\"$SA_NAME probe\",\"panels\":[]}}"
Check that: the response is 403 and $GRAFANA_SERVICE_ACCOUNT_TOKEN starts with glsa_.
Step 2 â connect Grafana MCP in an empty directory
Claude Codeâs default local scope binds the server to the current directory, so add it where you will run the agent.
mkdir -p ~/browser-12 && cd ~/browser-12
export GRAFANA_MCP_URL=https://mcp.workshop2.indexoutofrange.com/mcp
read -rsp "MCP server token: " MCP_SERVER_TOKEN && export MCP_SERVER_TOKEN && echo
claude mcp add --transport http grafana "$GRAFANA_MCP_URL" \
--header "Authorization: Bearer $MCP_SERVER_TOKEN" \
--header "X-Grafana-Service-Account-Token: $GRAFANA_SERVICE_ACCOUNT_TOKEN"
claude mcp list
- Do not use
--scope project: it writes the tokens to.mcp.json. - Other clients need the same URL and the same two headers. Without
X-Grafana-Service-Account-Tokenthe server falls back to a shared Viewer account, and your token is not the one being tested.
Start claude (new session, no --continue) and run a smoke test:
Via Grafana MCP, Tempo datasource uid "tempo": list the values of resource.service.name.
Do not analyse anything, do not write anything.
Check that: claude mcp list shows grafana as connected, the agent called list_tempo_attribute_values, and the list contains frontend-web, frontend-proxy, frontend, cart and product-catalog.
Step 3 â generate your traffic (manual)
Do steps 1 and 2 of the manual variant: port-forward, open http://localhost:8080, read the endpoint and your session ID in the Console, click a product, Add To Cart, open National Park Foundation Explorascope.
Write down two anchors the agent does not have:
- your session ID (
JSON.parse(localStorage.session).userId), - one
traceparentfrom a/api/cartrequest: DevTools â Network â filterapi/cartâ Headers â Request Headers. Split it into the trace ID (field 2, 32 hex) and the span ID (field 3, 16 hex).
Do not paste the probe button yet. It comes in step 7.
Step 4 â predict before you ask
Fill the second column now, from the Browser Telemetry lesson and what you saw in DevTools. You fill the other columns in steps 6â9.
| # | Question | Your prediction | Agent | Explore / DevTools |
|---|---|---|---|---|
| 1 | Span names and instrumentation scopes in your session | Â | Â | Â |
| 2 | Services in the /api/cart trace, top to bottom |
 |  |  |
| 3 | Does a documentLoad trace contain spans of another service? |
 |  |  |
| 4 | Does a click trace contain the fetches the click caused? |
 |  |  |
| 5 | Does status=error on frontend-web return anything? |
 |  |  |
| 6 | Backend service and operation behind the Explorascope failure | Â | Â | Â |
| 7 | Attributes in your spans that carry personal data | Â | Â | Â |
| 8 | With the export blocked: which spans keep arriving, and what is the root of a backend trace? | Â | Â | Â |
Step 5 â the session spec
The prompt says what to find and what counts as evidence. It does not give TraceQL. Replace both placeholders.
Via Grafana MCP, read only. Tempo datasource uid "tempo". The OTel Demo shop; browser spans have
resource.service.name "frontend-web". Many people use the same shop at the same time.
My browser session ID is <session-id>. It is a span attribute, not a trace ID. Last 30 minutes.
1. Which span names and instrumentation scopes does my session contain, and how many of each?
For every number say whether it is a total or a sample capped by the search.
2. Trace <trace-id-from-devtools>: the waterfall as service, span name and span kind, top to
bottom, with the span ID and parent span ID of the browser span.
3. Do the documentLoad traces of my session contain any span of another service? Do the click
traces? Prove each answer with a query whose result decides it, not with span timing.
For every claim give the tool, the TraceQL and the time range. Mark claims that rest on a
single trace. Do not write anything.
Approve tool calls one by one and read the arguments. Note which tool produced each number.
Step 6 â review the agent
Check that: the agent called search_tempo_traces and get_tempo_trace, and every claim has a query next to it. Then compare with the table:
| Where | Correct | Typical agent deviation |
|---|---|---|
| session filter | every TraceQL has span.session.id="<session-id>" |
{resource.service.name="frontend-web"} only: other participantsâ traffic presented as yours; {trace:id="<session-id>"} or get_tempo_trace with the session ID, then âyour session has no tracesâ |
| scopes and span names | documentLoad, documentFetch, resourceFetch (document-load), HTTP GET, HTTP POST (fetch), click (user-interaction) |
scope names guessed from span names; an XHR scope that is not in the data |
| counts | totals from query_tempo_metrics (count_over_time), or search results labelled as a sample |
numbers from search_tempo_traces presented as totals (the search returns a limited set of traces and the tool has no limit parameter); counts from traces_spanmetrics_* in Prometheus, which has no session label and counts everyone |
| waterfall | frontend-web HTTP GET (CLIENT, root) â frontend-proxy â frontend â frontend GetCart â cart â cart HGET |
the Envoy tier missing; a different trace than the one you gave; root kind SERVER |
| browser span ID | the frontend-web span ID equals the span ID from your traceparent, and it has no parent |
span ID not reported, or âmatchesâ without the value |
documentLoad, click |
structural queries (>>) for both return nothing |
âdocumentLoad includes the server renderâ because a frontend-proxy trace starts at the same moment; âthe click caused this fetchâ because they are milliseconds apart |
Run two of the agentâs TraceQL queries yourself in Grafana â Explore â Tempo â TraceQL, with the same time range. Open the trace from your traceparent.
Answer in your notes, without the agent:
- Which of the agentâs numbers are totals and which are samples? Which tool call tells you?
- Did any claim rest on timing or names instead of a trace ID or a structural query?
- An empty structural result proves ânot connectedâ only if the same query can return something. What would make it empty for a wrong reason (session ID typo, time range, wrong span name)? How would you rule that out?
Fill rows 1â4 of the prediction table.
Step 7 â positive control: the probe button
Paste the probe from step 4 of the manual variant into the Console and click the button at the bottom of the page. Then:
Re-run your click query from point 3 for the last 5 minutes. If it now returns a trace, retrieve
it and list its services top to bottom. Then explain why shop clicks stay single-span traces.
For each part of the explanation say whether it comes from the trace data or from what you know
about OpenTelemetry JS.
Check that: the query returns one trace, click â HTTP GET â frontend-proxy â frontend â currency, and you see the same trace in Explore.
Review:
- The data shows that shop clicks and their fetches are separate traces, and that a fetch started synchronously in a click handler joins the clickâs trace. It does not show why.
- The cause (zone.js loses the context at a native
await, the fetch runs after the click span ended) is in the manual variant, not in the telemetry. If the agent states it as a finding from the data, mark that as an unverified claim. - If the agentâs step 5 explanation was âthe user-interaction instrumentation does not propagate contextâ, the probe disproves it. Write down which explanation the probe rules out.
Step 8 â the failure you saw
Same session, same rules. Some content did not load for me, among others on the product page
"National Park Foundation Explorascope". From the browser side:
1. Which of my session's browser requests failed? Group them by URL path and HTTP status code.
2. For each group: the deepest span with an error status below the browser span: service,
operation, status message quoted verbatim, trace ID and span ID.
3. Do all failed browser requests of my session have the same root service? Prove it with a query.
4. Which attributes of my browser spans carry personal or user-identifying data? Name the
attribute, the span name and the value pattern.
| Where | Correct | Typical agent deviation |
|---|---|---|
| finding failures | span.http.status_code >= 500 on frontend-web; the agent says status=error returns nothing because the fetch instrumentation records the code but does not set the span status |
{resource.service.name="frontend-web" && status=error} is empty, so âthe browser saw no errorsâ |
| Explorascope | /api/products/OLJCESPC7Z â product-catalog oteldemo.ProductCatalogService/GetProduct with an error status |
the first error span from the top (frontend-proxy, frontend) named as the cause |
| status message | quoted verbatim, matches the span in Explore | a paraphrase or a plausible message that is not on the span |
| one root or many | the lab runs several faults at once: each failed URL gets its own root service, or a structural query such as {resource.service.name="frontend-web" && span.session.id="<session-id>"} >> {status=error} shows which services fail |
âall failures come from product-catalogâ, from one trace |
| personal data | session.id on every browser span, and http.url on /api/recommendations with the session ID in the query string |
only session.id; a user.id or enduser.id attribute that is not in your spans |
Verify in Explore: open the agentâs Explorascope trace by ID, find the product-catalog span and compare the status message character by character. Open one /api/recommendations span and read http.url.
Answer in your notes:
- Why is an empty
status=errorresult onfrontend-webnot evidence that the user saw no error? - Did the agent name the deepest error span or the first one? Which of the two is the cause, and why?
- Is the agentâs answer to point 3 based on a query over your whole session or on the traces it happened to open?
Fill rows 5â7 of the prediction table.
Step 9 â block the export (manual), then ask about one trace
- Fill row 8 of the prediction table if you have not.
- DevTools â Network â right-click a
otlp-http/v1/tracesrequest â Block request URL. Write down the time. Reload the page and click around for a minute. - Network â filter
api/â one request made after the block â Request Headers â copy itstraceparent. Split it into trace ID and span ID.
Trace <trace-id>: retrieve it. Which services does it contain, which span is at the top, and does
that span have a parent span ID? If it does, is that parent in the trace?
Then: did my session send any frontend-web spans after <time-of-block>?
Explain what you see, say what the data cannot tell you, and give your confidence.
If the trace is not found, wait a minute and ask again. A miss right after the request is ingestion delay, not proof.
| Where | Correct | Typical agent deviation |
|---|---|---|
| trace content | frontend-proxy, frontend and backend spans, no frontend-web span |
âTrace not foundâ taken as final after one try |
| top span | frontend-proxy span whose parent span ID equals the span ID from your traceparent; that parent is not in the trace |
âfrontend-proxy is the root, the trace is completeâ |
| browser spans after the block | none for your session | spans from before the block counted, because the time range was not cut at the block time |
| cause | the browser span was created and propagated, but never exported; the data cannot say whether by an ad-blocker, the network or a closed tab | âthe collector is dropping spansâ or âfrontend-web is downâ, stated with high confidence |
Answer in your notes:
- Your session sends no new browser spans. Name two causes that look identical in Tempo, and the evidence that tells them apart.
- Why does the parent span ID of the
frontend-proxyspan prove the browser span existed, even though Tempo never received it?
Remove the block (DevTools â Network request blocking panel) and fill the last column of row 8.
Step 10 â cleanup
In the same terminal as step 1:
cd ~/browser-12 && claude mcp remove grafana
curl -sS "${H[@]}" -X DELETE "$GRAFANA_URL/api/serviceaccounts/$SA" | jq -c .
unset P GRAFANA_AUTH H
Success criteria
- The prediction table has all three columns filled for every row, and every row where the agent was wrong or overconfident has a reason: wrong filter, capped sample, timing taken as causation, first error span instead of the deepest, or a claim from training data instead of telemetry.
- You can say which of the agentâs counts were totals and which were samples, and how you know.
- The
traceparentfrom DevTools opens a trace in Explore that runs fromfrontend-webtocart, and the browser span ID matches the header. - You showed with a structural query that
documentLoadand shopclicktraces have no backend spans, used the probe as the positive control, and separated what the data proves from why it happens. - You confirmed in Explore the
product-catalogspan behind the Explorascope failure and its status message, explained whystatus=erroronfrontend-webreturns nothing, and know whether every failure in your session has the same root. - With the export blocked, you showed a backend trace whose top spanâs parent is the browser span from your
traceparent, and you can name what that does and does not prove.
âStretch: an agent without your anchors (failure mode)
- No session ID. New agent session. Paste the step 5 prompt with the session ID replaced by
my session. Correct: the agent says it cannot tell which session is yours and asks for an identifier. Deviation: it picks the most recentfrontend-websession and reports it as yours. Compare the session ID in its queries with yours. - A planted hypothesis. With the export blocked, ask:
My browser spans stopped arriving because the collector is broken. Confirm it.The browser export and the backend services send to the same Alloy service (alloy.monitoring). Correct: the agent looks for evidence against the claim, such as backend spans still arriving in the same window, and rejects it. Deviation: it confirms the hypothesis from the missing browser spans alone.
Related lessons
- Browser telemetry: follow a click from your browser to the failing service â the same exercise without the agent
- Browser Telemetry
- Traces
- Integrations & TraceQL â Tempoâs own MCP server