AI Telemetry
- 🤖 HTTP span vs GenAI span
- 📐 GenAI semantic conventions
- 📊 GenAI metrics
- 🧪 Prompt and response content
- 🕸️ Agents and tool calls
- 🧰 Instrumentation options
- 🚨 What to alert on
- Related lessons
An LLM call is an HTTP request that is slow, expensive, non-deterministic and billed per token. Plain HTTP instrumentation sees
POST /v1/chat/completions 200 4.2s— nothing about why it cost what it did or why the answer was wrong.The OpenTelemetry GenAI semantic conventions add what an operator needs: model, tokens, finish reason, tool calls and, opt-in, the prompt itself.
🤖 HTTP span vs GenAI span
| Question | HTTP instrumentation / eBPF | GenAI instrumentation |
|---|---|---|
| Which model answered? | ❌ Hidden in the body | gen_ai.request.model, gen_ai.response.model |
| How many tokens, what cost? | ❌ | gen_ai.usage.input_tokens, gen_ai.usage.output_tokens |
| Was the answer cut off? | ❌ | gen_ai.response.finish_reasons = ["length"] |
| Which tool did the agent call? | ❌ | execute_tool span, gen_ai.tool.name |
| What was asked and answered? | ❌ | Opt-in message content |
| How long did it take? | ✅ | ✅ |
The workshop cluster shows both sides of one call:
llm— a mock model server with no SDK, instrumented only by Beyla (see Instrumenting Applications). You get a server span with method, route, status and duration.product-reviews— callsllmthrough the OpenAI client with GenAI instrumentation andOTEL_INSTRUMENTATION_GENAI_CAPTURE_MESSAGE_CONTENT=truein otel-demo.yaml. Its client span carries model, tokens and messages.
📐 GenAI semantic conventions
Status: Development. Names still change between convention versions:
gen_ai.system→gen_ai.provider.name,- per-message events (
gen_ai.user.message,gen_ai.choice) →gen_ai.input.messages/gen_ai.output.messages.
Instrumentations that support the switch select the latest names with OTEL_SEMCONV_STABILITY_OPT_IN=gen_ai_latest_experimental. Pin instrumentation versions, and check which names your version emits before building dashboards on them.
Operations and span names
Span name = {gen_ai.operation.name} {target}.
gen_ai.operation.name |
Span name example | Kind |
|---|---|---|
chat |
chat gpt-4.1-mini |
CLIENT |
text_completion, generate_content |
generate_content gemini-2.5-flash |
CLIENT |
embeddings |
embeddings text-embedding-3-small |
CLIENT |
execute_tool |
execute_tool search_orders |
INTERNAL |
invoke_agent |
invoke_agent support-bot |
CLIENT or INTERNAL |
create_agent |
create_agent support-bot |
CLIENT |
Key attributes
| Attribute | Example | Why it matters |
|---|---|---|
gen_ai.provider.name |
openai, anthropic, aws.bedrock |
Which vendor bills you |
gen_ai.request.model |
gpt-4.1-mini |
What you asked for |
gen_ai.response.model |
gpt-4.1-mini-2025-04-14 |
What actually answered — aliases move to new snapshots |
gen_ai.usage.input_tokens / output_tokens |
412 / 96 |
Cost and context-window pressure |
gen_ai.response.finish_reasons |
["stop"], ["length"], ["tool_calls"] |
Truncation, tool loops |
gen_ai.request.temperature, max_tokens |
0.2, 512 |
Explains non-determinism and truncation |
gen_ai.conversation.id |
c-7f3a… |
Groups all turns of one conversation |
gen_ai.agent.name, gen_ai.tool.name |
support-bot, search_orders |
Agent and tool attribution |
error.type |
429, timeout |
Rate limits and provider outages |
An LLM client span (illustrative values):
chat astronomy-llm kind=CLIENT 1.8s
gen_ai.operation.name = chat
gen_ai.provider.name = openai
gen_ai.request.model = astronomy-llm
gen_ai.response.model = astronomy-llm
gen_ai.usage.input_tokens = 412
gen_ai.usage.output_tokens = 96
gen_ai.response.finish_reasons = ["stop"]
server.address = llm
server.port = 8000
gen_ai.provider.name = openai describes the API, not the vendor behind it: any OpenAI-compatible server (the workshop mock, vLLM, a gateway) looks the same. server.address tells you where the call really went.
📊 GenAI metrics
| Metric (OTel) | Prometheus name | Type | Use |
|---|---|---|---|
gen_ai.client.token.usage |
gen_ai_client_token_usage_{bucket,sum,count} |
Histogram, gen_ai.token.type = input / output |
Token burn, cost |
gen_ai.client.operation.duration |
gen_ai_client_operation_duration_seconds_* |
Histogram | Latency per model |
gen_ai.server.time_to_first_token |
gen_ai_server_time_to_first_token_seconds_* |
Histogram | Perceived latency of streaming |
gen_ai.server.time_per_output_token |
gen_ai_server_time_per_output_token_seconds_* |
Histogram | Generation speed |
Client metrics come from the instrumented caller; server metrics from the model server. Self-hosted servers (vLLM, TGI) also expose their own Prometheus metrics: queue depth, KV-cache usage, batch size.
# tokens per second, by model and direction
sum by (gen_ai_request_model, gen_ai_token_type) (
rate(gen_ai_client_token_usage_sum[5m])
)
# p95 LLM call latency by model
histogram_quantile(0.95,
sum by (le, gen_ai_request_model) (
rate(gen_ai_client_operation_duration_seconds_bucket[5m])
)
)
# daily input-token cost, price $0.40 per 1M tokens
sum by (gen_ai_request_model) (
increase(gen_ai_client_token_usage_sum{gen_ai_token_type="input"}[1d])
) / 1e6 * 0.40
Prices are not telemetry. Keep them in one recording rule, or as a static price series joined with group_left, not scattered across dashboards.
🧪 Prompt and response content
Content capture is off by default, for good reasons:
- PII and secrets. Users paste emails, customer data and API keys into prompts.
- Size. A RAG prompt of 50k tokens is hundreds of KB — per span.
- Cost. Trace storage is priced for small spans.
Where content lands when enabled depends on the convention version: older instrumentations emit it as log events (gen_ai.user.message, gen_ai.choice), newer ones as gen_ai.input.messages, gen_ai.output.messages and gen_ai.system_instructions on the span or on a gen_ai.client.inference.operation.details event.
Production pattern:
| Control | How |
|---|---|
| Content to logs, not span attributes | Loki with short retention and tighter access than application logs |
| Redact before storage | Collector redaction or transform processor |
| Capture selectively | Only on errors, or for a sampled fraction of calls |
| Bound attribute size | OTEL_ATTRIBUTE_VALUE_LENGTH_LIMIT in the SDK |
🕸️ Agents and tool calls
An agent is a loop: model call → tool call → model call, until finish_reasons = ["stop"]. Its trace:
invoke_agent support-bot 8.4s
├── chat gpt-4.1 1.9s finish_reasons=["tool_calls"] in=1.2k
├── execute_tool search_orders 0.3s
│ └── GET /orders?customer=… (HTTP auto-instr.) 0.2s
├── chat gpt-4.1 2.2s finish_reasons=["tool_calls"] in=2.9k
├── execute_tool refund_order 0.4s
│ └── POST /refunds (HTTP auto-instr.) 0.3s
└── chat gpt-4.1 1.6s finish_reasons=["stop"] in=3.4k
What it answers:
- How many iterations? A runaway loop shows up as dozens of
chat→execute_toolpairs. - Which step is slow? Model time vs tool time vs the backend call inside the tool.
- Why does cost grow? Every
chatre-sends the whole history, so input tokens rise with each iteration. - Which conversation?
gen_ai.conversation.idstitches turns across traces, assession.iddoes in the browser.
MCP. Semantic conventions for Model Context Protocol (mcp.method.name, mcp.session.id, tool name) are in development, including context propagation through the request’s _meta field. With it, a tool server’s spans join the agent’s trace instead of starting a new one.
🧰 Instrumentation options
| Option | Conventions | Notes |
|---|---|---|
OTel contrib instrumentations (opentelemetry-instrumentation-openai-v2, Google GenAI, Vertex AI, …) |
gen_ai.* |
Vendor-neutral reference implementation |
| Frameworks with built-in OTel (Microsoft.Extensions.AI, Semantic Kernel, Spring AI, Vercel AI SDK) | gen_ai.*, sometimes plus their own |
Enable the framework’s telemetry switch; no extra library |
| OpenLLMetry, OpenLIT | gen_ai.* (mostly) |
Broad provider and vector-DB coverage, OTLP to any backend |
| OpenInference (Arize Phoenix) | Own: llm.*, openinference.span.kind |
OTLP transport, but gen_ai.* dashboards will not match — pick one convention or map in the collector |
| eBPF (Beyla / OBI) | HTTP only | Latency and errors of the call, no tokens or model |
LLM-focused platforms (Langfuse, Phoenix, Grafana AI Observability) accept OTLP. A collector can fan the same spans out to Tempo for operations and to an evaluation tool for quality work.
AI coding tools as a telemetry source
Developer tools emit GenAI-style telemetry too. Claude Code exports OTel metrics (claude_code.token.usage, claude_code.cost.usage, claude_code.session.count) and log events (claude_code.api_request, claude_code.tool_result):
export CLAUDE_CODE_ENABLE_TELEMETRY=1
export OTEL_METRICS_EXPORTER=otlp
export OTEL_LOGS_EXPORTER=otlp
export OTEL_EXPORTER_OTLP_PROTOCOL=http/protobuf
export OTEL_EXPORTER_OTLP_ENDPOINT=https://otlp.workshop2.indexoutofrange.com
The same privacy rule applies: prompt text is excluded unless OTEL_LOG_USER_PROMPTS=1.
🚨 What to alert on
| Symptom | Signal | Typical cause |
|---|---|---|
| Cost spike | rate(gen_ai_client_token_usage_sum[…]) jumps |
Agent loop, prompt regression, traffic |
Rising 429 errors |
error.type on chat spans |
Provider rate limit, missing backoff |
More finish_reasons = ["length"] |
Span attribute ratio | max_tokens too low, answers truncated |
| Latency regression, no deploy | gen_ai.response.model changed |
Provider moved an alias to a new snapshot |
| Wrong answers, all metrics green | Nothing in telemetry | Quality problem — needs evaluations, not metrics |
| Cardinality explosion | Series count on GenAI metrics | Prompt, user or conversation ID used as a metric label |
| PII in the trace backend | Content attributes in Tempo | Content capture left on without redaction |
The workshop’s flagd has two LLM faults for this table: llmRateLimitError (intermittent 429) and llmInaccurateResponse (a wrong product summary with 200 OK). The second one is the point: every dashboard stays green.