AI Telemetry

An LLM call is an HTTP request that is slow, expensive, non-deterministic and billed per token. Plain HTTP instrumentation sees POST /v1/chat/completions 200 4.2s — nothing about why it cost what it did or why the answer was wrong.

The OpenTelemetry GenAI semantic conventions add what an operator needs: model, tokens, finish reason, tool calls and, opt-in, the prompt itself.

🤖 HTTP span vs GenAI span

Question HTTP instrumentation / eBPF GenAI instrumentation
Which model answered? ❌ Hidden in the body gen_ai.request.model, gen_ai.response.model
How many tokens, what cost? ❌ gen_ai.usage.input_tokens, gen_ai.usage.output_tokens
Was the answer cut off? ❌ gen_ai.response.finish_reasons = ["length"]
Which tool did the agent call? ❌ execute_tool span, gen_ai.tool.name
What was asked and answered? ❌ Opt-in message content
How long did it take? ✅ ✅

The workshop cluster shows both sides of one call:

  • llm — a mock model server with no SDK, instrumented only by Beyla (see Instrumenting Applications). You get a server span with method, route, status and duration.
  • product-reviews — calls llm through the OpenAI client with GenAI instrumentation and OTEL_INSTRUMENTATION_GENAI_CAPTURE_MESSAGE_CONTENT=true in otel-demo.yaml. Its client span carries model, tokens and messages.

📐 GenAI semantic conventions

Status: Development. Names still change between convention versions:

  • gen_ai.system → gen_ai.provider.name,
  • per-message events (gen_ai.user.message, gen_ai.choice) → gen_ai.input.messages / gen_ai.output.messages.

Instrumentations that support the switch select the latest names with OTEL_SEMCONV_STABILITY_OPT_IN=gen_ai_latest_experimental. Pin instrumentation versions, and check which names your version emits before building dashboards on them.

Operations and span names

Span name = {gen_ai.operation.name} {target}.

gen_ai.operation.name Span name example Kind
chat chat gpt-4.1-mini CLIENT
text_completion, generate_content generate_content gemini-2.5-flash CLIENT
embeddings embeddings text-embedding-3-small CLIENT
execute_tool execute_tool search_orders INTERNAL
invoke_agent invoke_agent support-bot CLIENT or INTERNAL
create_agent create_agent support-bot CLIENT

Key attributes

Attribute Example Why it matters
gen_ai.provider.name openai, anthropic, aws.bedrock Which vendor bills you
gen_ai.request.model gpt-4.1-mini What you asked for
gen_ai.response.model gpt-4.1-mini-2025-04-14 What actually answered — aliases move to new snapshots
gen_ai.usage.input_tokens / output_tokens 412 / 96 Cost and context-window pressure
gen_ai.response.finish_reasons ["stop"], ["length"], ["tool_calls"] Truncation, tool loops
gen_ai.request.temperature, max_tokens 0.2, 512 Explains non-determinism and truncation
gen_ai.conversation.id c-7f3a… Groups all turns of one conversation
gen_ai.agent.name, gen_ai.tool.name support-bot, search_orders Agent and tool attribution
error.type 429, timeout Rate limits and provider outages

An LLM client span (illustrative values):

chat astronomy-llm                                  kind=CLIENT  1.8s
  gen_ai.operation.name          = chat
  gen_ai.provider.name           = openai
  gen_ai.request.model           = astronomy-llm
  gen_ai.response.model          = astronomy-llm
  gen_ai.usage.input_tokens      = 412
  gen_ai.usage.output_tokens     = 96
  gen_ai.response.finish_reasons = ["stop"]
  server.address                 = llm
  server.port                    = 8000

gen_ai.provider.name = openai describes the API, not the vendor behind it: any OpenAI-compatible server (the workshop mock, vLLM, a gateway) looks the same. server.address tells you where the call really went.

📊 GenAI metrics

Metric (OTel) Prometheus name Type Use
gen_ai.client.token.usage gen_ai_client_token_usage_{bucket,sum,count} Histogram, gen_ai.token.type = input / output Token burn, cost
gen_ai.client.operation.duration gen_ai_client_operation_duration_seconds_* Histogram Latency per model
gen_ai.server.time_to_first_token gen_ai_server_time_to_first_token_seconds_* Histogram Perceived latency of streaming
gen_ai.server.time_per_output_token gen_ai_server_time_per_output_token_seconds_* Histogram Generation speed

Client metrics come from the instrumented caller; server metrics from the model server. Self-hosted servers (vLLM, TGI) also expose their own Prometheus metrics: queue depth, KV-cache usage, batch size.

# tokens per second, by model and direction
sum by (gen_ai_request_model, gen_ai_token_type) (
  rate(gen_ai_client_token_usage_sum[5m])
)

# p95 LLM call latency by model
histogram_quantile(0.95,
  sum by (le, gen_ai_request_model) (
    rate(gen_ai_client_operation_duration_seconds_bucket[5m])
  )
)

# daily input-token cost, price $0.40 per 1M tokens
sum by (gen_ai_request_model) (
  increase(gen_ai_client_token_usage_sum{gen_ai_token_type="input"}[1d])
) / 1e6 * 0.40

Prices are not telemetry. Keep them in one recording rule, or as a static price series joined with group_left, not scattered across dashboards.

🧪 Prompt and response content

Content capture is off by default, for good reasons:

  • PII and secrets. Users paste emails, customer data and API keys into prompts.
  • Size. A RAG prompt of 50k tokens is hundreds of KB — per span.
  • Cost. Trace storage is priced for small spans.

Where content lands when enabled depends on the convention version: older instrumentations emit it as log events (gen_ai.user.message, gen_ai.choice), newer ones as gen_ai.input.messages, gen_ai.output.messages and gen_ai.system_instructions on the span or on a gen_ai.client.inference.operation.details event.

Production pattern:

Control How
Content to logs, not span attributes Loki with short retention and tighter access than application logs
Redact before storage Collector redaction or transform processor
Capture selectively Only on errors, or for a sampled fraction of calls
Bound attribute size OTEL_ATTRIBUTE_VALUE_LENGTH_LIMIT in the SDK

🕸️ Agents and tool calls

An agent is a loop: model call → tool call → model call, until finish_reasons = ["stop"]. Its trace:

invoke_agent support-bot                              8.4s
├── chat gpt-4.1                                      1.9s  finish_reasons=["tool_calls"]  in=1.2k
├── execute_tool search_orders                        0.3s
│   └── GET /orders?customer=…  (HTTP auto-instr.)    0.2s
├── chat gpt-4.1                                      2.2s  finish_reasons=["tool_calls"]  in=2.9k
├── execute_tool refund_order                         0.4s
│   └── POST /refunds  (HTTP auto-instr.)             0.3s
└── chat gpt-4.1                                      1.6s  finish_reasons=["stop"]        in=3.4k

What it answers:

  • How many iterations? A runaway loop shows up as dozens of chat → execute_tool pairs.
  • Which step is slow? Model time vs tool time vs the backend call inside the tool.
  • Why does cost grow? Every chat re-sends the whole history, so input tokens rise with each iteration.
  • Which conversation? gen_ai.conversation.id stitches turns across traces, as session.id does in the browser.

MCP. Semantic conventions for Model Context Protocol (mcp.method.name, mcp.session.id, tool name) are in development, including context propagation through the request’s _meta field. With it, a tool server’s spans join the agent’s trace instead of starting a new one.

🧰 Instrumentation options

Option Conventions Notes
OTel contrib instrumentations (opentelemetry-instrumentation-openai-v2, Google GenAI, Vertex AI, …) gen_ai.* Vendor-neutral reference implementation
Frameworks with built-in OTel (Microsoft.Extensions.AI, Semantic Kernel, Spring AI, Vercel AI SDK) gen_ai.*, sometimes plus their own Enable the framework’s telemetry switch; no extra library
OpenLLMetry, OpenLIT gen_ai.* (mostly) Broad provider and vector-DB coverage, OTLP to any backend
OpenInference (Arize Phoenix) Own: llm.*, openinference.span.kind OTLP transport, but gen_ai.* dashboards will not match — pick one convention or map in the collector
eBPF (Beyla / OBI) HTTP only Latency and errors of the call, no tokens or model

LLM-focused platforms (Langfuse, Phoenix, Grafana AI Observability) accept OTLP. A collector can fan the same spans out to Tempo for operations and to an evaluation tool for quality work.

AI coding tools as a telemetry source

Developer tools emit GenAI-style telemetry too. Claude Code exports OTel metrics (claude_code.token.usage, claude_code.cost.usage, claude_code.session.count) and log events (claude_code.api_request, claude_code.tool_result):

export CLAUDE_CODE_ENABLE_TELEMETRY=1
export OTEL_METRICS_EXPORTER=otlp
export OTEL_LOGS_EXPORTER=otlp
export OTEL_EXPORTER_OTLP_PROTOCOL=http/protobuf
export OTEL_EXPORTER_OTLP_ENDPOINT=https://otlp.workshop2.indexoutofrange.com

The same privacy rule applies: prompt text is excluded unless OTEL_LOG_USER_PROMPTS=1.

🚨 What to alert on

Symptom Signal Typical cause
Cost spike rate(gen_ai_client_token_usage_sum[…]) jumps Agent loop, prompt regression, traffic
Rising 429 errors error.type on chat spans Provider rate limit, missing backoff
More finish_reasons = ["length"] Span attribute ratio max_tokens too low, answers truncated
Latency regression, no deploy gen_ai.response.model changed Provider moved an alias to a new snapshot
Wrong answers, all metrics green Nothing in telemetry Quality problem — needs evaluations, not metrics
Cardinality explosion Series count on GenAI metrics Prompt, user or conversation ID used as a metric label
PII in the trace backend Content attributes in Tempo Content capture left on without redaction

The workshop’s flagd has two LLM faults for this table: llmRateLimitError (intermittent 429) and llmInaccurateResponse (a wrong product summary with 200 OK). The second one is the point: every dashboard stays green.

results matching ""

    No results matching ""