Exemplars from the SDK
- 🎯 What an exemplar is
- 🧰 How the SDK picks exemplars
- 🔗 The path from SDK to Grafana
- 🔍 Checking the workshop stack
- 📉 Where exemplars get lost
- 🚨 Common failure modes
- Related lessons
A p99 latency panel says that something is slow. An exemplar is one real request behind that point, with its trace ID, so one click opens the slow trace.
Exemplars are produced by the metrics SDK at the moment a measurement is recorded. Every hop between the SDK and Grafana can drop them without an error.
🎯 What an exemplar is
An exemplar is a sample measurement attached to a metric data point:
| Field | Example |
|---|---|
value |
2.43 (seconds) |
time_unix_nano |
When the measurement was recorded |
trace_id, span_id |
Span that was active when the measurement was recorded |
filtered_attributes |
Measurement attributes dropped by a View, kept here instead of on the series |
In Prometheus terms it becomes:
http_server_request_duration_seconds_bucket{le="2.5",http_route="/api/checkout"} 1043 # {trace_id="4bf92f35…",span_id="00f067aa…"} 2.43 1758700000.123
Exemplars cost almost nothing: a few per series per export interval, no new series.
🧰 How the SDK picks exemplars
Two components decide:
1. Exemplar filter — which measurements are candidates. Set with OTEL_METRICS_EXEMPLAR_FILTER:
| Value | Candidates |
|---|---|
trace_based |
Measurements recorded while a sampled span is active (spec default) |
always_on |
Every measurement (no trace link for those outside a span) |
always_off |
None |
2. Exemplar reservoir — which candidates are kept per data point:
| Instrument | Default reservoir |
|---|---|
| Explicit-bucket histogram | One exemplar per bucket, the latest measurement that fell into it |
| Exponential histogram | Small fixed-size random sample |
| Counters, gauges, up-down counters | Fixed-size random sample (size 1 by default) |
The trace link comes from the context at record time. The measurement must be recorded inside the span:
with tracer.start_as_current_span("checkout"):
process_order()
duration.record(time.monotonic() - start, {"http.route": "/api/checkout"}) # ✅ exemplar gets this span
duration.record(...) # ❌ span already ended → no trace link
// Go passes context explicitly — context.Background() means no exemplar
hist.Record(ctx, elapsed.Seconds(), metric.WithAttributes(attribute.String("http.route", route)))
Defaults differ by SDK. Java, Go and Python use trace_based in current versions. .NET needs an opt-in (OTEL_METRICS_EXEMPLAR_FILTER=trace_based or SetExemplarFilter(ExemplarFilterType.TraceBased) on the MeterProvider). Check the release notes of the version you run.
Auto-instrumented HTTP server metrics (http.server.request.duration) record inside the server span, so they get exemplars without code changes once the filter allows it.
🔗 The path from SDK to Grafana
Each hop has its own switch. Most are off by default.
| Hop | What must be true | Default |
|---|---|---|
| SDK | Exemplar filter allows it, measurement recorded inside a sampled span | Varies by language |
| OTLP export | Nothing — OTLP carries exemplars | ✅ |
| Collector / Alloy | Exporter translates them: prometheusremotewrite, otelcol.exporter.prometheus |
✅ |
| Remote write | send_exemplars on the sender |
Alloy prometheus.remote_write: on; Prometheus remoteWrite: off |
| Pull path | SDK Prometheus exporter + scraper negotiates OpenMetrics or protobuf; plain text format has no exemplars | Depends on scrape_protocols |
| Prometheus | --enable-feature=exemplar-storage (in-memory ring, storage.exemplars.max_exemplars) |
off |
| Mimir | max_global_exemplars_per_user > 0 |
0 = disabled |
| Grafana | Prometheus data source has an exemplar link whose label name matches the exemplar | Not configured |
Label names differ by producer:
| Producer | Trace ID label |
|---|---|
| OTel SDK via Collector / Alloy | trace_id |
| Tempo metrics-generator | traceID |
Grafana needs one exemplarTraceIdDestinations entry per label name:
jsonData:
exemplarTraceIdDestinations:
- name: traceID # Tempo metrics-generator
datasourceUid: tempo
- name: trace_id # OTel SDKs
datasourceUid: tempo
In Explore or a panel, enable Exemplars in the query options. They show up on histogram queries (histogram_quantile(... _bucket ...)) as dots on the time series.
🔍 Checking the workshop stack
The workshop stack configures the Grafana end, but not the storage chain:
| Hop | Setting in the repo | Result |
|---|---|---|
| Tempo metrics-generator | send_exemplars: true in tempo.values.yaml |
Sends exemplars to Prometheus |
| Alloy gateway | otelcol.exporter.prometheus → prometheus.remote_write |
Forwards SDK exemplars |
| Prometheus | enableFeatures: [native-histograms] in kp-stack.values.yaml |
No exemplar-storage — exemplars dropped on arrival |
| Prometheus → Mimir | remoteWrite without sendExemplars: true |
Not forwarded |
| Mimir | No max_global_exemplars_per_user in mimir.values.yaml |
Rejected |
| Grafana | exemplarTraceIdDestinations: [{name: traceID, …}] |
Matches Tempo’s label only |
To light up the whole path: add exemplar-storage to enableFeatures, sendExemplars: true to the Mimir remoteWrite, max_global_exemplars_per_user to the Mimir limits, and a trace_id destination to both Prometheus data sources.
Verify at the storage layer before blaming Grafana:
# Does Prometheus hold any exemplars for this metric?
curl -s "http://localhost:9090/api/v1/query_exemplars?query=http_server_request_duration_seconds_bucket&start=$(date -d '-1 hour' +%s)&end=$(date +%s)" \
| jq '[.data[].exemplars[]] | length'
📉 Where exemplars get lost
| Step | Why they disappear |
|---|---|
| Recording rules | A recorded series has no exemplars — dashboards built on job:…:rate5m lose the click-through |
| Collector aggregation | Processors that re-aggregate (metricstransform, interval, delta→cumulative) may drop or keep only some |
| Mimir limits | max_global_exemplars_per_user is a ring buffer per tenant: high-volume tenants push old exemplars out within minutes |
| Tail sampling | The exemplar is recorded before the trace is sampled. If tail_sampling drops the trace, the dot opens “trace not found” |
| Head sampling at 0% | With trace_based, unsampled requests produce no exemplars — correct, but the panel looks empty |
Tail sampling and exemplars partly align: “keep slow” and “keep errors” policies keep the traces behind the high-latency buckets, which are the dots you click.
🚨 Common failure modes
| Symptom | Cause |
|---|---|
| Exemplars toggle is on, no dots | Storage chain drops them (Prometheus feature flag, remote write, Mimir limit) |
| Dots for span metrics, none for SDK metrics | Grafana destination configured only for traceID, SDK sends trace_id |
| Dots appear, click shows “trace not found” | Tail sampling dropped the trace, or Tempo retention is shorter than metric retention |
| .NET service never produces exemplars | Exemplar filter left at the .NET default |
| Custom metric has no trace link | Recorded outside the span, or on a background thread without context |
| Exemplars only on some histogram panels | The panel queries a recording rule instead of raw buckets |