Head and Tail Sampling
- 🎯 Where the decision is made
- 🎲 Head sampling in the SDK
- 🔢 Consistent probability sampling
- 🧠 Tail sampling in the Collector
- 🔀 Combining head and tail
- 📊 Sampling and derived data
- 🚨 Common failure modes
- Related lessons
Keeping every span is the most expensive option in observability, and nobody reads 99% of them. Sampling decides which traces you keep.
The decision can be made at the start of a trace (head), after the trace is complete (tail), or both. Each option loses something different.
🎯 Where the decision is made
| Head sampling | Tail sampling | |
|---|---|---|
| Where | SDK, at the root span | Collector, after the trace is buffered |
| Knows | Trace ID, root span name, attributes at start | All spans: status, duration, every attribute |
| Can keep “all errors” | ❌ An error happens after the decision | ✅ |
| Cost saved | Network, Collector CPU, storage, SDK overhead | Storage only — every span still travels to the Collector |
| Infrastructure | None | Buffering tier + routing by trace ID |
| Propagates | Yes, via the sampled flag in traceparent |
No — the decision is local to the Collector |
🎲 Head sampling in the SDK
The root service decides. The decision travels downstream in the last byte of traceparent:
traceparent: 00-4bf92f3577b34da6a3ce929d0e0e4736-00f067aa0ba902b7-01
└─ 01 = sampled, 00 = not sampled
Built-in samplers, set with OTEL_TRACES_SAMPLER:
| Value | Behaviour |
|---|---|
always_on |
Keep everything (the default is parentbased_always_on) |
always_off |
Drop everything |
traceidratio |
Keep a fraction, decided from the trace ID; OTEL_TRACES_SAMPLER_ARG=0.1 = 10% |
parentbased_always_on |
Follow the parent’s decision; keep if this is a root |
parentbased_traceidratio |
Follow the parent; apply the ratio only at the root |
parentbased_jaeger_remote, jaeger_remote |
Fetch per-service ratios from a remote endpoint |
export OTEL_TRACES_SAMPLER=parentbased_traceidratio
export OTEL_TRACES_SAMPLER_ARG=0.1
- Always use a
parentbased_*sampler in non-root services. A plaintraceidratioin the middle of a call chain makes its own decision and cuts traces in half. - Non-recording spans still propagate. A dropped span keeps its trace ID and sends
-00downstream, so the whole trace is dropped consistently. - The ratio applies to traces, not spans. Setting 10% at the frontend removes 90% of the traces from every service behind it.
🔢 Consistent probability sampling
Different SDKs used to implement traceidratio with different algorithms. Two services with the same trace ID and different ratios could make contradictory decisions, and nothing recorded what probability produced a span.
The newer spec fixes both with two fields in the ot entry of tracestate:
tracestate: ot=th:c;rv:9b8233f7e3a151
| Field | Meaning |
|---|---|
rv |
56 bits of randomness; optional — W3C Trace Context Level 2 marks trace IDs as random with a random flag, and the SDK then reads the randomness from the trace ID |
th |
Rejection threshold in hex, trailing zeros dropped; the span is kept when randomness ≥ threshold |
th |
Probability | Adjusted count (each kept span represents…) |
|---|---|---|
0 |
100% | 1 span |
8 |
50% | 2 spans |
c |
25% | 4 spans |
e |
12.5% | 8 spans |
fd70a4 |
≈1% | ≈100 spans |
What this gives you:
- Consistency across services. Every stage compares the same randomness against its own threshold. A 1% stage keeps a subset of what a 10% stage keeps.
- Adjusted counts. A backend or connector can multiply by
1 / probabilityand estimate real request counts from sampled spans. - Collector support. The
probabilistic_samplerprocessor hasmode: proportional(andequalizing), which reads and writesthinstead of hashing the trace ID on its own.
SDK support is rolling out language by language. Check whether your SDK version writes th before you rely on adjusted counts.
🧠 Tail sampling in the Collector
The tail_sampling processor (otelcol.processor.tail_sampling in Alloy) holds spans in memory for decision_wait, then evaluates policies against the whole trace. A trace is kept if any policy says so.
processors:
tail_sampling:
decision_wait: 10s # how long to wait for late spans
num_traces: 100000 # traces held in memory at once
expected_new_traces_per_sec: 2000 # sizes internal structures
decision_cache:
sampled_cache_size: 100000 # remember "keep" for spans that arrive after the decision
policies:
- name: errors
type: status_code
status_code: {status_codes: [ERROR]}
- name: slow
type: latency
latency: {threshold_ms: 1000}
- name: checkout-always
type: string_attribute
string_attribute: {key: service.name, values: [checkout, payment]}
- name: slow-and-premium
type: and
and:
and_sub_policy:
- name: slow
type: latency
latency: {threshold_ms: 300}
- name: premium
type: string_attribute
string_attribute: {key: app.customer.tier, values: [premium]}
- name: baseline
type: probabilistic
probabilistic: {sampling_percentage: 5}
| Policy type | Keeps traces where… |
|---|---|
status_code |
Any span has the given status |
latency |
The trace duration exceeds a threshold |
string_attribute, numeric_attribute, boolean_attribute |
Any span has a matching attribute |
ottl_condition |
An OTTL expression matches a span or span event |
span_count |
The trace has between N and M spans |
rate_limiting |
Budget of spans per second not exceeded |
probabilistic |
Hash of the trace ID falls under a percentage |
trace_state |
tracestate contains a given value |
and, composite |
Sub-policies combined, or each given a share of the throughput |
Operational facts:
- All spans of a trace must reach the same replica. Put a stateless tier with the
loadbalancingexporter (routing_key: traceID) in front of the sampling tier. See Topologies on Kubernetes. - Memory = trace rate ×
decision_wait× spans per trace. Whennum_tracesis exceeded, the oldest traces are dropped before a decision — watchotelcol_processor_tail_sampling_sampling_trace_dropped_too_early. - Late spans. A span arriving after the decision is kept only if
decision_cacheremembers the trace. Long async flows (queues, batch jobs) exceed any sensibledecision_wait. - Latency is measured on the whole trace, from the earliest start to the latest end among the buffered spans.
🔀 Combining head and tail
Tail sampling can only keep what head sampling sent. With 10% head sampling, the errors policy sees 10% of the errors.
| Pattern | Head | Tail | Effect |
|---|---|---|---|
| Tail only | 100% | Errors + slow + 5% baseline | Best signal, highest network and Collector cost |
| Head only | 1–10% | — | Cheapest, errors sampled at the same rate as successes |
| Both | 25% on high-volume, 100% on critical services | Errors + slow + small baseline | Common production compromise |
| Head drop at source | always_off for health checks and synthetic traffic |
— | Removes noise before it costs anything |
Health checks are easier to drop in the Collector’s filter processor than with a custom SDK sampler. See Cost Optimization.
📊 Sampling and derived data
Everything computed from spans after the sampler sees only the kept traces.
| Derived data | What goes wrong | Fix |
|---|---|---|
RED metrics (Tempo metrics-generator, spanmetrics connector after the sampler) |
Rates undercounted; error rate inflated by the “keep all errors” policy | Generate span metrics before tail sampling, on the first Collector tier |
| Service graph | Edges missing for rare calls | Same: servicegraph connector before sampling |
| Exemplars | Metric points link to trace IDs that were dropped — “trace not found” | See Exemplars |
| Trace counts in TraceQL | count() counts kept traces, not requests |
Use metrics for volume; traces for examples |
Order in a two-tier setup:
# Tier 1 (stateless): metrics from 100% of spans, then route by trace ID
service:
pipelines:
traces:
receivers: [otlp]
processors: [memory_limiter, k8sattributes]
exporters: [spanmetrics, loadbalancing]
metrics/spanmetrics:
receivers: [spanmetrics]
exporters: [prometheusremotewrite]
# Tier 2 (stateful): tail_sampling → Tempo
🚨 Common failure modes
| Symptom | Cause |
|---|---|
| Traces cut in half, downstream spans missing | Non-root service uses traceidratio instead of parentbased_traceidratio |
| Orphan spans in Tempo with tail sampling on | Spans of one trace went to different sampling replicas — missing loadbalancing tier, or a scale event |
| Collector OOM after a traffic spike | num_traces / decision_wait sized for average, not peak, traffic |
| Error-rate dashboard shows 30% errors | RED metrics computed after “keep all errors” tail sampling |
| Async consumer spans missing | They arrived after decision_wait, and decision_cache is too small |
Errors missing despite an errors policy |
Head sampling dropped them first |