Head and Tail Sampling

Keeping every span is the most expensive option in observability, and nobody reads 99% of them. Sampling decides which traces you keep.

The decision can be made at the start of a trace (head), after the trace is complete (tail), or both. Each option loses something different.

🎯 Where the decision is made

  Head sampling Tail sampling
Where SDK, at the root span Collector, after the trace is buffered
Knows Trace ID, root span name, attributes at start All spans: status, duration, every attribute
Can keep “all errors” ❌ An error happens after the decision ✅
Cost saved Network, Collector CPU, storage, SDK overhead Storage only — every span still travels to the Collector
Infrastructure None Buffering tier + routing by trace ID
Propagates Yes, via the sampled flag in traceparent No — the decision is local to the Collector

🎲 Head sampling in the SDK

The root service decides. The decision travels downstream in the last byte of traceparent:

traceparent: 00-4bf92f3577b34da6a3ce929d0e0e4736-00f067aa0ba902b7-01
                                                                  └─ 01 = sampled, 00 = not sampled

Built-in samplers, set with OTEL_TRACES_SAMPLER:

Value Behaviour
always_on Keep everything (the default is parentbased_always_on)
always_off Drop everything
traceidratio Keep a fraction, decided from the trace ID; OTEL_TRACES_SAMPLER_ARG=0.1 = 10%
parentbased_always_on Follow the parent’s decision; keep if this is a root
parentbased_traceidratio Follow the parent; apply the ratio only at the root
parentbased_jaeger_remote, jaeger_remote Fetch per-service ratios from a remote endpoint
export OTEL_TRACES_SAMPLER=parentbased_traceidratio
export OTEL_TRACES_SAMPLER_ARG=0.1
  • Always use a parentbased_* sampler in non-root services. A plain traceidratio in the middle of a call chain makes its own decision and cuts traces in half.
  • Non-recording spans still propagate. A dropped span keeps its trace ID and sends -00 downstream, so the whole trace is dropped consistently.
  • The ratio applies to traces, not spans. Setting 10% at the frontend removes 90% of the traces from every service behind it.

🔢 Consistent probability sampling

Different SDKs used to implement traceidratio with different algorithms. Two services with the same trace ID and different ratios could make contradictory decisions, and nothing recorded what probability produced a span.

The newer spec fixes both with two fields in the ot entry of tracestate:

tracestate: ot=th:c;rv:9b8233f7e3a151
Field Meaning
rv 56 bits of randomness; optional — W3C Trace Context Level 2 marks trace IDs as random with a random flag, and the SDK then reads the randomness from the trace ID
th Rejection threshold in hex, trailing zeros dropped; the span is kept when randomness ≥ threshold
th Probability Adjusted count (each kept span represents…)
0 100% 1 span
8 50% 2 spans
c 25% 4 spans
e 12.5% 8 spans
fd70a4 ≈1% ≈100 spans

What this gives you:

  • Consistency across services. Every stage compares the same randomness against its own threshold. A 1% stage keeps a subset of what a 10% stage keeps.
  • Adjusted counts. A backend or connector can multiply by 1 / probability and estimate real request counts from sampled spans.
  • Collector support. The probabilistic_sampler processor has mode: proportional (and equalizing), which reads and writes th instead of hashing the trace ID on its own.

SDK support is rolling out language by language. Check whether your SDK version writes th before you rely on adjusted counts.

🧠 Tail sampling in the Collector

The tail_sampling processor (otelcol.processor.tail_sampling in Alloy) holds spans in memory for decision_wait, then evaluates policies against the whole trace. A trace is kept if any policy says so.

processors:
  tail_sampling:
    decision_wait: 10s                 # how long to wait for late spans
    num_traces: 100000                 # traces held in memory at once
    expected_new_traces_per_sec: 2000  # sizes internal structures
    decision_cache:
      sampled_cache_size: 100000       # remember "keep" for spans that arrive after the decision
    policies:
      - name: errors
        type: status_code
        status_code: {status_codes: [ERROR]}
      - name: slow
        type: latency
        latency: {threshold_ms: 1000}
      - name: checkout-always
        type: string_attribute
        string_attribute: {key: service.name, values: [checkout, payment]}
      - name: slow-and-premium
        type: and
        and:
          and_sub_policy:
            - name: slow
              type: latency
              latency: {threshold_ms: 300}
            - name: premium
              type: string_attribute
              string_attribute: {key: app.customer.tier, values: [premium]}
      - name: baseline
        type: probabilistic
        probabilistic: {sampling_percentage: 5}
Policy type Keeps traces where…
status_code Any span has the given status
latency The trace duration exceeds a threshold
string_attribute, numeric_attribute, boolean_attribute Any span has a matching attribute
ottl_condition An OTTL expression matches a span or span event
span_count The trace has between N and M spans
rate_limiting Budget of spans per second not exceeded
probabilistic Hash of the trace ID falls under a percentage
trace_state tracestate contains a given value
and, composite Sub-policies combined, or each given a share of the throughput

Operational facts:

  • All spans of a trace must reach the same replica. Put a stateless tier with the loadbalancing exporter (routing_key: traceID) in front of the sampling tier. See Topologies on Kubernetes.
  • Memory = trace rate × decision_wait × spans per trace. When num_traces is exceeded, the oldest traces are dropped before a decision — watch otelcol_processor_tail_sampling_sampling_trace_dropped_too_early.
  • Late spans. A span arriving after the decision is kept only if decision_cache remembers the trace. Long async flows (queues, batch jobs) exceed any sensible decision_wait.
  • Latency is measured on the whole trace, from the earliest start to the latest end among the buffered spans.

🔀 Combining head and tail

Tail sampling can only keep what head sampling sent. With 10% head sampling, the errors policy sees 10% of the errors.

Pattern Head Tail Effect
Tail only 100% Errors + slow + 5% baseline Best signal, highest network and Collector cost
Head only 1–10% — Cheapest, errors sampled at the same rate as successes
Both 25% on high-volume, 100% on critical services Errors + slow + small baseline Common production compromise
Head drop at source always_off for health checks and synthetic traffic — Removes noise before it costs anything

Health checks are easier to drop in the Collector’s filter processor than with a custom SDK sampler. See Cost Optimization.

📊 Sampling and derived data

Everything computed from spans after the sampler sees only the kept traces.

Derived data What goes wrong Fix
RED metrics (Tempo metrics-generator, spanmetrics connector after the sampler) Rates undercounted; error rate inflated by the “keep all errors” policy Generate span metrics before tail sampling, on the first Collector tier
Service graph Edges missing for rare calls Same: servicegraph connector before sampling
Exemplars Metric points link to trace IDs that were dropped — “trace not found” See Exemplars
Trace counts in TraceQL count() counts kept traces, not requests Use metrics for volume; traces for examples

Order in a two-tier setup:

# Tier 1 (stateless): metrics from 100% of spans, then route by trace ID
service:
  pipelines:
    traces:
      receivers: [otlp]
      processors: [memory_limiter, k8sattributes]
      exporters: [spanmetrics, loadbalancing]
    metrics/spanmetrics:
      receivers: [spanmetrics]
      exporters: [prometheusremotewrite]
# Tier 2 (stateful): tail_sampling → Tempo

🚨 Common failure modes

Symptom Cause
Traces cut in half, downstream spans missing Non-root service uses traceidratio instead of parentbased_traceidratio
Orphan spans in Tempo with tail sampling on Spans of one trace went to different sampling replicas — missing loadbalancing tier, or a scale event
Collector OOM after a traffic spike num_traces / decision_wait sized for average, not peak, traffic
Error-rate dashboard shows 30% errors RED metrics computed after “keep all errors” tail sampling
Async consumer spans missing They arrived after decision_wait, and decision_cache is too small
Errors missing despite an errors policy Head sampling dropped them first

results matching ""

    No results matching ""