OTel Arrow
- π― Why another protocol
- π§± Row vs column encoding
- π§° Components
- π Where it fits in a topology
- π Measuring the gain
- π¨ Common failure modes
- Related lessons
OTLP is simple and universal, but it repeats itself: every span carries the same attribute keys, and often the same values, as the span before it.
OTel Arrow (OTAP, the OpenTelemetry Arrow Protocol) sends the same data in a columnar Apache Arrow encoding over long-lived gRPC streams. It pays off on links where bandwidth costs money.
π― Why another protocol
| Link | Bandwidth cost | Arrow worth it? |
|---|---|---|
| SDK β node agent, same node | Free | β SDKs speak OTLP; keep it |
| Agent β gateway, same cluster | Cheap | β Usually not |
| Gateway β gateway, cross-region or cross-cloud | Egress billed per GB | β Main use case |
| Edge site, factory, ship β central cloud | Constrained or metered uplink | β |
| Gateway β SaaS vendor | Egress + vendor ingest | Only if the vendor accepts OTAP |
OTel Arrow does not replace OTLP. It is an optional transport between Collectors, and it falls back to plain OTLP when the peer does not support it.
π§± Row vs column encoding
OTLP (row-oriented protobuf): each span is a message with its own list of key-value attributes.
span 1: {name: "GET /api/cart", http.request.method: "GET", http.route: "/api/cart", service.name: "cart", β¦}
span 2: {name: "GET /api/cart", http.request.method: "GET", http.route: "/api/cart", service.name: "cart", β¦}
span 3: β¦
General-purpose compression (gzip, zstd) finds some of that repetition within one request, then starts over.
OTAP (columnar Arrow record batches):
| Technique | Effect |
|---|---|
| Columns | All http.route values together; similar values compress well |
| Dictionary encoding | Repeated strings ("GET", "/api/cart") become small integers |
| Stream state | Dictionaries and schemas persist for the life of the gRPC stream, so later batches send only new entries |
| Delta encoding | Timestamps and IDs stored as differences |
| zstd on top | Applied to the already-compact Arrow buffers |
The gain grows with repetition: many spans with the same attribute keys and values, large batches, long streams. Small, highly unique payloads gain little.
π§° Components
| Phase | What | Status |
|---|---|---|
| Phase 1 | otelarrow exporter and receiver in opentelemetry-collector-contrib: Arrow on the wire, converted back to OTLP data inside the Collector |
Usable; check the componentβs stability level |
| Phase 2 | A dataflow engine (Rust) that processes Arrow batches end-to-end without converting back to OTLP | Experimental |
# Regional gateway: export over Arrow to the central gateway
exporters:
otelarrow:
endpoint: central-gateway.example.com:4317
tls:
ca_file: /etc/tls/ca.pem
arrow:
num_streams: 4 # parallel long-lived streams
max_stream_lifetime: 10m # reconnect periodically so load balancers can rebalance
sending_queue:
enabled: true
# Central gateway: one port accepts both OTLP/gRPC and Arrow streams
receivers:
otelarrow:
protocols:
grpc:
endpoint: 0.0.0.0:4317
- Fallback is automatic. If the receiver answers as a plain OTLP endpoint, the exporter downgrades to standard OTLP/gRPC.
- Distribution matters.
otelcol-contribincludes both components; smaller distributions and vendor builds may not. List components withotelcol-contrib componentsbefore planning on it. - Config keys of these components still change between releases β check the README for your version.
π Where it fits in a topology
region A region B (central)
ββββββββββββββββββββββββββββ βββββββββββββββββββββββββββββββββ
β apps ββOTLPβββΆ node agent β β β
β node agent ββOTLPβββΆ gateway ββOTAPβββββββΆ gateway βββΆ Tempo / Mimir / Loki
ββββββββββββββββββββββββββββ (WAN, βββββββββββββββββββββββββββββββββ
egress billed)
- Terminate Arrow at a Collector, not at the backend. Tempo, Mimir and Loki ingest OTLP.
- Sampling and filtering come first. Dropping data before the WAN saves more than any encoding.
- Stateful streams and load balancing. A gRPC stream sticks to one backend replica for its lifetime. Behind an L4 load balancer, a new replica gets no traffic until streams reconnect β
max_stream_lifetimebounds that. - Per-stream memory. Dictionaries live in memory on both sides, per stream. More streams = more memory on the receiver.
π Measuring the gain
Published benchmarks show large reductions compared to zstd-compressed OTLP, but results depend on your data. Measure before and after on the same traffic:
| Measure | Where |
|---|---|
| Bytes on the wire | rate(container_network_transmit_bytes_total{pod=~"gateway.*"}[5m]) on the sender, or cloud egress billing |
| Items sent | otelcol_exporter_sent_spans / _metric_points / _log_records β must stay the same |
| CPU and memory | Both gateways: encoding and dictionaries cost CPU and RAM |
| Failures and fallbacks | otelcol_exporter_send_failed_*, exporter logs for downgrade to OTLP |
Bytes per span = network bytes / spans sent. Compare this ratio for OTLP + zstd and for OTAP. If it does not drop meaningfully, the extra complexity is not worth it.
π¨ Common failure modes
| Symptom | Cause |
|---|---|
| No bandwidth saving | Exporter silently fell back to OTLP (receiver is plain OTLP, or a proxy terminates gRPC streams) |
| One central replica overloaded, others idle | Long-lived streams pinned by the load balancer; no max_stream_lifetime |
| Receiver memory grows with sender count | Per-stream dictionary state Γ num_streams Γ senders |
| Streams reset every few seconds | Proxy or load-balancer idle timeout shorter than the streamβs quiet periods |
| Collector fails to start after upgrade | Renamed config keys in a non-stable component |
| Savings smaller than benchmarks | Small batches, high-cardinality attributes, or data already filtered heavily |