Database and Messaging Telemetry

Most latency lives in two places: waiting for a database and waiting for a queue. Plain HTTP instrumentation shows neither.

OpenTelemetry has separate semantic conventions for both. Database conventions are stable. Messaging conventions are still in development, and the async hop through a broker is where traces most often break.

💾 Database client spans

A database call is a CLIENT span created by the driver or ORM instrumentation. Status: stable for the core attributes and db.client.operation.duration.

The stable release renamed most attributes. Instrumentations switch with OTEL_SEMCONV_STABILITY_OPT_IN=database (new names) or database/dup (both, for a migration window):

Old (pre-stable) Stable
db.system db.system.name
db.name db.namespace
db.statement db.query.text
db.operation db.operation.name
db.sql.table, db.mongodb.collection, … db.collection.name
Attribute Example Notes
db.system.name postgresql, mysql, redis, mongodb Required
db.namespace otel Database / schema / keyspace
db.collection.name orders Table or collection
db.operation.name SELECT, findAndModify, HGET  
db.query.summary SELECT orders Low-cardinality description; preferred span name
db.query.text SELECT * FROM orders WHERE id = ? Sanitised by default
db.response.status_code 57014 (Postgres: query cancelled) Vendor error code
db.operation.batch.size 50 Batch calls
server.address, server.port postgresql, 5432 Which instance was called
error.type 57014, java.sql.SQLTimeoutException Set on failure

Span name: {db.query.summary} if the instrumentation produces it, otherwise {db.operation.name} {db.collection.name}, otherwise {db.system.name}. A span named with the full SQL text is a cardinality bug.

In the workshop stack, accounting (.NET) writes to PostgreSQL through Entity Framework Core, enabled with OTEL_DOTNET_AUTO_TRACES_ENTITYFRAMEWORKCORE_INSTRUMENTATION_ENABLED=true in otel-demo.yaml. cart calls Valkey (Redis protocol), and product-reviews (Python) calls PostgreSQL.

🔒 Query text and sanitisation

Case What the instrumentation records
Parameterised query The text as written: SELECT * FROM users WHERE email = $1
Literals inlined in the query Sanitised: literals replaced with ?
Parameter values Only with opt-in db.query.parameter.<key> — leave off in production
  • Sanitisation is best-effort. Parsers miss vendor syntax. Keep a Collector redaction or transform rule as a second line of defence.
  • db.query.text is a span attribute, never a metric label. Use db.query.summary or db.operation.name + db.collection.name for metric dimensions.
  • Size. ORMs generate multi-kilobyte queries. Cap with OTEL_ATTRIBUTE_VALUE_LENGTH_LIMIT.

📊 Database client metrics

Metric Type Status Use
db.client.operation.duration Histogram (s) Stable Latency per system, collection, operation
db.client.connection.count UpDownCounter, db.client.connection.state = idle / used Development Pool saturation
db.client.connection.max UpDownCounter Development Pool size
db.client.connection.pending_requests UpDownCounter Development Callers waiting for a connection
db.client.connection.wait_time Histogram (s) Development Time to get a connection from the pool
db.client.response.returned_rows Histogram Development N+1 queries, unbounded SELECT
# p95 query latency per table, as the application sees it
histogram_quantile(0.95,
  sum by (le, db_collection_name, db_operation_name) (
    rate(db_client_operation_duration_seconds_bucket[5m])
  )
)

# pool saturation: used / max
sum by (service_name, db_client_connection_pool_name) (db_client_connection_count{db_client_connection_state="used"})
/ sum by (service_name, db_client_connection_pool_name) (db_client_connection_max)

Pool wait is the most common invisible latency. A slow span with a fast query on the server side usually means the caller waited for a connection.

🔭 Client view vs server view

Source Sees Misses
Client spans / metrics (SDK) Latency incl. pool wait and network, which service and endpoint called What the database did: plans, locks, I/O
Server statistics (pg_stat_statements via postgres_exporter, Alloy database_observability.postgres) Per-query-shape time, rows, buffer hits, locks Which request or trace caused it
eBPF (Beyla / OBI, Pixie) Wire-protocol calls without an SDK Pool wait, ORM context

The workshop stack has all three for PostgreSQL: pg_stat_statements enabled through the postgresql-init ConfigMap, postgres-exporter for Prometheus metrics, and client spans from accounting and product-reviews. The connection string of accounting sets Application Name=accounting, so pg_stat_activity shows which service holds a connection.

Joining the two views. sqlcommenter, supported by several OTel instrumentations, appends the traceparent as an SQL comment:

SELECT * FROM orders WHERE id = $1 /*traceparent='00-4bf92f35…-00f067aa…-01'*/

The trace ID then appears in the database’s slow-query log and pg_stat_activity, so a slow query can be traced back to its request.

📨 Messaging spans

Status: Development. Names still change between convention versions.

messaging.operation.type Span kind Created by
create PRODUCER Creating a message that is sent later (batch send)
send PRODUCER (CLIENT for a batch of already-created messages) Publishing to the broker
receive CLIENT Pulling messages (poll)
process CONSUMER Handling a message
settle CLIENT Ack, nack, commit offset

Span name: {messaging.operation.name} {destination}, e.g. send orders, process orders.

Attribute Example
messaging.system kafka, rabbitmq, aws_sqs, servicebus, eventhubs
messaging.operation.name send, poll, process, ack
messaging.destination.name orders
messaging.consumer.group.name accounting
messaging.destination.partition.id 3
messaging.message.id Broker or client message ID
messaging.batch.message_count 50
messaging.kafka.offset, messaging.kafka.message.key Kafka-specific

In the OTel Demo, checkout (Go) sends to the Kafka topic orders, and accounting (.NET) and fraud-detection (Kotlin, Java agent) consume from it.

🔗 Context through the broker

The producer injects traceparent into the message headers (Kafka record headers, AMQP application properties). The consumer extracts it. The message itself is the carrier.

checkout: POST /checkout                                   120ms
└── send orders                   PRODUCER                   4ms
      ┆  (message sits in Kafka)
      ├── process orders          CONSUMER  accounting      35ms
      │   └── INSERT orders       CLIENT    postgresql      12ms
      └── process orders          CONSUMER  fraud-detection 18ms

Parent or link?

Consumer pattern Relationship Result
One message → one process span Child of the producer context (most instrumentations), plus a link One trace across the broker
Batch of N messages → one process span Links to each message’s context The batch span starts its own trace; each link leads to one producer trace
Receive span separate from process receive links to message contexts; process may be its child Depends on instrumentation

Consequences:

  • Trace duration includes queue time. Hours of backlog make the trace hours long. Do not alert on trace duration for async flows; use consumer lag.
  • Fan-out makes huge traces. A message consumed by 20 services puts 20 sub-trees in one trace. Tail sampling’s decision_wait will not cover a slow consumer.
  • Links need backend support. Tempo stores span links. Grafana shows them in the span details. See span links in Traces.

📈 Messaging metrics and consumer lag

Metric Type Use
messaging.client.operation.duration Histogram (s) Send / receive latency
messaging.client.sent.messages Counter Producer throughput
messaging.client.consumed.messages Counter Consumer throughput
messaging.process.duration Histogram (s) Handler latency

Consumer lag is not an SDK metric. Neither the producer nor the consumer knows how far behind it is. Measure it on the broker side:

Source Metric
Collector kafkametrics receiver kafka.consumer_group.lag, kafka.consumer_group.lag_sum
kafka_exporter / Strimzi Kafka Exporter kafka_consumergroup_lag
Cloud brokers Service Bus active message count, SQS ApproximateAgeOfOldestMessage

Alert on lag growth (deriv(...) over 10–15 minutes) or on the age of the oldest message, not on a fixed message count.

🚨 Common failure modes

Symptom Cause
Consumer spans start a new trace Headers not propagated: manual producer code bypasses the instrumented client, a proxy strips headers, or a bridge re-publishes the message
Span names contain full SQL Old instrumentation or custom spans named from the query text
Customer e-mails in Tempo Literals inlined into queries, sanitisation missed a vendor syntax, or parameter capture left on
Slow DB spans, fast queries in pg_stat_statements Connection-pool wait or network, not the database
Dashboard broke after an upgrade Instrumentation switched from db.statement to db.query.text (or old to new messaging names)
One trace with thousands of spans Batch consumer creates child spans instead of links, or a fan-out topic
Throughput fine, users see stale data Consumer lag growing — not visible in SDK metrics

results matching ""

    No results matching ""