CI/CD Observability

A pipeline is a distributed system: a workflow schedules jobs on runners, jobs call registries, clouds and clusters, and any step can be slow or flaky. Most teams only see a green or red icon and a log.

OpenTelemetry treats pipeline runs as traces and pipeline health as metrics, with dedicated cicd.* and vcs.* conventions.

🎯 Why pipelines need telemetry

Question Without telemetry With telemetry
Which step made the deploy 10 minutes slower? Compare logs of two runs by hand Span durations per step, across runs
Is the test suite flaky or broken? Re-run and hope Failure rate per task over time
Are we waiting for runners? Invisible Queue time = gap between run created and first job started
Which deploy broke production? Match timestamps by eye Deploy spans and version changes on the same dashboard
How often do we deploy, how long does a change take? Spreadsheet DORA metrics from pipeline and VCS data

πŸ“ CI/CD and VCS semantic conventions

Status: Development. Attribute and metric names may still change.

Trace shape: pipeline run = root span, task (job, stage, step) = child span.

Attribute Example On
cicd.pipeline.name Deploy Kubernetes Cluster Run span
cicd.pipeline.run.id 17234567890 Run span
cicd.pipeline.run.url.full Link to the run Run span
cicd.pipeline.result success, failure, error, timeout, cancellation, skip Run span
cicd.pipeline.task.name Deploy cluster Task span
cicd.pipeline.task.type build, test, deploy Task span
cicd.pipeline.task.run.result failure Task span
cicd.worker.name ubuntu-24.04 runner 12 Task span
vcs.repository.url.full https://github.com/ProtopiaTech/training_Observability_v2 Both
vcs.ref.head.name, vcs.ref.head.revision main, e92e3d5… Both
vcs.change.id Pull request number Both
Metric Type Use
cicd.pipeline.run.duration Histogram, by cicd.pipeline.run.state (pending, executing, finalizing) Queue time vs run time
cicd.pipeline.run.active UpDownCounter Concurrency, runner pressure
cicd.pipeline.run.errors Counter Failed runs by error.type
cicd.system.errors Counter Failures of the CI system itself, not the code
cicd.worker.count UpDownCounter, by cicd.worker.state Runner pool capacity
vcs.change.duration, vcs.change.time_to_approval, vcs.change.time_to_merge Gauge Review and merge latency
vcs.change.count UpDownCounter, by vcs.change.state Open pull requests
vcs.ref.count, vcs.ref.time, vcs.ref.lines_delta UpDownCounter / Gauge Branch count, age and size

deployment.environment.name, deployment.id and deployment.status connect a pipeline run to what it deployed.

🧰 Three ways to get pipeline telemetry

Approach How Good for Limits
1. Collector receiver from CI events CI system sends webhooks; the receiver builds traces after the run GitHub Actions, GitLab CI, no change to pipelines Spans reconstructed from API timestamps; step granularity only
2. CI-native plugin Plugin inside the CI server exports OTLP Jenkins (OpenTelemetry plugin), Tekton, Buildkite Per-platform, different attribute sets
3. Instrument from inside the job otel-cli or an SDK wraps commands and exports spans Detail inside a step: each helm upgrade, each test You maintain the wrapping

Collector github receiver (contrib, alpha): a scraper for VCS metrics through the GitHub API, and a webhook endpoint that turns workflow_run / workflow_job events into traces.

receivers:
  github:
    collection_interval: 300s
    scrapers:
      scraper:
        github_org: ProtopiaTech
        search_query: "org:ProtopiaTech topic:training"
        auth:
          authenticator: bearertokenauth/github
    webhook:
      endpoint: 0.0.0.0:19418
      path: /events
      health_path: /health
      secret: ${env:GITHUB_WEBHOOK_SECRET}   # HMAC check of every delivery

extensions:
  bearertokenauth/github:
    token: ${env:GITHUB_PAT}

service:
  extensions: [bearertokenauth/github]
  pipelines:
    traces:  {receivers: [github], exporters: [otlp/tempo]}
    metrics: {receivers: [github], exporters: [prometheusremotewrite]}

Check the receiver README for your Collector version β€” the config layout of alpha components changes. The gitlab receiver follows the same pattern for GitLab pipeline webhooks.

The webhook endpoint must be reachable from the CI platform. Expose only that path, validate the secret, and run it on a separate Collector from your internal OTLP ingest.

otel-cli inside a job:

- name: Deploy cluster
  env:
    OTEL_EXPORTER_OTLP_ENDPOINT: https://otlp.example.com
    OTEL_SERVICE_NAME: ci-deploy
  run: |
    otel-cli exec --name "helm upgrade tempo" -- \
      helm upgrade --install tempo grafana/tempo -f helm_values/tempo/tempo.values.yaml

otel-cli exec creates a span around the command and passes TRACEPARENT to the child process.

πŸ”— Propagating context into build tools

Processes have no HTTP headers. The OTel spec defines environment variables as context carriers:

Variable Carries
TRACEPARENT W3C trace context
TRACESTATE Vendor trace state
BAGGAGE Baggage

A parent process sets them; an instrumented child reads them and continues the trace. Tools that support it (OpenTelemetry Maven extension, pytest and Gradle plugins, otel-cli) produce one trace per run instead of disconnected spans per tool:

Deploy Kubernetes Cluster                 run      14m02s   result=success
β”œβ”€β”€ Deploy to AKS                         job      13m40s
β”‚   β”œβ”€β”€ Azure Login                       step        8s
β”‚   β”œβ”€β”€ Install CLI tools                 step       41s
β”‚   β”œβ”€β”€ Deploy cluster                    step    10m55s
β”‚   β”‚   β”œβ”€β”€ az deployment group create    otel-cli  6m10s
β”‚   β”‚   β”œβ”€β”€ helm upgrade tempo            otel-cli    52s
β”‚   β”‚   └── helm upgrade otel-demo        otel-cli  2m31s
β”‚   β”œβ”€β”€ Configure Cloudflare DNS          step        6s
β”‚   └── Create Grafana workshop users     step       12s

Job and step names are from this repo’s .github/workflows/deploy.yml; the otel-cli spans are illustrative. With only the webhook receiver, the tree stops at the step level.

πŸš€ Deployments as signals for runtime telemetry

A pipeline’s most useful output for operations is the fact that a deploy happened.

Technique How
Version as a resource attribute Set service.version (and vcs.ref.head.revision) at build time; every span, metric and log carries it
Annotation from version changes Grafana annotation query on count by (service_name, service_version) (target_info)
Annotation from deploy spans Grafana annotation from a Tempo or Loki query for cicd.pipeline.task.type = "deploy"
Compare by version sum by (service_version) (rate(http_server_request_duration_seconds_count{http_response_status_code=~"5.."}[5m])) during a rollout

service.version must be low-cardinality: a release tag or short commit SHA, not a build timestamp.

πŸ“Š DORA metrics from telemetry

DORA metric Source Telemetry-only?
Deployment frequency Count of successful deploy runs on the main branch βœ…
Lead time for changes vcs.change.time_to_merge + deploy run duration βœ… Approximate: merge to deployed, not first commit
Change failure rate Deploys followed by a rollback, hotfix or incident ❌ Needs incident or rollback data
Time to restore Incident open to resolved ❌ Comes from the incident tool
# successful production deploys per day
sum(increase(cicd_pipeline_run_duration_seconds_count{
  cicd_pipeline_name="Deploy Kubernetes Cluster", cicd_pipeline_result="success"}[1d]))

🚨 Common failure modes

Symptom Cause
Pipeline spans far in the past or overlapping Spans reconstructed from webhook timestamps; runner clocks differ
Series count explodes cicd.pipeline.run.id or commit SHA used as a metric label
Secrets in Tempo Command lines or env dumps recorded as span attributes by a wrapper
Spans per tool, no single trace TRACEPARENT not passed to child processes, or the tool does not read it
Webhook endpoint abused Public endpoint without HMAC secret validation
Dashboards break after a receiver upgrade Development-status conventions renamed attributes

results matching ""

    No results matching ""