CI/CD Observability
- π― Why pipelines need telemetry
- π CI/CD and VCS semantic conventions
- π§° Three ways to get pipeline telemetry
- π Propagating context into build tools
- π Deployments as signals for runtime telemetry
- π DORA metrics from telemetry
- π¨ Common failure modes
- Related lessons
A pipeline is a distributed system: a workflow schedules jobs on runners, jobs call registries, clouds and clusters, and any step can be slow or flaky. Most teams only see a green or red icon and a log.
OpenTelemetry treats pipeline runs as traces and pipeline health as metrics, with dedicated
cicd.*andvcs.*conventions.
π― Why pipelines need telemetry
| Question | Without telemetry | With telemetry |
|---|---|---|
| Which step made the deploy 10 minutes slower? | Compare logs of two runs by hand | Span durations per step, across runs |
| Is the test suite flaky or broken? | Re-run and hope | Failure rate per task over time |
| Are we waiting for runners? | Invisible | Queue time = gap between run created and first job started |
| Which deploy broke production? | Match timestamps by eye | Deploy spans and version changes on the same dashboard |
| How often do we deploy, how long does a change take? | Spreadsheet | DORA metrics from pipeline and VCS data |
π CI/CD and VCS semantic conventions
Status: Development. Attribute and metric names may still change.
Trace shape: pipeline run = root span, task (job, stage, step) = child span.
| Attribute | Example | On |
|---|---|---|
cicd.pipeline.name |
Deploy Kubernetes Cluster |
Run span |
cicd.pipeline.run.id |
17234567890 |
Run span |
cicd.pipeline.run.url.full |
Link to the run | Run span |
cicd.pipeline.result |
success, failure, error, timeout, cancellation, skip |
Run span |
cicd.pipeline.task.name |
Deploy cluster |
Task span |
cicd.pipeline.task.type |
build, test, deploy |
Task span |
cicd.pipeline.task.run.result |
failure |
Task span |
cicd.worker.name |
ubuntu-24.04 runner 12 |
Task span |
vcs.repository.url.full |
https://github.com/ProtopiaTech/training_Observability_v2 |
Both |
vcs.ref.head.name, vcs.ref.head.revision |
main, e92e3d5β¦ |
Both |
vcs.change.id |
Pull request number | Both |
| Metric | Type | Use |
|---|---|---|
cicd.pipeline.run.duration |
Histogram, by cicd.pipeline.run.state (pending, executing, finalizing) |
Queue time vs run time |
cicd.pipeline.run.active |
UpDownCounter | Concurrency, runner pressure |
cicd.pipeline.run.errors |
Counter | Failed runs by error.type |
cicd.system.errors |
Counter | Failures of the CI system itself, not the code |
cicd.worker.count |
UpDownCounter, by cicd.worker.state |
Runner pool capacity |
vcs.change.duration, vcs.change.time_to_approval, vcs.change.time_to_merge |
Gauge | Review and merge latency |
vcs.change.count |
UpDownCounter, by vcs.change.state |
Open pull requests |
vcs.ref.count, vcs.ref.time, vcs.ref.lines_delta |
UpDownCounter / Gauge | Branch count, age and size |
deployment.environment.name, deployment.id and deployment.status connect a pipeline run to what it deployed.
π§° Three ways to get pipeline telemetry
| Approach | How | Good for | Limits |
|---|---|---|---|
| 1. Collector receiver from CI events | CI system sends webhooks; the receiver builds traces after the run | GitHub Actions, GitLab CI, no change to pipelines | Spans reconstructed from API timestamps; step granularity only |
| 2. CI-native plugin | Plugin inside the CI server exports OTLP | Jenkins (OpenTelemetry plugin), Tekton, Buildkite | Per-platform, different attribute sets |
| 3. Instrument from inside the job | otel-cli or an SDK wraps commands and exports spans |
Detail inside a step: each helm upgrade, each test |
You maintain the wrapping |
Collector github receiver (contrib, alpha): a scraper for VCS metrics through the GitHub API, and a webhook endpoint that turns workflow_run / workflow_job events into traces.
receivers:
github:
collection_interval: 300s
scrapers:
scraper:
github_org: ProtopiaTech
search_query: "org:ProtopiaTech topic:training"
auth:
authenticator: bearertokenauth/github
webhook:
endpoint: 0.0.0.0:19418
path: /events
health_path: /health
secret: ${env:GITHUB_WEBHOOK_SECRET} # HMAC check of every delivery
extensions:
bearertokenauth/github:
token: ${env:GITHUB_PAT}
service:
extensions: [bearertokenauth/github]
pipelines:
traces: {receivers: [github], exporters: [otlp/tempo]}
metrics: {receivers: [github], exporters: [prometheusremotewrite]}
Check the receiver README for your Collector version β the config layout of alpha components changes. The gitlab receiver follows the same pattern for GitLab pipeline webhooks.
The webhook endpoint must be reachable from the CI platform. Expose only that path, validate the secret, and run it on a separate Collector from your internal OTLP ingest.
otel-cli inside a job:
- name: Deploy cluster
env:
OTEL_EXPORTER_OTLP_ENDPOINT: https://otlp.example.com
OTEL_SERVICE_NAME: ci-deploy
run: |
otel-cli exec --name "helm upgrade tempo" -- \
helm upgrade --install tempo grafana/tempo -f helm_values/tempo/tempo.values.yaml
otel-cli exec creates a span around the command and passes TRACEPARENT to the child process.
π Propagating context into build tools
Processes have no HTTP headers. The OTel spec defines environment variables as context carriers:
| Variable | Carries |
|---|---|
TRACEPARENT |
W3C trace context |
TRACESTATE |
Vendor trace state |
BAGGAGE |
Baggage |
A parent process sets them; an instrumented child reads them and continues the trace. Tools that support it (OpenTelemetry Maven extension, pytest and Gradle plugins, otel-cli) produce one trace per run instead of disconnected spans per tool:
Deploy Kubernetes Cluster run 14m02s result=success
βββ Deploy to AKS job 13m40s
β βββ Azure Login step 8s
β βββ Install CLI tools step 41s
β βββ Deploy cluster step 10m55s
β β βββ az deployment group create otel-cli 6m10s
β β βββ helm upgrade tempo otel-cli 52s
β β βββ helm upgrade otel-demo otel-cli 2m31s
β βββ Configure Cloudflare DNS step 6s
β βββ Create Grafana workshop users step 12s
Job and step names are from this repoβs .github/workflows/deploy.yml; the otel-cli spans are illustrative. With only the webhook receiver, the tree stops at the step level.
π Deployments as signals for runtime telemetry
A pipelineβs most useful output for operations is the fact that a deploy happened.
| Technique | How |
|---|---|
| Version as a resource attribute | Set service.version (and vcs.ref.head.revision) at build time; every span, metric and log carries it |
| Annotation from version changes | Grafana annotation query on count by (service_name, service_version) (target_info) |
| Annotation from deploy spans | Grafana annotation from a Tempo or Loki query for cicd.pipeline.task.type = "deploy" |
| Compare by version | sum by (service_version) (rate(http_server_request_duration_seconds_count{http_response_status_code=~"5.."}[5m])) during a rollout |
service.version must be low-cardinality: a release tag or short commit SHA, not a build timestamp.
π DORA metrics from telemetry
| DORA metric | Source | Telemetry-only? |
|---|---|---|
| Deployment frequency | Count of successful deploy runs on the main branch | β |
| Lead time for changes | vcs.change.time_to_merge + deploy run duration |
β Approximate: merge to deployed, not first commit |
| Change failure rate | Deploys followed by a rollback, hotfix or incident | β Needs incident or rollback data |
| Time to restore | Incident open to resolved | β Comes from the incident tool |
# successful production deploys per day
sum(increase(cicd_pipeline_run_duration_seconds_count{
cicd_pipeline_name="Deploy Kubernetes Cluster", cicd_pipeline_result="success"}[1d]))
π¨ Common failure modes
| Symptom | Cause |
|---|---|
| Pipeline spans far in the past or overlapping | Spans reconstructed from webhook timestamps; runner clocks differ |
| Series count explodes | cicd.pipeline.run.id or commit SHA used as a metric label |
| Secrets in Tempo | Command lines or env dumps recorded as span attributes by a wrapper |
| Spans per tool, no single trace | TRACEPARENT not passed to child processes, or the tool does not read it |
| Webhook endpoint abused | Public endpoint without HMAC secret validation |
| Dashboards break after a receiver upgrade | Development-status conventions renamed attributes |