OpAMP — Managing Agents Remotely

A hundred Collectors, each with a config file, each a slightly different version. Which config is actually running? Which agents are unhealthy? How do you change a sampling rate everywhere without redeploying?

OpAMP (Open Agent Management Protocol) is the OpenTelemetry protocol between agents and a management server: status up, configuration down.

🎯 What OpAMP solves

Need Without OpAMP With OpAMP
Inventory Grep deployments, hope labels are right Every agent reports identity, version, host and available components
Effective config The file in git, maybe The config the agent actually loaded, reported by the agent
Health Scrape each agent’s metrics Health reported over the management connection
Config change Redeploy or restart Server pushes a new config; agent reports applied or failed
Agent’s own telemetry Configure each agent’s self-monitoring Server tells agents where to send their own metrics, logs and traces
Credentials Rotate by redeploy Server offers new connection settings and certificates

OpAMP is vendor-neutral. The same agent can be managed by any OpAMP server; the same server can manage different agent types.

🔌 Protocol model

Aspect Detail
Direction Agent connects to the server; the server never dials agents
Transport WebSocket (server can push at any time) or plain HTTP (agent polls)
Encoding Protobuf: AgentToServer and ServerToAgent messages
Agent identity instance_uid, 16 bytes (UUID v7 recommended), stable across restarts
Agent description identifying_attributes (service.name, service.instance.id, service.version) and non_identifying_attributes (host.name, os.type, …)
Efficiency Agents send only changed fields; a sequence number lets the server request a full status if it lost state

The agent description uses resource attributes with the identifying/descriptive split from Entities and the Resource Model.

Agent ──AgentToServer──▶ Server
        instance_uid, description, capabilities,
        health, effective_config, remote_config_status

Agent ◀──ServerToAgent── Server
        remote_config (hash + files), connection_settings,
        packages_available, command (restart)

🧰 Capabilities

An agent declares what it supports; the server must not ask for more. Capabilities are flags:

Capability Meaning
ReportsStatus Sends description and status (mandatory)
ReportsEffectiveConfig Sends the config it is running
AcceptsRemoteConfig Applies config sent by the server
ReportsRemoteConfig Reports APPLYING / APPLIED / FAILED for a remote config
ReportsHealth Sends component health
ReportsOwnMetrics, ReportsOwnTraces, ReportsOwnLogs Can send its own telemetry to a destination the server offers
AcceptsOpAMPConnectionSettings Accepts a new server endpoint or credentials
AcceptsOtherConnectionSettings Accepts credentials for its exporters
AcceptsPackages, ReportsPackageStatuses Downloads and installs packages (plugins, binaries)
AcceptsRestartCommand Can be restarted by the server
ReportsAvailableComponents Lists the receivers, processors and exporters it was built with
ReportsHeartbeat Sends periodic heartbeats

Remote config is a map of named files plus a hash. The agent applies it, reports the result with the same hash, and includes the resulting effective config. The server compares desired and effective config to detect drift.

🏭 OpAMP for the Collector

Two components, often used together:

Component Runs Does
opamp extension Inside the Collector Reports description, effective config, health and available components. Does not apply remote config
OpAMP Supervisor (opampsupervisor) Separate process that starts the Collector Talks to the server, writes the merged config, restarts the Collector on change, reports config status, keeps the last known config if the server is unreachable
# supervisor.yaml
server:
  endpoint: wss://opamp.example.com/v1/opamp
  tls:
    ca_file: /etc/opamp/ca.pem
capabilities:
  accepts_remote_config: true
  reports_effective_config: true
  reports_remote_config: true
  reports_health: true
  reports_own_metrics: true
agent:
  executable: /otelcol-contrib
storage:
  directory: /var/lib/otelcol/supervisor   # last received config, instance UID

The Supervisor exists because a process cannot reliably replace its own config and restart itself: a bad config would kill the only component able to roll it back.

On Kubernetes, the OTel Operator’s OpAMP Bridge (OpAMPBridge resource) connects OpenTelemetryCollector resources to an OpAMP server. The server manages the CRs, and the Operator does the rollout.

Alloy is managed differently. Grafana Alloy uses its remotecfg block and Grafana Fleet Management, not OpAMP. See Alloy vs OpenTelemetry Collector vs Vector.

📦 Beyond the Collector

  • SDKs. OpAMP clients for SDKs and language agents are emerging (e.g. in the Java contrib repo). Use case: change sampling or log level of a running service without a redeploy.
  • Packages. The protocol can distribute binaries and plugins. On Kubernetes this conflicts with image-based deploys. Most teams manage config through OpAMP and binaries through their normal pipeline, as Config Lifecycle and Upgrades recommends.
  • Servers. opamp-go ships the protocol library and an example server. Commercial agent-management products implement OpAMP; choose one that shows effective config and config status, not only “sent”.

🔒 Operating a fleet safely

Remote config is remote control of your telemetry pipeline. A pushed config can add an exporter that copies all data elsewhere.

Control How
Authenticate both sides mTLS or bearer tokens on the OpAMP connection; AcceptsOpAMPConnectionSettings for rotation
Limit what the server can change Supervisor merges remote config with local files; keep receivers, auth and TLS in local config
Review config like code Keep the source of truth in git; the server deploys from it, not from a UI edit box
Stage rollouts Push to a canary group by identifying attributes first; watch FAILED statuses and data-loss metrics
Alert on drift Effective config hash differs from desired for more than N minutes
Alert on silence Agents that stopped reporting are as important as agents that report errors

🚨 Common failure modes

Symptom Cause
Server shows “config sent”, agent runs the old one Agent lacks AcceptsRemoteConfig (extension only, no Supervisor)
Fleet restarts at once after a config push No staged rollout; every agent applies immediately
Collector crash-loops after a remote config Invalid config; Supervisor reports FAILED, but no one alerts on it
Duplicate agents in the inventory instance_uid not persisted; each restart registers a new agent
Agents managed by two tools fight OpAMP and a GitOps controller both write the same config
Management outage stops telemetry Agent configured to wait for remote config at startup, no local fallback

results matching ""

    No results matching ""