OpAMP — Managing Agents Remotely
- 🎯 What OpAMP solves
- 🔌 Protocol model
- 🧰 Capabilities
- 🏭 OpAMP for the Collector
- 📦 Beyond the Collector
- 🔒 Operating a fleet safely
- 🚨 Common failure modes
- Related lessons
A hundred Collectors, each with a config file, each a slightly different version. Which config is actually running? Which agents are unhealthy? How do you change a sampling rate everywhere without redeploying?
OpAMP (Open Agent Management Protocol) is the OpenTelemetry protocol between agents and a management server: status up, configuration down.
🎯 What OpAMP solves
| Need | Without OpAMP | With OpAMP |
|---|---|---|
| Inventory | Grep deployments, hope labels are right | Every agent reports identity, version, host and available components |
| Effective config | The file in git, maybe | The config the agent actually loaded, reported by the agent |
| Health | Scrape each agent’s metrics | Health reported over the management connection |
| Config change | Redeploy or restart | Server pushes a new config; agent reports applied or failed |
| Agent’s own telemetry | Configure each agent’s self-monitoring | Server tells agents where to send their own metrics, logs and traces |
| Credentials | Rotate by redeploy | Server offers new connection settings and certificates |
OpAMP is vendor-neutral. The same agent can be managed by any OpAMP server; the same server can manage different agent types.
🔌 Protocol model
| Aspect | Detail |
|---|---|
| Direction | Agent connects to the server; the server never dials agents |
| Transport | WebSocket (server can push at any time) or plain HTTP (agent polls) |
| Encoding | Protobuf: AgentToServer and ServerToAgent messages |
| Agent identity | instance_uid, 16 bytes (UUID v7 recommended), stable across restarts |
| Agent description | identifying_attributes (service.name, service.instance.id, service.version) and non_identifying_attributes (host.name, os.type, …) |
| Efficiency | Agents send only changed fields; a sequence number lets the server request a full status if it lost state |
The agent description uses resource attributes with the identifying/descriptive split from Entities and the Resource Model.
Agent ──AgentToServer──▶ Server
instance_uid, description, capabilities,
health, effective_config, remote_config_status
Agent ◀──ServerToAgent── Server
remote_config (hash + files), connection_settings,
packages_available, command (restart)
🧰 Capabilities
An agent declares what it supports; the server must not ask for more. Capabilities are flags:
| Capability | Meaning |
|---|---|
ReportsStatus |
Sends description and status (mandatory) |
ReportsEffectiveConfig |
Sends the config it is running |
AcceptsRemoteConfig |
Applies config sent by the server |
ReportsRemoteConfig |
Reports APPLYING / APPLIED / FAILED for a remote config |
ReportsHealth |
Sends component health |
ReportsOwnMetrics, ReportsOwnTraces, ReportsOwnLogs |
Can send its own telemetry to a destination the server offers |
AcceptsOpAMPConnectionSettings |
Accepts a new server endpoint or credentials |
AcceptsOtherConnectionSettings |
Accepts credentials for its exporters |
AcceptsPackages, ReportsPackageStatuses |
Downloads and installs packages (plugins, binaries) |
AcceptsRestartCommand |
Can be restarted by the server |
ReportsAvailableComponents |
Lists the receivers, processors and exporters it was built with |
ReportsHeartbeat |
Sends periodic heartbeats |
Remote config is a map of named files plus a hash. The agent applies it, reports the result with the same hash, and includes the resulting effective config. The server compares desired and effective config to detect drift.
🏭 OpAMP for the Collector
Two components, often used together:
| Component | Runs | Does |
|---|---|---|
opamp extension |
Inside the Collector | Reports description, effective config, health and available components. Does not apply remote config |
OpAMP Supervisor (opampsupervisor) |
Separate process that starts the Collector | Talks to the server, writes the merged config, restarts the Collector on change, reports config status, keeps the last known config if the server is unreachable |
# supervisor.yaml
server:
endpoint: wss://opamp.example.com/v1/opamp
tls:
ca_file: /etc/opamp/ca.pem
capabilities:
accepts_remote_config: true
reports_effective_config: true
reports_remote_config: true
reports_health: true
reports_own_metrics: true
agent:
executable: /otelcol-contrib
storage:
directory: /var/lib/otelcol/supervisor # last received config, instance UID
The Supervisor exists because a process cannot reliably replace its own config and restart itself: a bad config would kill the only component able to roll it back.
On Kubernetes, the OTel Operator’s OpAMP Bridge (OpAMPBridge resource) connects OpenTelemetryCollector resources to an OpAMP server. The server manages the CRs, and the Operator does the rollout.
Alloy is managed differently. Grafana Alloy uses its remotecfg block and Grafana Fleet Management, not OpAMP. See Alloy vs OpenTelemetry Collector vs Vector.
📦 Beyond the Collector
- SDKs. OpAMP clients for SDKs and language agents are emerging (e.g. in the Java contrib repo). Use case: change sampling or log level of a running service without a redeploy.
- Packages. The protocol can distribute binaries and plugins. On Kubernetes this conflicts with image-based deploys. Most teams manage config through OpAMP and binaries through their normal pipeline, as Config Lifecycle and Upgrades recommends.
- Servers.
opamp-goships the protocol library and an example server. Commercial agent-management products implement OpAMP; choose one that shows effective config and config status, not only “sent”.
🔒 Operating a fleet safely
Remote config is remote control of your telemetry pipeline. A pushed config can add an exporter that copies all data elsewhere.
| Control | How |
|---|---|
| Authenticate both sides | mTLS or bearer tokens on the OpAMP connection; AcceptsOpAMPConnectionSettings for rotation |
| Limit what the server can change | Supervisor merges remote config with local files; keep receivers, auth and TLS in local config |
| Review config like code | Keep the source of truth in git; the server deploys from it, not from a UI edit box |
| Stage rollouts | Push to a canary group by identifying attributes first; watch FAILED statuses and data-loss metrics |
| Alert on drift | Effective config hash differs from desired for more than N minutes |
| Alert on silence | Agents that stopped reporting are as important as agents that report errors |
🚨 Common failure modes
| Symptom | Cause |
|---|---|
| Server shows “config sent”, agent runs the old one | Agent lacks AcceptsRemoteConfig (extension only, no Supervisor) |
| Fleet restarts at once after a config push | No staged rollout; every agent applies immediately |
| Collector crash-loops after a remote config | Invalid config; Supervisor reports FAILED, but no one alerts on it |
| Duplicate agents in the inventory | instance_uid not persisted; each restart registers a new agent |
| Agents managed by two tools fight | OpAMP and a GitOps controller both write the same config |
| Management outage stops telemetry | Agent configured to wait for remote config at startup, no local fallback |