Collecting Azure Metrics and Logs
- What Azure exposes
- The pipeline
- Authentication: workload identity
- Metrics: prometheus.exporter.azure
- Logs: loki.source.azure_event_hubs
- Monitoring the hub itself
- The hub’s own logs
- Failure modes
- Alloy vs azure-metrics-exporter vs Vector
- Decision table
- In our stack
- Related lessons
Managed services have no node to put an agent on. The AKS control plane, a PostgreSQL Flexible Server and a storage account report only to Azure Monitor.
Collecting them means two different paths: pull metrics from the Azure Monitor API, and push logs out of Azure through Event Hubs.
What Azure exposes
| Signal | Source in Azure | How you get it out | Latency |
|---|---|---|---|
| Platform metrics | Azure Monitor, per resource; 1-minute minimum grain (some metrics every 20 min or hourly) | Pull: Azure Monitor metrics REST API | Under a minute in Azure, plus the scrape interval |
| Resource logs | Per-resource categories (e.g. kube-audit-admin, PostgreSQLLogs) |
Push: a diagnostic setting per resource streams them to Log Analytics, a storage account, Event Hubs or a partner | Usually 3–10 min; a new diagnostic setting can take up to 90 min to deliver |
| Activity log | Subscription-level control-plane operations | Collected automatically; a diagnostic setting on the subscription exports it | 3–20 min |
- Metrics are not pushed. Something has to poll the API. Every poll is an ARM API call that counts against throttling limits.
- Resource logs are not collected by default. Without a diagnostic setting they are not stored anywhere. Event Hubs is the only destination a collector can read as a stream.
- Idle often means NULL, not 0. Each resource provider decides whether an idle minute is
0or no value, and a dimension value has no series until it had data. Error and throttling series often appear only with the first error — query withor vector(0).
The pipeline
flowchart LR
subgraph az ["Azure"]
res["AKS · PostgreSQL"]
sto["Storage account<br/>(metrics only)"]
mon["Azure Monitor metrics API"]
diag["Diagnostic settings"]
eh[("Event Hubs<br/>main namespace")]
eh2[("Event Hubs<br/>-logs namespace<br/>(metrics only)")]
res --> mon
sto --> mon
eh --> mon
eh2 --> mon
res --> diag --> eh
eh -- "its own logs" --> eh2
end
subgraph k8s ["AKS — alloy-azure (1 replica, workload identity)"]
exp["prometheus.exporter.azure"] --> scr["prometheus.scrape"]
src["loki.source.azure_event_hubs ×2"] --> proc["loki.process"]
end
mon -- "pull, every 60 s" --> exp
eh -- "consume (Kafka, OAuth)" --> src
eh2 -- "consume (Kafka, OAuth)" --> src
scr --> prom[(Prometheus)]
proc --> loki[(Loki)]
Authentication: workload identity
No secret in the cluster. The pod proves its identity with a Kubernetes token; Entra ID trusts that token through a federated credential.
- Pod label
azure.workload.identity/use: "true"+ ServiceAccount annotationazure.workload.identity/client-id→ the AKS webhook injectsAZURE_CLIENT_ID,AZURE_TENANT_ID,AZURE_AUTHORITY_HOST,AZURE_FEDERATED_TOKEN_FILEand mounts a projected ServiceAccount token. - Both components use the Azure SDK
DefaultAzureCredential, which exchanges that token for Entra ID tokens of the managed identity — one for ARM (metrics), one forhttps://<ns>.servicebus.windows.net/.default(Event Hubs). - Entra ID accepts it because of a federated credential: issuer = the cluster’s OIDC issuer URL, subject =
system:serviceaccount:monitoring:alloy-azure, audience =api://AzureADTokenExchange. - Event Hubs receives the token over Kafka
SASL_SSL+OAUTHBEARER; Azure Monitor receives it as a bearer token.
| Role | Scope | Needed for |
|---|---|---|
| Monitoring Reader | Resource group | Metrics API + Resource Graph discovery |
| Azure Event Hubs Data Receiver | Event Hubs namespace | Reading the hub (not sending, not managing) |
The only key involved is the namespace’s SAS rule (Manage, Send, Listen) that diagnostic settings use to write into the hub. Alloy never sees it.
Metrics: prometheus.exporter.azure
Excerpt — one of ten exporter blocks; the real config has more metrics and concatenates all exporter targets:
prometheus.exporter.azure "aks_nodes" {
subscriptions = [sys.env("AZURE_SUBSCRIPTION_ID")]
resource_type = "Microsoft.ContainerService/managedClusters"
resource_graph_query_filter = string.format("where resourceGroup =~ '%s'", sys.env("AZURE_RESOURCE_GROUP"))
metrics = ["node_cpu_usage_percentage", "node_memory_working_set_percentage"]
included_dimensions = ["node"]
}
prometheus.scrape "azure" {
targets = prometheus.exporter.azure.aks_nodes.targets // array.concat(...) of all blocks
job_name = "integrations/azure"
scrape_interval = "60s" // matches the 1-minute grain; each scrape = API calls
scrape_timeout = "50s"
forward_to = [prometheus.relabel.azure_dimensions.receiver]
}
| Behaviour | Consequence |
|---|---|
| Discovery via Azure Resource Graph, on every scrape | Query = Resources \| where type =~ "<type>" \| <filter> \| project id, tags. The exporter adds the \| itself — write where …; a leading \| gives a parser error |
regions instead of a filter |
Subscription-scope mode: one metrics call per subscription and region instead of one per resource — the answer to throttling with many resources. Mutually exclusive with resource_graph_query_filter, so it cannot be limited to one resource group |
included_dimensions becomes one $filter for every metric in the block |
With validate_dimensions = false (default) Azure ignores dimensions a metric does not have. We still use one block per dimension set (AKS: 6 blocks) so each metric family has one label shape |
| Generic dimension labels | One dimension → label dimension; several → dimensionNode, dimensionCondition, … Rename them in prometheus.relabel (node, request_kind, entity) |
metric_aggregations not set |
Each metric uses its primary aggregation; the name says which: azure_<type>_<metric>_<aggregation>_<unit> |
| “Total” metrics are per-minute values, not counters | The exporter queries a 5-minute timespan and exposes the last non-null minute. sum_over_time(x[$__range]) is an approximate total — an idle or late minute can be exposed on two scrapes. Never rate() |
| No caching | Every scrape runs the Resource Graph query plus the metric calls (one call per resource per 20 metrics). More blocks or a shorter interval = more API calls |
Names get long: azure_microsoft_containerservice_managedclusters_apiserver_cpu_usage_percentage_average_percent. That is the price of a name that says type, metric, aggregation and unit.
Logs: loki.source.azure_event_hubs
loki.source.azure_event_hubs "resource_logs" {
fully_qualified_namespace = sys.env("EVENTHUB_KAFKA_ENDPOINT") // <ns>.servicebus.windows.net:9093
event_hubs = [sys.env("EVENTHUB_NAME")]
group_id = "alloy-azure"
use_incoming_timestamp = true
labels = { job = "integrations/azure_event_hubs", source = "azure" }
relabel_rules = loki.relabel.azure.rules // __azure_event_hubs_category → category
authentication { mechanism = "oauth" }
forward_to = [loki.process.azure.receiver]
}
All Azure lines are under {source="azure"}, with category, resource and resource_type labels.
- Event Hubs Standard tier or higher. Basic has no Kafka endpoint.
- One Event Hubs message = many records. Diagnostic settings wrap logs in
{"records": [...]}; the component emits one Loki line per record and sets__azure_event_hubs_category. - Strict schema check. A record needs
timeortimeStamp,category,resourceIdandoperationName(field names match case-insensitively). Otherwise it becomes a raw line without a category label. Some services use an older schema (Event Hubs:EventTimeString, notime, nooperationName) — takecategoryfrom the JSON inloki.process. - Derive labels from
resourceId(resource_type,resource) — low cardinality, and the same names work across categories. - Set
levelat ingest. Azure severities (LOG, klogI/W/E, HTTP status in audit events) are not names Loki recognises, sodetected_levelbecomesunknown. Alevelstructured-metadata field fixes Explore, Logs Drilldown and Grafana MCP at once. - Choose categories by volume.
- Storage account
StorageRead/StorageWritelogs for a Loki/Tempo/Mimir backend flood the hub. kube-auditis the main AKS cost driver;kube-audit-adminexcludesgetandlistevents.- Of the categories enabled here,
guard(Entra ID authorization checks) is the noisiest.
- Storage account
- Know what audit events contain. The AKS audit policy logs request and response bodies (
RequestResponse) for writes on core API groups. Secrets, ConfigMaps, ServiceAccounts and TokenReviews are logged atMetadatalevel — name and user, no body. Filter inloki.processwhat must not reach Loki. - Never send a namespace’s logs into itself. Microsoft: sending logs “from a resource to the same resource would generate an infinite loop of generating and writing data”. The portal blocks it; ARM, Bicep and the CLI do not. Send them to a second namespace (below) and monitor the hub with metrics (next section).
Monitoring the hub itself
The hub cannot report on itself through itself, and Standard tier has no consumer-lag metric (ConsumerLag exists only in ApplicationMetricsLogs, Premium and Dedicated). Four independent views:
| View | Source | Answers |
|---|---|---|
| Platform metrics | prometheus.exporter.azure on Microsoft.EventHub/namespaces, split by EntityName |
Traffic in/out, server and user errors, throttling, quota, connections, retained size |
| Throughput-unit utilization | IncomingBytes/60 ÷ (1 MB/s × TU), IncomingMessages/60 ÷ (1000/s × TU), OutgoingBytes/60 ÷ (2 MB/s × TU); highest of the three |
How close the namespace is to throttling. The TU count is not a metric — it comes from the SKU (a dashboard variable here). Per-minute averages hide sub-second bursts: act at ~70% |
| Backlog estimate | sum_over_time(IncomingMessages[30m]) − sum_over_time(OutgoingMessages[30m]) per hub |
A stand-in for consumer lag: around 0 = the consumer keeps up, growing = it falls behind or is down |
| Consumer metrics | Alloy’s own /metrics (loki_write_sent_entries_total, loki_write_dropped_entries_total, loki_process_dropped_lines_total) |
Did the consumer write what it read, and what did it drop on purpose |
| End-to-end arrival | count_over_time({source="azure"}[5m]) by category in Loki |
Does each diagnostic setting still deliver? One category stopping while others continue points at that setting, not at the hub |
Outside the cluster:
- Activity log — who changed or deleted the namespace, hub or a diagnostic setting (management-plane operations). Collected automatically; export it with a subscription diagnostic setting.
- Resource Health / Service Health — Azure-side incidents for the namespace and region.
- The hub’s own logs — through a second namespace, see below.
- Kafka consumer-group tooling — Kafka clients and exporters can read committed offsets of the consumer group through the Kafka endpoint (Event Hubs supports only part of the Kafka admin API — test before relying on it).
The hub’s own logs
The namespace that carries every other resource’s logs cannot carry its own (infinite loop). Yet its operational logs, Kafka coordinator logs (consumer group joins, rebalances) and Kafka user error logs (failed client API calls) are what explains a failure: metrics say that a client failed or a consumer group rebalanced, not which or why.
Every option sends the logs somewhere else. The chain ends at the second destination: it has no diagnostic setting of its own and is monitored by metrics only — metrics are pulled from Azure Monitor and need no hub, so nothing recurses.
| Option | How | Cost / effort | Trade-off |
|---|---|---|---|
| Second Event Hubs namespace (“monitoring hub”) | Diagnostic setting on the main namespace → event hub in a second Standard namespace; a second loki.source.azure_event_hubs block (or event_hubs entry) reads it |
One more namespace and TU; one more role assignment and Kafka endpoint | Same pipeline, same labels, logs in Loki. The second namespace’s own logs are not collected — metrics only |
| Log Analytics workspace | Diagnostic setting → workspace (resource-specific tables AZMSOperationalLogs, AZMSKafkaCoordinatorLogs, AZMSKafkaUserErrorLogs, AZMSDiagnosticErrorLogs); Grafana Azure Monitor data source queries them with KQL |
No collector; pay per GB ingested and retained | Logs stay in Azure, not in Loki: no LogQL, no correlation with the rest in one query, not visible to Grafana MCP through Loki |
| Storage account (archive) | Diagnostic setting → blob container | Cheapest | Neither Alloy nor Vector has an Azure Blob source — audit and post-mortem only, nothing live |
| Same namespace, another event hub | Diagnostic setting → a second hub in the same namespace | Free | Still the namespace writing into itself — the loop Microsoft warns about; the portal blocks it |
| Do nothing | Rely on metrics (UserErrors, ThrottledRequests, ConnectionsClosed) |
— | Detects that something fails, not what |
Our stack: a second namespace (<ns>-logs, Standard, 1 TU, one event hub, 1 partition). The main namespace’s diagnostic setting writes into it; a second loki.source.azure_event_hubs block in the same alloy-azure release reads it with the same identity (Data Receiver on both namespaces). The logs land in Loki under resource="<main namespace>", next to every other Azure log. Log Analytics is the alternative when the logs are rarely needed and a KQL query from Grafana is enough.
Failure modes
| Symptom | Cause | Fix |
|---|---|---|
Pod restarts right after deploy; log: could not perform the initial load successfully |
The Kafka client is created when the component starts. If the login fails (role assignment not propagated yet, wrong FQDN), Alloy exits — and the metrics path stops with it | Wait: role assignments take minutes to reach the Event Hubs data plane. Split metrics and logs into two releases if one must not take the other down |
ParserFailure in Alloy logs from Resource Graph |
Leading \| in resource_graph_query_filter |
Start the filter with where |
Metrics present, label dimension everywhere |
Exporter’s naming convention | prometheus.relabel per metric family |
| Series appear and disappear | Azure reports nothing for idle minutes (IOPS, scheduling attempts) | spanNulls: false (lines or bars) so gaps stay gaps; or vector(0) for counts |
| Logs stop, metrics fine | Diagnostic setting removed, hub throttled, consumer group stuck | Event Hubs IncomingMessages vs OutgoingMessages, Alloy loki_write_sent_entries_total |
| Lines lost although Event Hubs delivered them | The Kafka offset is committed once the entry is handed to Alloy’s pipeline, not when Loki accepts it. A short Loki outage only blocks (backpressure); loss happens when loki.write gives up retrying or the process dies with entries in memory |
Hub retention (1 day here) is the replay window: reset the consumer group’s offset |
| First start replays a day of logs | A new consumer group starts at the oldest offset | Expected once; old lines keep their original timestamps (use_incoming_timestamp) |
| Template stage fails to load | Helm tpl renders in the Alloy config | Wrap Alloy templates in a Helm raw string: }} ` |
Alloy vs azure-metrics-exporter vs Vector
prometheus.exporter.azure uses the webdevops azure-metrics-exporter as a library — same prober, but Alloy’s own handler and defaults. Switching between them is not drop-in: dashboards and alerts change. Vector has no Azure source at all.
Alloy (alloy-azure) |
webdevops azure-metrics-exporter | Vector | |
|---|---|---|---|
| Azure Monitor metrics | prometheus.exporter.azure |
✅ (the original) | ❌ — needs the exporter + prometheus_scrape |
| Defaults | Name per metric (azure_{type}_{metric}_{aggregation}_{unit}), timespan PT5M, validate_dimensions off |
One azurerm_resource_metric with labels, timespan PT1M, validation on |
n/a |
| Event Hubs logs | loki.source.azure_event_hubs, Azure record parsing built in |
❌ metrics only | Generic kafka source; unpack records[] in VRL yourself |
| Auth | Workload identity for both paths | Workload identity (Azure SDK chain) | Event Hubs takes SASL PLAIN with a SAS connection string. Since librdkafka 2.12 (Vector 0.58), OAUTHBEARER via librdkafka_options with azure_imds = node-level managed identity, not per-pod workload identity |
| Engine version | v1.16–1.17 embed a July 2023 commit; v1.18–1.20 a July 2025 commit | Latest release 25.12.0 (Dec 2025) | n/a |
| Features | Resource Graph discovery, subscription scope (regions) |
+ /probe/metrics/resource, /list, tag-driven /scrape, dimension lowercasing |
n/a |
| Caching | None — every scrape queries Azure | Service discovery cache (default 30 min), optional metric cache | n/a |
| Config change | New component block → config reload | New Prometheus scrape job (config in URL parameters) | — |
| Transform language | loki.process stages + Go templates (sprig) |
— | VRL: compiled, type-checked |
| Delivery | Offset committed before Loki accepts; loki.write WAL experimental |
— | End-to-end acknowledgements, disk buffers |
| Blast radius | Metrics + logs in one process | Metrics only | Logs only (metrics elsewhere) |
| Project | Grafana Labs, Apache 2.0, 1.x with a backward-compatibility promise for GA components | Essentially one maintainer, MIT, two releases in 2025 | Datadog, MPL 2.0, still 0.x |
Alloy — pros
- One binary for both paths. No extra Deployment, Service and ServiceMonitor for the exporter; same config language, UI and upgrade cadence as the rest of the collector fleet.
- Secretless Event Hubs with per-pod identity.
mechanism = "oauth"+ workload identity; nothing to rotate. - Azure log format handled.
records[]unpacking, timestamps and category without custom code. - Supported, GA components. Both Azure components are generally available in Alloy.
Alloy — cons
- Behind upstream and only partly exposed. Even v1.20 pins a July 2025 commit (upstream 25.12.0). No caching, no tag-driven or per-resource probe endpoints.
- Shared crash domain. An Event Hubs failure at startup stops metrics too; a panic in the embedded library stops the whole process.
- Weaker log delivery than Vector (offset before Loki accepts, no production disk buffer for
loki.write). - Limited transforms. Mixed Azure schemas need
stage.match+ templates where VRL would be one function. - No
metrics:getBatch. Neither Alloy (issue #4651, open) nor the webdevops exporter uses the batch API; Promitor and the OTel Collectorazuremonitorreceiver(use_batch_api) have it as an experimental option.
Decision table
| Situation | Choose |
|---|---|
| Grafana stack, moderate volume, workload identity required | Alloy for both paths |
| Hundreds of resources, API throttling | regions (subscription scope) in Alloy or the exporter; beyond that a batch-API collector (Promitor, OTel azuremonitorreceiver, both experimental) |
| Scrape more often than the Azure grain without more API calls | Standalone azure-metrics-exporter with its metric cache |
| Metrics must survive a log-pipeline failure | Two Alloy releases (metrics / logs), or exporter + Alloy |
| Heavy log reshaping, many destinations (archive to blob, SIEM), durable delivery | Vector for logs (SAS connection string or node-level managed identity) + exporter for metrics |
| Only dashboards needed, no copy of the data | No collector: Grafana Azure Monitor data source queries Azure Monitor metrics, Log Analytics and Resource Graph directly |
In our stack
| Piece | File |
|---|---|
Event Hubs (main + -logs namespace), managed identity, federated credential, roles, diagnostic settings |
infrastructure/azure-monitoring.bicep |
| Workload identity on the cluster | infrastructure/aks-cluster.bicep |
Deploy module (step 5b of deploy/run.sh) |
run_deploy_scripts/deploy-azure-monitoring-module.sh |
The alloy-azure release |
helm_values/alloy/alloy-azure.values.yaml |
| Dashboards (Grafana folder Azure) | AKS control plane, Event Hubs, PostgreSQL |
| Post-deploy check | tests/test_azure_monitoring.sh |
Diagnostic settings:
| Resource | Categories |
|---|---|
| AKS control plane | kube-audit-admin, kube-controller-manager, kube-scheduler, cluster-autoscaler, guard |
PostgreSQL (only with install_grafana_database=true) |
PostgreSQLLogs |
| Event Hubs namespace | OperationalLogs, KafkaCoordinatorLogs, KafkaUserErrorLogs, DiagnosticErrorLogs → the second namespace <ns>-logs (why) |
Second namespace <ns>-logs |
none — the chain ends here; metrics only |
| Storage account | none (metrics only) |
loki.process drops every kube-audit-admin line that mentions flagd (the incident exercise must not be solvable by reading the audit log) and every lease renewal (coordination.k8s.io, pure noise).