Collecting Azure Metrics and Logs

Managed services have no node to put an agent on. The AKS control plane, a PostgreSQL Flexible Server and a storage account report only to Azure Monitor.

Collecting them means two different paths: pull metrics from the Azure Monitor API, and push logs out of Azure through Event Hubs.

What Azure exposes

Signal Source in Azure How you get it out Latency
Platform metrics Azure Monitor, per resource; 1-minute minimum grain (some metrics every 20 min or hourly) Pull: Azure Monitor metrics REST API Under a minute in Azure, plus the scrape interval
Resource logs Per-resource categories (e.g. kube-audit-admin, PostgreSQLLogs) Push: a diagnostic setting per resource streams them to Log Analytics, a storage account, Event Hubs or a partner Usually 3–10 min; a new diagnostic setting can take up to 90 min to deliver
Activity log Subscription-level control-plane operations Collected automatically; a diagnostic setting on the subscription exports it 3–20 min
  • Metrics are not pushed. Something has to poll the API. Every poll is an ARM API call that counts against throttling limits.
  • Resource logs are not collected by default. Without a diagnostic setting they are not stored anywhere. Event Hubs is the only destination a collector can read as a stream.
  • Idle often means NULL, not 0. Each resource provider decides whether an idle minute is 0 or no value, and a dimension value has no series until it had data. Error and throttling series often appear only with the first error — query with or vector(0).

The pipeline

flowchart LR
    subgraph az ["Azure"]
        res["AKS · PostgreSQL"]
        sto["Storage account<br/>(metrics only)"]
        mon["Azure Monitor metrics API"]
        diag["Diagnostic settings"]
        eh[("Event Hubs<br/>main namespace")]
        eh2[("Event Hubs<br/>-logs namespace<br/>(metrics only)")]
        res --> mon
        sto --> mon
        eh --> mon
        eh2 --> mon
        res --> diag --> eh
        eh -- "its own logs" --> eh2
    end
    subgraph k8s ["AKS — alloy-azure (1 replica, workload identity)"]
        exp["prometheus.exporter.azure"] --> scr["prometheus.scrape"]
        src["loki.source.azure_event_hubs ×2"] --> proc["loki.process"]
    end
    mon -- "pull, every 60 s" --> exp
    eh -- "consume (Kafka, OAuth)" --> src
    eh2 -- "consume (Kafka, OAuth)" --> src
    scr --> prom[(Prometheus)]
    proc --> loki[(Loki)]

Authentication: workload identity

No secret in the cluster. The pod proves its identity with a Kubernetes token; Entra ID trusts that token through a federated credential.

  1. Pod label azure.workload.identity/use: "true" + ServiceAccount annotation azure.workload.identity/client-id → the AKS webhook injects AZURE_CLIENT_ID, AZURE_TENANT_ID, AZURE_AUTHORITY_HOST, AZURE_FEDERATED_TOKEN_FILE and mounts a projected ServiceAccount token.
  2. Both components use the Azure SDK DefaultAzureCredential, which exchanges that token for Entra ID tokens of the managed identity — one for ARM (metrics), one for https://<ns>.servicebus.windows.net/.default (Event Hubs).
  3. Entra ID accepts it because of a federated credential: issuer = the cluster’s OIDC issuer URL, subject = system:serviceaccount:monitoring:alloy-azure, audience = api://AzureADTokenExchange.
  4. Event Hubs receives the token over Kafka SASL_SSL + OAUTHBEARER; Azure Monitor receives it as a bearer token.
Role Scope Needed for
Monitoring Reader Resource group Metrics API + Resource Graph discovery
Azure Event Hubs Data Receiver Event Hubs namespace Reading the hub (not sending, not managing)

The only key involved is the namespace’s SAS rule (Manage, Send, Listen) that diagnostic settings use to write into the hub. Alloy never sees it.

Metrics: prometheus.exporter.azure

Excerpt — one of ten exporter blocks; the real config has more metrics and concatenates all exporter targets:

prometheus.exporter.azure "aks_nodes" {
  subscriptions               = [sys.env("AZURE_SUBSCRIPTION_ID")]
  resource_type               = "Microsoft.ContainerService/managedClusters"
  resource_graph_query_filter = string.format("where resourceGroup =~ '%s'", sys.env("AZURE_RESOURCE_GROUP"))
  metrics                     = ["node_cpu_usage_percentage", "node_memory_working_set_percentage"]
  included_dimensions         = ["node"]
}

prometheus.scrape "azure" {
  targets         = prometheus.exporter.azure.aks_nodes.targets   // array.concat(...) of all blocks
  job_name        = "integrations/azure"
  scrape_interval = "60s"   // matches the 1-minute grain; each scrape = API calls
  scrape_timeout  = "50s"
  forward_to      = [prometheus.relabel.azure_dimensions.receiver]
}
Behaviour Consequence
Discovery via Azure Resource Graph, on every scrape Query = Resources \| where type =~ "<type>" \| <filter> \| project id, tags. The exporter adds the \| itself — write where …; a leading \| gives a parser error
regions instead of a filter Subscription-scope mode: one metrics call per subscription and region instead of one per resource — the answer to throttling with many resources. Mutually exclusive with resource_graph_query_filter, so it cannot be limited to one resource group
included_dimensions becomes one $filter for every metric in the block With validate_dimensions = false (default) Azure ignores dimensions a metric does not have. We still use one block per dimension set (AKS: 6 blocks) so each metric family has one label shape
Generic dimension labels One dimension → label dimension; several → dimensionNode, dimensionCondition, … Rename them in prometheus.relabel (node, request_kind, entity)
metric_aggregations not set Each metric uses its primary aggregation; the name says which: azure_<type>_<metric>_<aggregation>_<unit>
“Total” metrics are per-minute values, not counters The exporter queries a 5-minute timespan and exposes the last non-null minute. sum_over_time(x[$__range]) is an approximate total — an idle or late minute can be exposed on two scrapes. Never rate()
No caching Every scrape runs the Resource Graph query plus the metric calls (one call per resource per 20 metrics). More blocks or a shorter interval = more API calls

Names get long: azure_microsoft_containerservice_managedclusters_apiserver_cpu_usage_percentage_average_percent. That is the price of a name that says type, metric, aggregation and unit.

Logs: loki.source.azure_event_hubs

loki.source.azure_event_hubs "resource_logs" {
  fully_qualified_namespace = sys.env("EVENTHUB_KAFKA_ENDPOINT")   // <ns>.servicebus.windows.net:9093
  event_hubs                = [sys.env("EVENTHUB_NAME")]
  group_id                  = "alloy-azure"
  use_incoming_timestamp    = true
  labels                    = { job = "integrations/azure_event_hubs", source = "azure" }
  relabel_rules             = loki.relabel.azure.rules   // __azure_event_hubs_category → category
  authentication { mechanism = "oauth" }
  forward_to = [loki.process.azure.receiver]
}

All Azure lines are under {source="azure"}, with category, resource and resource_type labels.

  • Event Hubs Standard tier or higher. Basic has no Kafka endpoint.
  • One Event Hubs message = many records. Diagnostic settings wrap logs in {"records": [...]}; the component emits one Loki line per record and sets __azure_event_hubs_category.
  • Strict schema check. A record needs time or timeStamp, category, resourceId and operationName (field names match case-insensitively). Otherwise it becomes a raw line without a category label. Some services use an older schema (Event Hubs: EventTimeString, no time, no operationName) — take category from the JSON in loki.process.
  • Derive labels from resourceId (resource_type, resource) — low cardinality, and the same names work across categories.
  • Set level at ingest. Azure severities (LOG, klog I/W/E, HTTP status in audit events) are not names Loki recognises, so detected_level becomes unknown. A level structured-metadata field fixes Explore, Logs Drilldown and Grafana MCP at once.
  • Choose categories by volume.
    • Storage account StorageRead/StorageWrite logs for a Loki/Tempo/Mimir backend flood the hub.
    • kube-audit is the main AKS cost driver; kube-audit-admin excludes get and list events.
    • Of the categories enabled here, guard (Entra ID authorization checks) is the noisiest.
  • Know what audit events contain. The AKS audit policy logs request and response bodies (RequestResponse) for writes on core API groups. Secrets, ConfigMaps, ServiceAccounts and TokenReviews are logged at Metadata level — name and user, no body. Filter in loki.process what must not reach Loki.
  • Never send a namespace’s logs into itself. Microsoft: sending logs “from a resource to the same resource would generate an infinite loop of generating and writing data”. The portal blocks it; ARM, Bicep and the CLI do not. Send them to a second namespace (below) and monitor the hub with metrics (next section).

Monitoring the hub itself

The hub cannot report on itself through itself, and Standard tier has no consumer-lag metric (ConsumerLag exists only in ApplicationMetricsLogs, Premium and Dedicated). Four independent views:

View Source Answers
Platform metrics prometheus.exporter.azure on Microsoft.EventHub/namespaces, split by EntityName Traffic in/out, server and user errors, throttling, quota, connections, retained size
Throughput-unit utilization IncomingBytes/60 ÷ (1 MB/s × TU), IncomingMessages/60 ÷ (1000/s × TU), OutgoingBytes/60 ÷ (2 MB/s × TU); highest of the three How close the namespace is to throttling. The TU count is not a metric — it comes from the SKU (a dashboard variable here). Per-minute averages hide sub-second bursts: act at ~70%
Backlog estimate sum_over_time(IncomingMessages[30m]) − sum_over_time(OutgoingMessages[30m]) per hub A stand-in for consumer lag: around 0 = the consumer keeps up, growing = it falls behind or is down
Consumer metrics Alloy’s own /metrics (loki_write_sent_entries_total, loki_write_dropped_entries_total, loki_process_dropped_lines_total) Did the consumer write what it read, and what did it drop on purpose
End-to-end arrival count_over_time({source="azure"}[5m]) by category in Loki Does each diagnostic setting still deliver? One category stopping while others continue points at that setting, not at the hub

Outside the cluster:

  • Activity log — who changed or deleted the namespace, hub or a diagnostic setting (management-plane operations). Collected automatically; export it with a subscription diagnostic setting.
  • Resource Health / Service Health — Azure-side incidents for the namespace and region.
  • The hub’s own logs — through a second namespace, see below.
  • Kafka consumer-group tooling — Kafka clients and exporters can read committed offsets of the consumer group through the Kafka endpoint (Event Hubs supports only part of the Kafka admin API — test before relying on it).

The hub’s own logs

The namespace that carries every other resource’s logs cannot carry its own (infinite loop). Yet its operational logs, Kafka coordinator logs (consumer group joins, rebalances) and Kafka user error logs (failed client API calls) are what explains a failure: metrics say that a client failed or a consumer group rebalanced, not which or why.

Every option sends the logs somewhere else. The chain ends at the second destination: it has no diagnostic setting of its own and is monitored by metrics only — metrics are pulled from Azure Monitor and need no hub, so nothing recurses.

Option How Cost / effort Trade-off
Second Event Hubs namespace (“monitoring hub”) Diagnostic setting on the main namespace → event hub in a second Standard namespace; a second loki.source.azure_event_hubs block (or event_hubs entry) reads it One more namespace and TU; one more role assignment and Kafka endpoint Same pipeline, same labels, logs in Loki. The second namespace’s own logs are not collected — metrics only
Log Analytics workspace Diagnostic setting → workspace (resource-specific tables AZMSOperationalLogs, AZMSKafkaCoordinatorLogs, AZMSKafkaUserErrorLogs, AZMSDiagnosticErrorLogs); Grafana Azure Monitor data source queries them with KQL No collector; pay per GB ingested and retained Logs stay in Azure, not in Loki: no LogQL, no correlation with the rest in one query, not visible to Grafana MCP through Loki
Storage account (archive) Diagnostic setting → blob container Cheapest Neither Alloy nor Vector has an Azure Blob source — audit and post-mortem only, nothing live
Same namespace, another event hub Diagnostic setting → a second hub in the same namespace Free Still the namespace writing into itself — the loop Microsoft warns about; the portal blocks it
Do nothing Rely on metrics (UserErrors, ThrottledRequests, ConnectionsClosed) — Detects that something fails, not what

Our stack: a second namespace (<ns>-logs, Standard, 1 TU, one event hub, 1 partition). The main namespace’s diagnostic setting writes into it; a second loki.source.azure_event_hubs block in the same alloy-azure release reads it with the same identity (Data Receiver on both namespaces). The logs land in Loki under resource="<main namespace>", next to every other Azure log. Log Analytics is the alternative when the logs are rarely needed and a KQL query from Grafana is enough.

Failure modes

Symptom Cause Fix
Pod restarts right after deploy; log: could not perform the initial load successfully The Kafka client is created when the component starts. If the login fails (role assignment not propagated yet, wrong FQDN), Alloy exits — and the metrics path stops with it Wait: role assignments take minutes to reach the Event Hubs data plane. Split metrics and logs into two releases if one must not take the other down
ParserFailure in Alloy logs from Resource Graph Leading \| in resource_graph_query_filter Start the filter with where
Metrics present, label dimension everywhere Exporter’s naming convention prometheus.relabel per metric family
Series appear and disappear Azure reports nothing for idle minutes (IOPS, scheduling attempts) spanNulls: false (lines or bars) so gaps stay gaps; or vector(0) for counts
Logs stop, metrics fine Diagnostic setting removed, hub throttled, consumer group stuck Event Hubs IncomingMessages vs OutgoingMessages, Alloy loki_write_sent_entries_total
Lines lost although Event Hubs delivered them The Kafka offset is committed once the entry is handed to Alloy’s pipeline, not when Loki accepts it. A short Loki outage only blocks (backpressure); loss happens when loki.write gives up retrying or the process dies with entries in memory Hub retention (1 day here) is the replay window: reset the consumer group’s offset
First start replays a day of logs A new consumer group starts at the oldest offset Expected once; old lines keep their original timestamps (use_incoming_timestamp)
Template stage fails to load Helm tpl renders in the Alloy config | Wrap Alloy templates in a Helm raw string: }} `  

Alloy vs azure-metrics-exporter vs Vector

prometheus.exporter.azure uses the webdevops azure-metrics-exporter as a library — same prober, but Alloy’s own handler and defaults. Switching between them is not drop-in: dashboards and alerts change. Vector has no Azure source at all.

  Alloy (alloy-azure) webdevops azure-metrics-exporter Vector
Azure Monitor metrics prometheus.exporter.azure ✅ (the original) ❌ — needs the exporter + prometheus_scrape
Defaults Name per metric (azure_{type}_{metric}_{aggregation}_{unit}), timespan PT5M, validate_dimensions off One azurerm_resource_metric with labels, timespan PT1M, validation on n/a
Event Hubs logs loki.source.azure_event_hubs, Azure record parsing built in ❌ metrics only Generic kafka source; unpack records[] in VRL yourself
Auth Workload identity for both paths Workload identity (Azure SDK chain) Event Hubs takes SASL PLAIN with a SAS connection string. Since librdkafka 2.12 (Vector 0.58), OAUTHBEARER via librdkafka_options with azure_imds = node-level managed identity, not per-pod workload identity
Engine version v1.16–1.17 embed a July 2023 commit; v1.18–1.20 a July 2025 commit Latest release 25.12.0 (Dec 2025) n/a
Features Resource Graph discovery, subscription scope (regions) + /probe/metrics/resource, /list, tag-driven /scrape, dimension lowercasing n/a
Caching None — every scrape queries Azure Service discovery cache (default 30 min), optional metric cache n/a
Config change New component block → config reload New Prometheus scrape job (config in URL parameters) —
Transform language loki.process stages + Go templates (sprig) — VRL: compiled, type-checked
Delivery Offset committed before Loki accepts; loki.write WAL experimental — End-to-end acknowledgements, disk buffers
Blast radius Metrics + logs in one process Metrics only Logs only (metrics elsewhere)
Project Grafana Labs, Apache 2.0, 1.x with a backward-compatibility promise for GA components Essentially one maintainer, MIT, two releases in 2025 Datadog, MPL 2.0, still 0.x

Alloy — pros

  • One binary for both paths. No extra Deployment, Service and ServiceMonitor for the exporter; same config language, UI and upgrade cadence as the rest of the collector fleet.
  • Secretless Event Hubs with per-pod identity. mechanism = "oauth" + workload identity; nothing to rotate.
  • Azure log format handled. records[] unpacking, timestamps and category without custom code.
  • Supported, GA components. Both Azure components are generally available in Alloy.

Alloy — cons

  • Behind upstream and only partly exposed. Even v1.20 pins a July 2025 commit (upstream 25.12.0). No caching, no tag-driven or per-resource probe endpoints.
  • Shared crash domain. An Event Hubs failure at startup stops metrics too; a panic in the embedded library stops the whole process.
  • Weaker log delivery than Vector (offset before Loki accepts, no production disk buffer for loki.write).
  • Limited transforms. Mixed Azure schemas need stage.match + templates where VRL would be one function.
  • No metrics:getBatch. Neither Alloy (issue #4651, open) nor the webdevops exporter uses the batch API; Promitor and the OTel Collector azuremonitorreceiver (use_batch_api) have it as an experimental option.

Decision table

Situation Choose
Grafana stack, moderate volume, workload identity required Alloy for both paths
Hundreds of resources, API throttling regions (subscription scope) in Alloy or the exporter; beyond that a batch-API collector (Promitor, OTel azuremonitorreceiver, both experimental)
Scrape more often than the Azure grain without more API calls Standalone azure-metrics-exporter with its metric cache
Metrics must survive a log-pipeline failure Two Alloy releases (metrics / logs), or exporter + Alloy
Heavy log reshaping, many destinations (archive to blob, SIEM), durable delivery Vector for logs (SAS connection string or node-level managed identity) + exporter for metrics
Only dashboards needed, no copy of the data No collector: Grafana Azure Monitor data source queries Azure Monitor metrics, Log Analytics and Resource Graph directly

In our stack

Piece File
Event Hubs (main + -logs namespace), managed identity, federated credential, roles, diagnostic settings infrastructure/azure-monitoring.bicep
Workload identity on the cluster infrastructure/aks-cluster.bicep
Deploy module (step 5b of deploy/run.sh) run_deploy_scripts/deploy-azure-monitoring-module.sh
The alloy-azure release helm_values/alloy/alloy-azure.values.yaml
Dashboards (Grafana folder Azure) AKS control plane, Event Hubs, PostgreSQL
Post-deploy check tests/test_azure_monitoring.sh

Diagnostic settings:

Resource Categories
AKS control plane kube-audit-admin, kube-controller-manager, kube-scheduler, cluster-autoscaler, guard
PostgreSQL (only with install_grafana_database=true) PostgreSQLLogs
Event Hubs namespace OperationalLogs, KafkaCoordinatorLogs, KafkaUserErrorLogs, DiagnosticErrorLogs → the second namespace <ns>-logs (why)
Second namespace <ns>-logs none — the chain ends here; metrics only
Storage account none (metrics only)

loki.process drops every kube-audit-admin line that mentions flagd (the incident exercise must not be solvable by reading the audit log) and every lease renewal (coordination.k8s.io, pure noise).

results matching ""

    No results matching ""