Browser Telemetry

The backend trace starts when the request reaches your ingress. The user’s experience started earlier, in a browser you do not control.

Browser telemetry closes that gap. The price: you accept data from a client that is untrusted, unreliable and has a wrong clock.

🌐 Why the browser is different

Concern Backend service Browser
Environment Your container, your runtime The user’s device, extensions, ad-blockers, flaky network
Transport OTLP/gRPC or HTTP on a private network OTLP/HTTP only, over the public internet, subject to CORS
Lifetime Long-running process, graceful shutdown The tab can close at any moment; an unsent batch is lost
Clock NTP-synced Whatever the device says
Trust Authenticated workload Anyone can POST to the endpoint; no secret survives in a JS bundle
Volume Scales with requests Scales with users × page views × interactions

📡 What to collect

Data Source Signal Notes
Page load Navigation and Resource Timing APIs Span per load, child spans per resource DNS, TLS, TTFB, DOM timings
Outgoing requests fetch / XMLHttpRequest hooks Client span + traceparent header The link to the backend trace
User interactions DOM event listeners Span per click/submit Root of the trace that follows
Errors window.onerror, unhandledrejection Log/event with stack trace Unreadable without source maps
Core Web Vitals web-vitals library Metric or event LCP, INP, CLS. INP replaced FID in March 2024
Session SDK-generated ID session.id attribute on everything Groups all telemetry of one visit

Web Vitals are per-page measurements, not operations with a start and an end. Store them as metrics or events, not spans.

🧰 Two SDK families

OpenTelemetry JS for the web

  • Packages: @opentelemetry/sdk-trace-web, @opentelemetry/context-zone, @opentelemetry/auto-instrumentations-web (document-load, fetch, xml-http-request, user-interaction).
  • Status: tracing works and is widely used; browser support is still marked experimental. Browser logs, metrics, sessions and Web Vitals conventions are work in progress in the OTel client-side SIG.
  • Output: plain OTLP/HTTP to any collector.
  • Setup: see the browser section of Traces.

The workshop’s OTel Demo frontend uses this SDK: its browser code reports as its own service (frontend-web) and adds session.id to every span with a custom span processor.

Grafana Faro

  • Packages: @grafana/faro-web-sdk (errors, Web Vitals, sessions, page views, console) and @grafana/faro-web-tracing (OTel JS underneath).
  • Output: Faro’s own protocol to Alloy’s faro.receiver, or to Grafana Cloud Frontend Observability.
  • Strength: full RUM out of the box, including source-map deobfuscation in the receiver.
import { initializeFaro, getWebInstrumentations } from '@grafana/faro-web-sdk';
import { TracingInstrumentation } from '@grafana/faro-web-tracing';

initializeFaro({
  url: 'https://shop.example.com/faro/collect',
  app: { name: 'shop-web', version: '1.24.3', environment: 'prod' },
  instrumentations: [...getWebInstrumentations(), new TracingInstrumentation()],
});

The receiver splits the payload: logs, errors, events and measurements go to Loki; traces go to Tempo.

faro.receiver "web" {
  server {
    listen_address           = "0.0.0.0"
    listen_port              = 12347
    cors_allowed_origins     = ["https://shop.example.com"]
    max_allowed_payload_size = "5MiB"

    rate_limiting {
      rate       = 50
      burst_size = 100
    }
  }

  sourcemaps {
    download = true   // fetch *.map next to the minified bundle
  }

  output {
    logs   = [loki.write.default.receiver]
    traces = [otelcol.exporter.otlp.tempo.input]
  }
}
  OpenTelemetry JS Grafana Faro
Vendor neutrality Full, OTLP to anything Faro protocol, Grafana-oriented backends
Traces ✅ ✅ (via OTel JS)
Errors, Web Vitals, sessions Build it yourself ✅ Out of the box
Source maps Your pipeline ✅ In faro.receiver
Maturity in the browser Experimental Stable, production RUM

🔗 Joining browser and backend traces

Outgoing requests. The fetch/XHR instrumentation injects traceparent automatically for same-origin calls. For other origins, list them in propagateTraceHeaderCorsUrls, and the backend must allow the headers in CORS:

Access-Control-Allow-Headers: traceparent, tracestate

Without both, the browser span and the backend span land in two separate traces.

The first HTML request. The browser fetches the document before any JavaScript runs, so nothing can inject a header. The reverse works: the server renders its own trace context into the page, and the document-load instrumentation uses it as the parent:

<meta name="traceparent" content="00-4bf92f3577b34da6a3ce929d0e0e4736-00f067aa0ba902b7-01">

Sessions. session.id on every browser span lets you pull every trace of one visit: “show me what this user did before the checkout failed”.

🚪 Ingest from the public internet

Protocol. Browsers cannot speak gRPC. Expose OTLP/HTTP (/v1/traces, /v1/logs, /v1/metrics, port 4318), protobuf or JSON.

CORS on the receiver (OTel Collector):

receivers:
  otlp:
    protocols:
      http:
        endpoint: 0.0.0.0:4318
        cors:
          allowed_origins: ["https://shop.example.com"]
          max_age: 7200

Alloy has the same block: otelcol.receiver.otlp → http { cors { allowed_origins = [...] } }.

Same-origin proxy. Serve the ingest path from the application’s own domain (https://shop.example.com/otlp-http/v1/traces) and route it to the collector at the edge. This avoids CORS preflights entirely, and ad-blocker lists that target well-known telemetry hostnames do not match.

The OTel Demo does exactly this: its Envoy frontend-proxy routes /otlp-http/ to the collector. In this workshop the browser exporter endpoint comes from NEXT_PUBLIC_OTEL_EXPORTER_OTLP_TRACES_ENDPOINT, which defaults to http://localhost:8080/otlp-http/v1/traces in deploy-otel-module.sh. Browser spans therefore arrive when you open the shop through kubectl port-forward on port 8080.

Hardening a public endpoint:

Risk Control
API key in the bundle Treat it as an identifier, not authentication — anyone can read it
Flood / abuse Rate limit at the ingress and in the receiver
Huge payloads Cap body size at the ingress; set SDK span and attribute limits
Spoofed resource attributes Overwrite service.namespace and similar in the collector for this pipeline
Browser flood starving backend telemetry Separate receiver, pipeline or tenant for browser data

🔒 Privacy and cardinality

  • URLs leak data. url.full carries query strings with tokens, emails and search terms. Strip them in the SDK (span processor, ignoreUrls) or in the collector (transform / redaction processor).
  • Telemetry with session.id or a user ID is personal data. Tie SDK initialization to your consent logic, or drop the identifiers.
  • Client IP. The collector sees it; do not add it as an attribute unless you need geo data.
  • Cardinality. Never turn url.full, user_agent.original or element selectors into metric labels. Normalize to routes (/product/{id}) first.

📉 Delivery, clocks and sampling

Page unload. The last batch is lost when the tab closes. Flush when the page is hidden:

document.addEventListener('visibilitychange', () => {
  if (document.visibilityState === 'hidden') provider.forceFlush();
});

Unload-time sends go through sendBeacon or fetch with keepalive, both limited to roughly 64 KiB per request. Keep batches small and the export delay short.

Clock skew. Browser timestamps come from the device clock. A backend child span can appear to start before its browser parent. Durations within one tier are reliable; offsets between tiers are not.

Sampling. Head sampling in the browser is the only way to cut egress from millions of clients. The sampled flag travels in traceparent, and a ParentBased sampler on the backend honours it. The browser’s sampling ratio therefore decides how many user-initiated backend traces are complete. Tail sampling in a gateway still works on top.

Volume estimate. 100k daily users × 20 page views × 30 spans = 60M spans per day, before any backend span.

📱 Mobile clients

OpenTelemetry Android and OpenTelemetry Swift apply the same model to native apps. The problems grow: offline periods require on-device buffering, battery and data plans limit export frequency, and old app versions stay in the field for years. The request still has to leave the device with a traceparent — the service mesh cannot add it, see Instrumenting Applications.

🚨 Common failure modes

Symptom Cause Fix
Browser and backend spans in separate traces traceparent not sent cross-origin propagateTraceHeaderCorsUrls + backend CORS Access-Control-Allow-Headers
No browser spans, CORS error on /v1/traces in the console Receiver has no CORS config cors.allowed_origins on the OTLP HTTP receiver
Spans only from some users Ad-blockers block the telemetry host Same-origin proxy path
Last click before navigation missing Batch not flushed on unload forceFlush() on visibilitychange
Stack traces show main.3f2a.js:1:48213 Minified bundle, no source maps Upload source maps / faro.receiver sourcemaps
413 from the ingress Body size limit Raise the ingress limit, shrink batches
Exporter fails to load at all gRPC exporter bundled Use the OTLP/HTTP exporter

results matching ""

    No results matching ""