Measure Once, Distribute Everywhere: A Central Telemetry Engine for GenAI
How to handle high-cardinality GenAI traces by ingesting once with OpenTelemetry and forking into APM, billing, and ClickHouse.
Monitoring traditional web microservices is straightforward.
You track HTTP status codes, request latency, CPU usage, and database query durations. Spans are small, metrics are simple counters or gauges, and logs are short strings.
Generative AI applications break this model completely.
A single agent interaction can generate dozens of complex, verbose events:
- Multi-step reasoning chains, intermediate reflections, and planning thoughts.
- Complete prompt text, system messages, and generated completions.
- Tool call inputs and multi-megabyte JSON outputs.
- Granular token counts: prompt tokens, cached prompt tokens, reasoning tokens, and completion tokens.
- Financial cost attribution broken down by tenant, business unit, and project.
When you scale an enterprise platform to hundreds or thousands of agents, this generates between 100 GB and 1 TB of verbose, semi-structured telemetry every day.
If you dump all of this data directly into a primary Application Performance Monitoring (APM) tool, your monthly observability bill will quickly rival your AI compute bill.
To solve this, we designed the “Measure Once, Distribute Everywhere” telemetry architecture.
The Idea: The Telemetry Fork
Instead of building separate logging pipelines for developers, finance teams, and operations engineers, we capture a single, unified telemetry stream using OpenTelemetry (OTEL).
Our collector pipeline ingests each event once, validates its schema, and “forks” it into three purpose-built destinations:
- Real-Time Operational APM: Dedicated to service health, uptime, and alerting. The collector strips large prompt and response bodies, forwarding only operational metrics: Time To First Token (TTFT), Time Per Output Token (TPOT), error codes, and latency percentiles.
- Immutable Financial Ledger: Formatted as strict, append-only events. Captures exact token consumption and cost center attribution to reconcile internal chargebacks, deduplicated on unique event IDs.
- Columnar OLAP Data Store (ClickHouse): Dedicated to deep prompt inspection, full reasoning chains, and historical latency analytics. Stores the complete semi-structured payloads at high compression for long-term auditability.
How It Worked Well
- Massive APM Cost Savings: By stripping heavy text payloads before sending data to our real-time APM tool, we reduced operational monitoring ingestion volume by over 85% while retaining instant alerting for outages and latency spikes.
- Sub-Second Multi-Gigabyte Analytics: Using a columnar database like ClickHouse allowed us to run fast aggregations and percentile calculations across hundreds of millions of daily rows without pre-aggregating or losing detail.
- Client-Side Trace Reassembly: By storing multi-step agent executions in a flattened trace model, developers can query by a single trace identifier and instantly reconstruct the full hierarchical reasoning tree (root prompt, subagents, tool calls, and final synthesis) in the debugging UI.
- Transparent Cost Attribution: Finance and platform teams gained automated, granular visibility into token consumption across different business divisions, enabling accurate chargeback accounting.
What to Watch Out For
- Personally Identifiable Information (PII) Leakage: Because GenAI prompts can contain sensitive user information, logs stored in your columnar data warehouse must be sanitized. Run automated redaction filters at the collector edge to scrub credentials, personal identification numbers, and payment details before data reaches long-term storage.
- High-Cardinality Indexing Pitfalls: In columnar databases, sorting or partitioning by high-cardinality columns (such as unique user IDs or trace IDs) can degrade performance. Partition data by time ranges and sort primarily by tenant and date, using sparse primary keys to keep queries fast and memory usage low.
- Tiered Cluster Telemetry: In large Kubernetes environments, running full-featured APM tracing agents inside ephemeral CI/CD runners or batch worker pods wastes observability licenses. Restrict full APM agents to customer-facing workload clusters, routing ephemeral batch runners to basic cloud logs instead.
- Collector Buffer Saturation: During massive burst workloads, collector pipelines can drop events if internal memory buffers fill up. Configure persistent file-backed buffering or queue-backed buffers between your applications and your telemetry collectors to guarantee zero data loss.