Compare · updated 2026-08-24
Observability data ownership: vendor-hosted vs a lake on Iceberg
Vendor-hosted telemetry is retained on the vendor's terms and read through the vendor's interfaces, while a lake on open table formats holds full-fidelity history that every engine and every model you ever run can read.
01The mechanics
What actually differs
The unit each option charges on, and who owns what afterwards, matter more than any single quoted figure.
Every platform hosts your telemetry as part of the service, and every platform prices retention so that keeping everything is irrational. The result is an estate whose history is scattered across vendor retention windows, none of it complete, all of it readable only through the interfaces each vendor provides, exportable on each vendor's terms.
The alternative is structural: full-fidelity history in your own object storage, enriched at write time, partitioned the way investigations move, on open table formats with Iceberg as the market's direction. Retention becomes a storage-class decision instead of a licensing negotiation, replay into any destination is a query, and the history is readable by whatever query engine or model comes next. Sourced cells are being assembled row by row.
02The honest row
When each is the right answer
Every option in this comparison is the right answer for somebody, and saying when is the part most comparisons leave out.
Vendor-hosted
The right answer for the hot window: the recent data your detections and dashboards actively read, where the platform's query speed and integrations earn the hosting. The mistake is not hosting data there, it is letting that window define what you keep.
Lake on Iceberg
The right answer for everything else: complete history at object-storage cost, compliance evidence on demand, replay into any tool, and the foundation every AI ambition stands on. Not a replacement for the platforms, the ground beneath them.
03Questions
Asked about this comparison
Is a telemetry lake a SIEM replacement?
No, and it should never alert. Detection stays on the SIEM, which is built for it. The lake holds the full-fidelity history the SIEM should not have to, serves audits and investigations, and feeds the agents. The two are complements priced for different jobs.
Why do open table formats matter for AI?
Because the model you run next year does not exist yet, and neither does its vendor. History held in a format one vendor controls is readable on that vendor's roadmap. History on open table formats is readable by every engine and every model, which is what makes it a foundation rather than an archive.
Model it against your estate
The comparison that matters is the one run on your volumes and your contracts. The review reads them with you, and you leave with your version of the Logmetry Blueprint. No system access, no obligation.