What you run · A Data Lake
The data without the explanation is a haystack
A telemetry lake is object storage you own, holding logs, metrics, and traces as Parquet on open table formats like Iceberg, queried by any engine and priced at storage plus what each query scans rather than per GB ingested. It never fires an alert, because detection belongs in the real-time store, and it pays off only when the data in it is enriched, partitioned the way investigations move, and read by an agent. Until then it is cheap storage nobody opens.
01Today
How a data lake estate looks today
A lake that stores telemetry raw and unpartitioned answers no questions: every search is a scan, every investigation an export job, and the expensive tools still receive everything.
What a lake does best
If telemetry already lands in your own storage, you got the most important decision right: history in a place you control, at object-storage cost, outside any vendor's retention math. That position compounds, and it is the asset every AI ambition in this market ultimately stands on.
The catch is that raw storage answers no questions. Data written as it arrived, unenriched and unpartitioned, makes every search a scan and every investigation an export job. The lake most estates have is an archive. The Lake the Blueprint builds is a working asset: enriched in flight, partitioned the way investigations move, held on open table formats so every engine and every model can read it.
02The five verbs
What we do to a data lake estate
Keep it and govern what flows in, shrink its footprint, cut what it costs, extend it with an agent, and replace only where replacement is honest.
Your storage, your keys, your history. Nothing moves out. The lake you built is the right foundation.
What still flows to expensive destinations gets governed by the control layer, now that there is a credible cheap path beside them that keeps every byte replayable on demand.
Retrieval cost collapses when partitioning matches how queries actually move: by time, by entity, by source.
Enrichment lands context at write time, compliance replay becomes a query, and the agent phases read the Lake as their working memory.
Nothing. This page is about making an asset you already own start paying.
The letters mark where each verb acts in the drawing above
03The approach
From archive to working asset
The difference between an archive and an asset is enrichment at write time, partitioning for retrieval, and open table formats every future reader can use.
Events land carrying context that never existed before: which team owns the subnet, whether the IP is internal or external, which service the host belongs to. Partitioning follows investigation patterns so a year of history comes back in seconds at pennies. Open table formats keep the data readable by the query engine you use today and the model you point at it next year.
Then the agent phases put it to work: triage and root cause reading full-fidelity history, and the hunt across everything nobody thought to alert on.
04Questions
Asked about data lake estates
We already store everything in S3. Is that not enough?
Storage is the prerequisite, not the payoff. Unenriched, unpartitioned data makes every question an export job, so nobody asks. Enrichment, partitioning, and open table formats are what turn kept data into answered questions, and they are retrofittable onto the lake you have.
Which table format do you build on?
Open ones, with Iceberg as the direction the market has chosen. The specific layout follows your query engines and your cloud in the review. The principle does not move: no format a single vendor controls gets to hold your history.
Can the Lake replace our SIEM?
No, and we will not pitch that. The Lake never alerts. Detection stays on the SIEM, which is built for it. The Lake holds the full-fidelity history the SIEM should not have to, serves compliance and investigations, and feeds the agents. Each layer does the job it is priced for. What the Lake does change is the lock-in: with your history in your own storage, swapping the SIEM becomes a routing change and a migration in weeks, not months.
Start with the review
You share your diagrams, we review them with you, and you leave with your version of the Logmetry Blueprint drawn on your data lake estate. No system access, no obligation.