HNHacker News
TopNewBestAskShowJobs

nikhil4usinha

5 karma · joined April 17, 2025

submissionscomments
nikhil4usinha··on Show HN: Parseable, an open observability datalake, handles 100M time-series/min
Fair question. 100M isn't a ceiling, it's what we have seen in that deployment. We have not run a billion series test yet.

The reason we think it scales differently - labels are just columns in Parquet, so there is no per series index that grows with cardinality. In that deployment one label alone has ~2.5M distinct values among 500+ labels, which would be painful for an index based TSDB but here is just a high cardinality column. What drives cost for us is ingestion rate (data points/s) and how much data a query has to scan for a particular time range not series count. Ingest scales horizontally by adding ingestors, and queries prune by time partition and column stats.

A billion series benchmark is on our list, and we'll publish the numbers when we run it.

nikhil4usinha··on Show HN: Parseable, an open observability datalake, handles 100M time-series/min
scrape interval is 15s and sustained ingestion we have seen is ~3M samples/sec that is ~300 TB/day of raw ingest payload, when stored on object store as parquet, the data gets compressed to 99% which makes it 3 TB/day. The 100M figure is total unique series seen over time. For a sense of per metric cardinality, one metric that has the highest cardinality label (2.5 M distinct values) shows ~6M active series per hour.
nikhil4usinha··on Zero-Shot Forecasting: Our Search for a Time-Series Foundation Model
Interesting, what are the usecases youre using the models for? Would like to know more on that, like anomaly detection