Fast drilldown dashboards from a single Parquet file
hamiltonulmer.com
hamiltonulmer.com
For a 40MB file I suggest hosting it directly on GitHub Pages - that's effectively a free CORS-enabled CDN and supports HTTP range requests, so you should be able to get that demo working without needing to involve Cloudflare Workers at all.
If you are okay with it being down regularly
Edit: Ironically, that would be the case now: https://www.githubstatus.com/incidents/hcbtzksccj2f
> your pipeline has to rebuild each customer’s file fast enough to meet the update cadence. ... data that updates on a coarse schedule rather than in realtime
You're making it seem like there's hard limits to what can be done but while there definitely is, you can do incredible stuff.
I expect that if your overall data is less than a GB this trick will work really well for you.
The next step in this journey is Iceberg (and a proper incremental pipeline), which can also be read directly in the browser via WASM either via DuckDB or without. This is from the same author as the parquet library mentioned in the OP https://github.com/hyparam/icebird
Traditional observability is ill-suited for observability around business events. What if you forget to instrument a counter or gauge for something? In my experience it's far easier to log wide events with as much context as possible instead of agonizing over anticipating the dimensionality of metrics upfront (you're going to miss something).
Yes this was also my experience working at a large tech co. I work in fintech now and data volumes are low enough to maintain 2-3 minute up to a few hour data freshness.
What's the benefit of "data cubes" over caching?
I provide caveats for when this would work vs. when it doesn't in the post. For a lot of customer-facing dashboards, I think it's probably pretty good.
It will cost (slightly) more to write this frequently to R2 since you are charged per-write, but this is something you can tune.
https://duckdb.org/docs/current/sql/query_syntax/grouping_se...