Put all of that together, and you get a website that queries S3 with no backend at all. Amazing.
Put all of that together, and you get a website that queries S3 with no backend at all. Amazing.
Cloudflare actually has built in iceberg support for R2 buckets. It's quite nice.
Combine that with their pipelines it's a simple http request to ingest, then just point duckdb to the iceberg enabled R2 bucket to analyze.
There's no egress data transfer fees, but you still pay for the GET request operations. Lots of little range requests can add up quick.
It is time like this that makes self-hosting a lot more attractive.
You have to store the data somehow anyway, and you have to retrieve some of it to service a query. If egress costs too much you could always change later to put the browser code on a server. Also it would presumably be possible to quantify the trade-off between processing the data client side and on the server.
But yeah - this is pretty neat. Easily seems like the future of static datasets should wind up in something like this. Just data, with some well chosen indices.
Lack of server/dynamic code qualifies as no backend.
I’m just bemused that we all refer to one of the larger, more sophisticated storage systems on the plant, composed of dozens of subsystems and thousands of servers as “no backend at all.” Kind of a “draw the rest of the owl”.
They let you easily abstract over storage.
https://2019.splashcon.org/details/splash-2019-Onward-papers...