HNHacker News
TopNewBestAskShowJobs

rabernat

30 karma · joined February 15, 2016

Startup Founder, Scientist and Software Developer

- CEO and co-founder of Earthmover: https://earthmover.io/ - Associate professor in the Columbia University Department of Earth and Environmental Science and Lamont Doherty Earth Observatory: https://ocean-transport.github.io/ - Co-founder and community leader of the Pangeo Project: https://pangeo.io/

submissionscomments
rabernat··on Loading a trillion rows of weather data into TimescaleDB
> It's possible but not very cost-effective to maintain separately-chunked versions of these large geospatial datasets.

Like all things in tech, it's about tradeoffs. S3 storage costs about $275 TB a year. Typical weather datasets are ~10 TB. If you're running a business that uses weather data in operations to make money, you could easily afford to make 2-3 copies that are optimized for different query patterns. We see many teams doing this today in production. That's still much cheaper (and more flexible) than putting the same volume of data in a RDBMS, given the relative cost of S3 vs. persistent disks.

The real hidden costs of all of these solutions is the developer time operating the data pipelines for the transformation.

rabernat··on Loading a trillion rows of weather data into TimescaleDB
True, but in fact, the Google ERA5 public data suffers from the exact chunking problem described in the post: it's optimized for spatial queries, not timeseries queries. I just ran a benchmark, and it took me 20 minutes to pull a timeseries of a single variable at a single point!

This highlights the needs for timeseries-optimized chunking if that is your anticipated usage pattern.

rabernat··on Loading a trillion rows of weather data into TimescaleDB
Great post! Hi Ali!

I think what's missing here is an analysis of what is gained by moving the weather data into a RDBMS. The motivation is to speed up queries. But what's the baseline?

As someone very familiar with this tech landscape (maintainer of Xarray and Zarr, founder of https://earthmover.io/), I know that serverless solutions + object storage can deliver very low latency performance (sub second) for timeseries queries on weather data--much faster than the 30 minutes cited here--_if_ the data are chunked appropriately in Zarr. Given the difficulty of data ingestion described in this post, it's worth seriously evaluating those solutions before going down the RDBMS path.

rabernat··on We're wasting money by only supporting gzip for raw DNA files
The Zarr format is used in some genomics workflows (see https://github.com/zarr-developers/community/issues/19) and supports a wide range of modern compressors (e.g. Zstd, Zlib, BZ2, LZMA, ZFPY, Blosc, as well as many filters.)
rabernat··on Request for Startups: Climate Tech
PyTorch and JAX are used heavily in climate science on the ML side. For more general analytics, not so much. Many of our users like to use Xarray as a high-level API. There has been some work to integrate Xarray with PyTorch (https://github.com/pydata/xarray/issues/3232) but we're not there yet.

The Python Array API standard should help align these different back-ends: https://data-apis.org/array-api/latest/

rabernat··on Request for Startups: Climate Tech
Agree 100%. This is big part of the motivation behind our new startup Earthmover: https://earthmover.io/

Our mission is to make it easier to work with scientific data at scale in the cloud, focusing mainly on the climate, weather, and geospatial vertical.

My cofounder Joe Hamman and I are climate scientists who helped create the Pangeo project. We are also core devs on the Python packages Xarray and Zarr. We think that a layer of managed services (think a "modern data stack" oriented around the multidimensional array data model) is exactly what this ecosystem needs to make it easier for teams to build data-intensive products in the climate-tech space.

And we're hiring! https://earthmover.io/posts/earthmover-is-hiring/