HNHacker News
TopNewBestAskShowJobs

tomnicholas1

160 karma · joined August 19, 2024

OSS python software for science, currently employed at earthmover.io. I help maintain Xarray, Zarr, Icechunk, VirtualiZarr, and other pangeo.io projects.

https://github.com/TomNicholas

submissionscomments
tomnicholas1··on Property-Based Testing for the People
The python package Hypothesis[0] already does a great job bringing property-based testing to the people! I've used it and it's extremely powerful.

[0]: https://github.com/HypothesisWorks/hypothesis

tomnicholas1··on Spectral Imaging Made Easy: A Powerful Python Library
Interesting - I'm curious whether you feel that Xarray covers these use cases already?

https://xarray.dev/

Especially as I've said before that Hyperspy shares so many features in common with Xarray that Hyperspy should just use Xarray under the hood.

https://github.com/hyperspy/hyperspy/discussions/3405

tomnicholas1··on The Parker Solar Probe will make its closest approach yet to the Sun
No not necessarily - it will keep growing hotter until the black body radiation emitted by the probe matches the power of the radiation hitting the probe. Then it will stay at constant temperature.

It's a standard undergraduate problem to work out what this equilibrium temperature is for a flat plate at a distance from the sun equal to the Earth's orbital radius.

Interestingly the result is only a few 10's of degrees less than the average temperature of the real Earth - the difference is due to the Greenhouse Effect.

For the probe one could easily do the maths but I could believe that at 4 million miles that equilibrium temperature is 2,500F.

tomnicholas1··on Show HN: Gribstream.com – Historical Weather Forecast API
That's pretty cool! Quite specific to this file format/workload, but this is an important enough problem that people might well be interested in a tailored solution like this :)
tomnicholas1··on Show HN: Gribstream.com – Historical Weather Forecast API
> It is sort of analogous to what GribStream is doing already.

The difference is presumably that you are doing some large rechunking operation on your server to hide from the user the fact that the data is actually in multiple files?

Cool project btw, would love to hear a little more about how it works underneath :)

tomnicholas1··on Show HN: Gribstream.com – Historical Weather Forecast API
> Why is weather data stored in netcdf instead of tensors or sparse tensors?

NetCDF is a "tensor", at least in the sense of being a self-describing multi-dimensional array format. The bigger problem is that it's not a Cloud-Optimized format, which is why Zarr has become popular.

> Also, SQLite supports virtual tables that can be backed by Content Range requests

The multi-dimensional equivalent of this is "virtual Zarr". I made this library to create virtual Zarr stores pointing at archival data (e.g. netCDF and GRIB)

https://github.com/zarr-developers/VirtualiZarr

> xarray and thus NetCDF-style multidimensional arrays in WASM in the browser with HTTP Content Range requests to fetch and cache just the data requested

Pretty sure you can do this today already using Xarray and fsspec.

tomnicholas1··on Humans have caused 1.5 °C of long-term global warming according to new estimates
That's partly because the warming experienced over land can be ~50-100% larger than the globally-averaged warming, with the temperatures over the oceans increasing more slowly to make up the difference.

https://www.carbonbrief.org/guest-post-why-does-land-warm-up...

tomnicholas1··on Data Version Control
You should look at Icechunk. Your imaging data is structured (it's a multidimensional array), so it should be possible be to represent it as "Virtual Zarr". Then you could commit it to an Icechunk store.

https://earthmover.io/blog/icechunk

tomnicholas1··on Data Version Control
If you're wondering this you should look at Icechunk too, which was open-sourced just this week. It's Apache Iceberg but for multidimensional data (e.g. Zarr).

https://earthmover.io/blog/icechunk

https://news.ycombinator.com/item?id=41850352

tomnicholas1··on Launch HN: Sorcerer (YC S24) – Weather balloons that collect more data
So the equivalent of these balloons in oceanography are called ARGO floats, which similarly cannot be driven laterally but can control their own depth like a submarine. So far millions of timeseries have been collected across the world ocean using these floats.

https://argo.ucsd.edu/

One difference though is that the ARGO floats are unfortunately not recycled, and just wash up on various beaches. (I'm curious whether you think you can realistically collect many of these mini balloons?)

If you do want to control the lateral position of fleets of sensors, oceanographers also now have "gliders", which are basically small powered drone submarines. These are used by a few groups, but most of the gliders in the world are operated by the US Navy, who launch them out of torpedo tubes to survey local ocean conditions (which is badass).

https://oceanservice.noaa.gov/facts/ocean-gliders.html

The recorded measurements present an interesting data assimilation challenge - they record data along 3D trajectories (4D including time), sampling jagged and twisting lines through the 4D space. But we normally prefer to think of weather/ocean data as gridded, so you need to interpolate the trajectory data onto the grid, whilst keeping the result physically-consistent. Oceanographers use systems like ECCO for ocean state estimation, which effectively find the "ocean of best fit" to various data sources.

https://www.ecco-group.org/

Interestingly ECCO uses an auto-differentiable form of the governing equations for the ocean flow to ensure that updates stay physically consistent. This works by using a differentiable ocean fluid model called [MITgcm](https://github.com/MITgcm/MITgcm) to perform runs which match experimental data as closely as possible, and minimizing a loss function through gradient descent. The gradient is of a loss function (error) with respect to model input parameters + forcings, which is calculated by running MITgcm in adjoint mode - i.e. automatic differentation. Therefore this approach is sort of ML before it was cool (they were doing all this well before the new batch of AI weather models). See slides 9-18 of this deck for a nice explanation

https://firebasestorage.googleapis.com/v0/b/firescript-577a2...

The trajectory data is also interesting because it's sort of tabular, but also you often want to query it in an array-like 4D space. You could also call it a "ragged" array. We have nice open-source tools for gridded (non-ragged) arrays (e.g. xarray and zarr, and the pangeo.io project) but I think we could provide scientists with better tools for trajectory-like data in general. If that seems relevant to you I would love to chat.

P.S: Sorceror seems awesome, and I applaud you for working on something hard-tech & climate-tech!

← PreviousPage 2 of 2