HNHacker News
TopNewBestAskShowJobs

tomnicholas1

160 karma · joined August 19, 2024

OSS python software for science, currently employed at earthmover.io. I help maintain Xarray, Zarr, Icechunk, VirtualiZarr, and other pangeo.io projects.

https://github.com/TomNicholas

submissionscomments
tomnicholas1··on Evidence of Fraud in an Influential Study About Procrastination
> The problem in climatology is that when the models don't fit the data they just change the data and claim victory

That is... not remotely true, nor supported by either of the links you shared. (Nor by any cursory look at the state of the world today, with well-anticipated Himalayan glacial floods as front-page news.) You therefore appear to be leaping over the line from healthy skepticism to irrational conspiratorial thinking, so I'm going to stop engaging, sorry.

tomnicholas1··on Evidence of Fraud in an Influential Study About Procrastination
Whilst I weakly agree with this characterisation, it's not at all fair or accurate to put epidemiology and climatology in the same category as your other examples. Yes they have some weaknesses in replication practices (especially for purely computational work), yes the feedback loop for self-correction can be slow, but it would be a mistake to conclude or imply that the major results of those fields are therefore unreliable.

That's because those fields do still make objective empirical predictions about the world, which might take years to be testable but are still ultimately testable. For example, it's been convincingly shown that climate models from decades ago predicted our current climate pretty well. [0]

These fields are more comparable to subfields of physics which only get to do big experiments a handful of times, e.g. because they can only observe so many planets/supernovae/universes, or because they have to spend decades building a new billion-dollar machine to test new hypotheses.

We should also not confuse these arguments with uselessness - soft social science work can still be societally valuable without being easily empirically testable.

[0] https://agupubs.onlinelibrary.wiley.com/doi/full/10.1029/201...

tomnicholas1··on Changes at Google DeepMind: Demis Hassabis from CEO to Chair, Jeff Dean departs
> Data from physics, chemistry, biology experiments

There are PetaBytes of important scientific data locked in archival file formats. The first step is to make this efficiently readable.

https://www.earthmover.io/blog/virtual-zarr

https://news.ycombinator.com/item?id=46659254

tomnicholas1··on Goes-19 weather satellite enters Safe Hold mode
Yes! But if you want to share it with anyone else that would be great, since we're advocating for fairly radical changes within a big bureaucracy here, as I'm sure you will appreciate :)
tomnicholas1··on Goes-19 weather satellite enters Safe Hold mode
Anyone interested in accessing GOES data at scale will find this interesting - I created a Zarr index over the 7 billion chunks of data in the GOES-16 archive.

https://www.earthmover.io/blog/virtual-zarr

tomnicholas1··on Cloud-optimizing the GOES-16 satellite data archive without copying data
The last 2 years of my professional life have in some sense all been working towards this blog post.

It shows how you can make an enormous archive of public scientific data[1] much easier to use for everyone, without having to copy all the data, which would be prohibitively expensive at this scale.

I also made some sweet gifs and images of the Earth from space from the raw data so check those out!

[1] In this case satellite data, but this could work for data from many other fields - my background is original nuclear fusion plasma physics then oceanography.

tomnicholas1··on Matadisco – Decentralized Data Discovery
Awesome to see this project here - it was partly inspired by my blog post (original is linked from the OP, but there's a slightly newer version on my personal site here[0]).

[0]: https://tom-nicholas.com/blog/2025/science-needs-a-social-ne...

tomnicholas1··on A distributed queue in a single JSON file on object storage
What you describe is very similar to how Icechunk[1] works. It works beautifully for transactional writes to "repos" containing PBs of scientific array data in object storage.

[1]: https://icechunk.io/en/latest/

tomnicholas1··on Show HN: Streaming gigabyte medical images from S3 without downloading them
People have literally used Zarr for this - at one point Gemini used Zarr for checkpointing model weights. Not sure what the current fashion in that space is though.

It's definitely one of many fields that see convergent evolution towards something that just looks like Zarr. In fact you can use VirtualiZarr to parse HuggingFace's "SafeTensors" format [0].

[0]: https://github.com/zarr-developers/VirtualiZarr/pull/555

tomnicholas1··on Show HN: Streaming gigabyte medical images from S3 without downloading them
IMO Zarr is that newer format. It abstracts over the features of all these other formats so neatly that it can literally subsume them.

I feel that we no longer really need TIFF etc. - for scientific use cases in the cloud Zarr is all that's needed going forwards. The other file formats become just archival blobs that either are converted to Zarr or pointed at by virtual Zarr stores.

tomnicholas1··on Show HN: Streaming gigabyte medical images from S3 without downloading them
The generalized form of this range-request-based streaming approach looks something like my project VirtualiZarr [0].

Many of these scientific file formats (HDF5, netCDF, TIFF/COG, FITS, GRIB, JPEG and more) are essentially just contiguous multidimensional array(/"tensor") chunks embedded alongside metadata about what's in the chunks. Efficiently fetching these from object storage is just about efficiently fetching the metadata up front so you know where the chunks you want are [1].

The data model of Zarr [2] generalizes this pattern pretty well, so that when backed by Icechunk [3], you can store a "datacube" of "virtual chunk references" that point at chunks anywhere inside the original files on S3.

This allows you to stream data out as fast as the S3 network connection allows [4], and then you're free to pull that directly, or build tile servers on top of it [5].

In the Pangeo project and at Earthmover we do all this for Weather and Climate science data. But the underlying OSS stack is domain-agnostic, so works for all sorts of multidimensional array data, and VirtualiZarr has a plugin system for parsing different scientific file formats.

I would love to see if someone could create a virtual Zarr store pointing at this WSI data!

[0]: https://virtualizarr.readthedocs.io/en/stable/

[1]: https://earthmover.io/blog/fundamentals-what-is-cloud-optimi...

[2]: https://earthmover.io/blog/what-is-zarr

[3]: https://earthmover.io/blog/icechunk-1-0-production-grade-clo...

[4]: https://earthmover.io/blog/i-o-maxing-tensors-in-the-cloud

[5]: https://earthmover.io/blog/announcing-flux

tomnicholas1··on Programmers and software developers lost the plot on naming their tools
God this article is 10000% better than the posted one. This is great:

> Names should not describe what you currently think the thing you’re naming is for. Imagine naming your newborn child "Doctor", or "SupportsMeInMyOldAge". Poor kid.

tomnicholas1··on F3: Open-source data file format for the future [pdf]
Thank you for the explanation! But what a mess.

I would love to bring these benefits to the multidimensional array world, via integration with the Zarr/Icechunk formats somehow (which I work on). But this fragmentation of formats makes it very hard to know where to start.

tomnicholas1··on F3: Open-source data file format for the future [pdf]
The pitch for this sounds very similar to the pitch for Vortex (i.e. obviating the need to create a new format every time a shift occurs in data processing and computing by providing a data organization structure and a general-purpose API to allow developers to add new encoding schemes easily).

But I'm not totally clear what the relationship between F3 and Vortex is. It says their prototype uses the encoding implementation in Vortex, but does not use the Vortex type system?

tomnicholas1··on Progress toward fusion energy gain as measured against the Lawson criteria
The really depressing part is if you plot rate of new delays against real time elapsed, the projected finishing date is even further.

This is why much of the fusion research community feel disillusioned with ITER, and so are more interested in these smaller (and supposedly more "agile") machines with high-temperature superconductors instead.

tomnicholas1··on Progress toward fusion energy gain as measured against the Lawson criteria
Presumably because everyone in MCF has been waiting for ITER for decades, and JET is being decommissioned after a last gasp. Every other tokamak is considerably smaller (or similar size like DIII-D or JT-60SA).

Much of the interesting tokamak engineering ideas were on small (so low-power) machines or just concepts using high-temperature superconducting magnets.

tomnicholas1··on What Is Cloud-Optimized Scientific Data?
I wrote the article I wish I could have read back when I first heard of Zarr and cloud-native science back in 2018.

This explains how object storage and conventional filesystems are different, and the key properties that make Zarr work so well in cloud object storage.

tomnicholas1··on What Is Entropy?
Isn't that more about enumerating the microstates? The Pauli exclusion principle just ends up forbidding some of the microstates (forbidding a significant fraction of them if you're in the low-temperature regime).
tomnicholas1··on What Is Entropy?
Yes, that assumption is called the Ergodic Hypothesis, and generally justified in undergraduate statistical mechanics courses by proving and appealing to Liouville's theorem.

[1] https://en.wikipedia.org/wiki/Ergodic_hypothesis

tomnicholas1··on Tensors vs. Tables: Why tabular tools trip over gridded data
The scientific community works primarily with array (or "tensor") data, using tools like numpy, xarray, and zarr. People familiar with modern relational database tools such as DuckDB and Parquet often ask why can't we just use those? This article explains why: it's massively inefficient to use tabular tools on array data, and demonstrates with a benchmark showing a 10x difference in query speed.
tomnicholas1··on Apache iceberg the Hadoop of the modern-data-stack?
I think the posted article was generated from this one - the structure of the content is so similar.
tomnicholas1··on Building an Open, Multi-Engine Data Lakehouse with S3 and Python
This entire stack also now exists for arrays as well as for tabular data. It's still S3 for storage, but Zarr instead of parquet, Icechunk instead of Iceberg, and Xarray for queries in python.
tomnicholas1··on Apache Iceberg now supports geospatial data types natively
Icechunk can handle growing dimensions with ACID transactions!

For irregular shapes in some cases using multiple groups + xarray.DataTree can help you, but in general yeah ragged data is hard.

tomnicholas1··on Apache Iceberg now supports geospatial data types natively
Surely Zarr is already a long-term storage format for multidimensional data? It can even be mapped directly to netCDF, GRIB and geoTIFF via VirtualiZarr[0].

Also if you like Iceberg and you like arrays you will really like Icechunk[1], which is Version-controlled Zarr!

[0] https://github.com/zarr-developers/VirtualiZarr

[1] https://icechunk.io/en/latest/

tomnicholas1··on Julia and JuliaHub: Advancing Innovation and Growth
> The future of Python's main open source data science ecosystem, numfocus, does not seem bright. Despite performance improvements, Python will always be a glue language.

Your first sentence is a scorching hot take, but I don't see how it's justified by your second sentence.

The community always understood that python is a glue language, which is why the bottleneck interfaces (with IO or between array types) are implemented in lower-level languages or ABIs. The former was originally C but often is now Rust, and Apache Arrow is a great example of the latter.

The strength of using Python is when you want to do anything beyond pure computation (e.g. networking) the rest of the world already built a package for that.

tomnicholas1··on [dead]
I’ve been thinking a lot recently about how one thing science needs is a social network for sharing big data.

One thing the post gets at is that providing a decentralized global subscribable data catalog is fundamentally a network protocol problem, somewhat similar to RSS.

The social network analogy is particularly generative here - because the desired network structure is similar to that of Federated social media, the protocol I want would be very similar in structure to those protocols underlying attempts to decentralize social media platforms. I therefore think it might well be possible to build what I’m suggesting by piggybacking off of BlueSky’s AT protocol or the Fediverse/Mastodon’s ActivityPub protocol.

I’ve made a repo[1] for discussing ideas for how one might implement such a protocol.

Curious what people think of any of this!

[1] https://github.com/TomNicholas/FROST

tomnicholas1··on [dead]
Reporter imagines how we'd cover overseas what's happening to the U.S. right now.

This is by far the most clear-eyed description I've seen of the severity of what is happening.

Relevant to HackerNews partly because it includes descriptions of how Big Tech executives are effectively bribing the Trump Administration in exchange for removing federal watchdogs.

tomnicholas1··on Hydro: Distributed Programming Framework for Rust
I was also going to say this looks similar to one layer of dask - dask takes arbitrary python code and uses cloudpickle to serialise it in order to propagate dependencies to workers, this seems to be an equivalent layer for rust.
tomnicholas1··on GitHub Git Operations Are Down
Is there a world in which GitHub used an open protocol for the social network part of their product like BlueSky's AT protocol[0]?

[0] https://docs.bsky.app/docs/advanced-guides/atproto

tomnicholas1··on Why aren't we all serverless yet?
In python there is lithops, which provides nice Executor primitives that can run on a wide range of cloud services (AWS lambda, GCF etc.)

https://github.com/lithops-cloud/lithops

Page 1 of 2Next →