Open-source could finally get the world’s microscopes speaking the same language
nature.com
nature.com
In particular the Pangeo project (https://pangeo.io/architecture.html) uses large Zarr stores as a performant format in the cloud which we can analyse in parallel at scale using distributed computing frameworks like dask.
More and more climate science data is being made publicly available as Zarr in the cloud, often through open data partnerships with cloud providers (e.g. on AWS (https://aws.amazon.com/blogs/publicsector/decrease-geospatia...) ERA-5 on GCP(https://cloud.google.com/storage/docs/public-datasets/era5)).
I personally think that the more that common tooling can be shared between scientific disciplines the better.
For scientific instruments just the opposite took place in the 1990's with chromatography.
They adapted a climate file type for laboratory instrument use.
It was quite a saga but in a nutshell it was the manufacturers who well understood the benefit of a common file type once computers had taken enough of a hold. Most lab analyses didn't need a computer except for the most data-intensive, which could never wait and were being handled by mainframes or minis before the PC appeared. Only by those that could afford such costly investments like oil companies and drug companies. Which are some of the most desirable customers for instrumentation.
At the time HP's lab instrument division (spun off decades ago as part of Agilent) had the most agressive sales approach but if a potential customer was already entrenched with one of the few other vendors there was no efficient way to carry their past proprietary data forward into a new system.
So the vendors got together, formed the Analytical Instrument Association (AIA) and began work in earnest toward a common file type they could all agree on for data interchange, without having to abandon each of the file types they had developed earlier which remained in everyday use within their proprietary software.
Once this had been accomplished it would be more of a battle of the salespeople with one less (big) technical obstacle to overcome.
They ended up choosing a well-established extensible file type, netCDF, which was in wide use at NOAA without encumbrance. Tailored it to the task and after a number of years the new standard was almost complete. Major vendors were already supporting it.
Well, by that time HP had already become dominant so they stealthily lost interest for a while before pulling out of AIA completely.
Fortunately with no further progress in sight the whole thing was dropped in the lap of ASTM, where it has remained unchanged for over two decades. That is even more fortunate because hindsight has shown the difference between an ideal standard and a nonideal standard is that an ideal standard never changes, and this is one of the best examples. These are often referred to as CDF files even though that does not stand for Chromatography Data File, they were just named netCDF by the weather service during the mainframe days.
So for more than 20 years each PC chromatography data system has had the capability of handling AIA files for either input or output. The underlying proprietary files which each software package utilizes under the hood for real-time data acquisition may have some unique features, but the common netCDF interchange covers all the essential bases needed for post-processing. Anything less is sub-standard.
Only in the last few years has a major software package appeared which for the first time does not fully support CDF's and that is the latest offering from Agilent of all places.
It can only create optional CDF files from the data acquired by their own instruments, so far there is no allowance for inputting a standard CDF file for processing by their software any more.
Right not there's basically zero funding for people who work on making our data open. It's kind of a hobby project that you do if you have free time and no one has found a real job for you.
I think it's important to acknowledge here that there are degrees of open access. Simply providing data files is relatively easy. Making reproducible workflows as a non-developer can be nearly impossible in some cases.
Why is that the case? Size?
I am very familiar with dealing with old proprietary things that have to be fought tooth and nail to give you a simple text format, but that doesn't seem to be nearly the scale of trouble as you describe.
Just getting old software to compile can be a challenge in itself. You could say to just put everything in a Docker container. Then when you have it working, anyone can just use that container. That's not so easy for someone who doesn't even know what Docker is. I'm not saying that to suggest people are incompetent. As developers, we really should do better to make things easy.
What is ingrained is the idea among my collaborators that open data is kind of an "us vs them" problem: it's our data, and the public hasn't been demanding access, but if we gave it to them they would just scoop us and steal all the science, or they would do something stupid and force us to disprove it.
It's all quite speculative but the limited data releases so far haven't caused the field to collapse as some doomsayers predicted. Meanwhile I'd just like more open data because it would make collaboration easier.
https://lists.andrew.cmu.edu/pipermail/openslide-users/2012-...
Here's hoping this eventually gets fixed in a similar sort of way the .mkv format "fixed" video.
I am not sure if the video actually says the name of the microscope, though, because he was mad at them for not testing it at high altitude where he lived at the time.
How much does one unit cost?
Mochii starts at $48K for the imaging only unit. The spectroscopy-enabled version, which provides full featured x-ray spectroscopy and spectrum imaging, is $65K.
It has an integrated metal coater option available for $5,000, and we offer a variety of optical cartridge exchange programs that can fit your consumables utilization and pricing needs.
https://en.wikipedia.org/wiki/Optical_microscope#Alternative...
I was hoping for an easy entry:
https://www.hackster.io/michalin70/3d-printed-laser-scanning...
I'm probably not being cynical enough here, but this doesn't necessarily imply a desire for obfuscation.
Codecs are hard, I doubt there are many which are better than the well known ones.
That said, it's certainly worthwhile benchmarking (de)compression speed and the compression ratio to see if the difference is actually meaningful enough to be worthwhile.
S3 supports the range header. Google Cloud, Azure, and I'm sure other object stores support it too.
It's ironic that they mention FAIR but don't see how much better it would be to have a serialized byte-stream you can cite and share with checksums etc.
Research labs literally have a thousand other things to do... besides play wack-a-mole with vendor IT shenanigans. =)
Each pixel? Why? All of those except the 3D position apply to all the pixels in a given image, and the (2D) position of a pixel can be inferred from its location in the image.
Wait - are there optical microscopes that can create 3D images? I know you can see a 3D image if you peer into a binocular microscope, but AFAIK cameras for those things are always 2D cameras.
There are also confocal "laser-scanning" microscopes which effectively illuminate a single point in space and image with what is effectively a single monochrome pixel sensor. In this case, the raw signal is pretty much a time series like you might imagine with digital audio. You need metadata to tell you how the optics were being shifted around, to interpret each sample as representing a point in space at a particular time and with what filters.
Normally, the first round of interpretation is done during the initial file save, packing measurements into a file format that still expects planar, rectangular pixel arrays with a regular grid spacing for the neighboring pixels. You pretend the whole image plane was captured at a time point, but it was really measured sequentially somewhat like a cathode-ray tube displays an image one pixel at a time. A session might produce many such planes, and metadata is still needed to understand whether these planes represent the same plane observed over time and/or parallel planes spanning a volume and/or different channels of the same plane.
On the other end of the spectrum, there are slide scanners which image a much larger stage area by shifting around a 2D camera and taking many overlapping images. So a 2048x2048 sensor might be used to produce a 100,000 x 60,000 pixel image plane. Each sub-image would record its position within the stage area and these do not necessarily line up at precise multiples of image pixels. There might also be gaps where the scanner first detects the overall shape of specimen(s) on the slide and then plans a set of images to cover the specimen while skipping over the background area. This sparsity saves time for imaging as well as storage and transfer times by skipping empty areas.
I use the Zarr format (for climate science data rather than microscope data), and I think this is just poor wording in the article. In the Zarr specification the metadata is stored separately from the actual chunks of compressed array data. So the metadata applies at the array level, not the pixel level.
> Wait - are there optical microscopes that can create 3D images?
I think so - they do it by scanning lots of images at different focal lengths to create a 3D section (I think?). There are whole projects just for visualizing the multi-TeraByte 3D image files produced - Napari is an open-source image viewer which opens OME-Zarr data.
Source: PhD in biolab with microscopes, and napari dev.
Flat fields, dark fields, and light fields can all be stored but would be their own arrays (structure of arrays rather than array of structures).
But that's post-processing; isn't the TFA's argument that it's hard for researchers to share raw data?
Instead, they have a collection of dense arrays representing the image data itself, then have metadata at the per-array, or overall level.
A typical dataset I work with is multidimensional, it starts as: 1) 2D planes of multiple channel image intensities, typically 5K x 5K pixels, each covering just part of an overall field of view. These are like patches when you do panoramas- take 20 partly overlapping shots. Each plane contains multiple channels- that could be "red green and blue" or more complicated spectral distributions.
2) 3D information- the microscope takes photos at various depths, only the "in-focus" (within the volume of view) information. These can be stacked (like depth stacking) or turned into a 3D "volume".
3) Maybe the data was collected over multiple time points, so (1) and (2) repeat every hour. Other parameters- like temperature, etc, could also represent an entire dimension.
4) Every 2D plane has its own key-value metadata, such as "what color channels were used", "what objective was used" (magnification), and lots of other high-dimensional attributes (that's what they mean by "each pixel must be labelled with metadata"- the 3D position is the same for every pixel in a 2D plane.
5)
Generally all of this is modelled as structures, arrays, and structures of arrays/arrays of structures. In the case of OME-zarr, it's modelled as an n-Dimensional array with dimensions expressed in a filesystem hierarchy (first dimension typicall the outermost directory, innermost dimension usually a flat file containing a block of scalar points using some compressed numpy storage. Then at each level of the directory you have additional .json files which contain attributes at that level of the data.
Those partly overlapping 2D planes are often assembled into panoramas, which can be a lot more convenient to work with. There are various tools for working with this- I've used navigation map javascript, but napari is a desktop app wiht full suport for sectioned viewing of high-dimensional (7d) data.
OME-zarr is nice because it sort of uses the same underlying tech that the machine learning folks use, and it's ostensibly "optimized for object storage", but I still have lots of complaints with the implementation details, but it's important for me not to distract the OME-zarr team from making the standard successful.
Thinking out loud, this looks like the consolidation of several formats and projects. Bioformats, OME, Zarr...
Back in the day I had the feeling that the OME/Bioformats people were more centered on research, and the DICOM crowd were mostly from the clinical side.
Don't get me started on Euler angles.
Caveat being I'm more used to electron microscopy, maybe these things aren't as important with light microscopy because the resolutions are lower?
It's one thing for a highly-skilled user with a decade of experience to be able to import it eventually, another for an unskilled user to just have off-the-shelf tooling do it the same way for everyone automatically. This is more about meta-studies and unlocking new use cases. It sucks that the state-of-the-art for sharing academic microscopy data right now is basically looking at raster images embedded in PDFs, or email the authors, and then write a custom Scikit-Image script. Imagine if you had to read a PDF catalog and then email someone to order something off Amazon, or if your favorite CRUD app consisted instead of having an expert read a PDF and email a screenshots to you. What if sending those very emails to different recipients required implementing each users custom IMAP-like mail client. That sounds absurd, but it's kind of the way academic data sharing works now, lots of people are re-inventing the wheel and creating custom file formats.
Consider, for example, the work of Dr. Bik (example at [1]) who identifies cloned sections from microscopy data. Or what if, instead of each researcher having to generate their own images, or get lucky and remember a particular image there was a Getty Images/AP Newsroom platform where you could just filter for your particular subject and imaging parameters and share your data. A collection of proprietary RAW files with randomly-formatted Excel documents for metadata would allow individual researchers to get their work done, but would be pretty worthless in comparison.
[1] https://scienceintegritydigest.com/2023/06/27/concerns-about...
We support DICOM supp 145 too, but it's no panacea. There are still vendor specific quirks. The "surface" is larger (cause you expect all the metadata to be there in the standard format) so you still sometimes see differences.
unless scope manufacturers unite behind a format, biologists are going to keep writing bespoke analysis pipelines for the formats their images come in. it's going to take some top down regulation from the NSF/NIH requiring that all scopes support the open format of choice by a certain date to be eligible for purchase with grant money.
Not always a very open solution but converge they do. See matlab etc.
We probably need a multidimensional lossy compression scheme for 2D images, 3D volumes/video, and 4+D metadata involving multiple spectrums, exposures, etc. I'd probably start with a multidimensional DCT with a way to specify the coefficients/weights of each layer and metadata (optionally losslessly):
https://en.wikipedia.org/wiki/Discrete_cosine_transform#Mult...
Looks like that's built into MATLAB with dct(x,n,dim):
https://www.mathworks.com/help/signal/ref/dct.html
https://www.mathworks.com/help/images/ref/blockproc.html (coefficients)
https://www.mathworks.com/help/images/discrete-cosine-transf... (2D example)
Then store the data in zipped JSON or BSON, or maybe TAR then ZIP.
Then I'd write an importer and exporter from/to all of the other formats, with a single 0-1 quality argument or maybe an array of qualities for each layer, with a way to specify coefficients or load them from files.
Once a known-good implementation worked, I'd add additional compression schemes that can compress across dimensions, so that a level of zoom only has to save a diff from a previous layer at a slightly different scale or focus (like a 3+D MPEG format, this is probably the standard that we're currently missing, does anyone know of one?).
This all seems fairly straightforward though, so I wonder if proprietary software patents are the real blocker..
https://github.com/ome/ome-zarr-py
https://ngff.openmicroscopy.org/latest/
https://zarr.readthedocs.io/en/v2.1.0/api/codecs.html
https://www.researchgate.net/publication/356622607_OME-NGFF_...
https://www.biorxiv.org/content/10.1101/2023.02.17.528834v1....
i.munro (Ian Munro) August 15, 2023, 8:54am
On the original point about size I suspect your answer is compression.
I have found that a lot of image formats use lossy compression.
I’m not sure about CZI specifically.
AFAIK ome-zarr (sensibly IMO) uses only lossless compression.
Best
Ian
and: sebi06 (Sebi06) August 15, 2023, 1:42pm
Hi all,
for CZI we offer currentyl two compression methods:
JPG-XR
ZSTD
So we have lossless and lossy options depending on your needs. All of them are implemented in our C++, #NET or Python APIs and also supported from BioFormats.
Looks like CZI is a competitor:https://www.zeiss.com/microscopy/en/products/software/zeiss-...
https://github.com/cgohlke/czifile
http://www.physics.hmc.edu/~gerbode/wppriv/wp-content/upload...
CZI is basically JPEG compression and is probably missing the multidimensional compressor which is the secret sauce missing from the various formats. It also mentions licensing and legal so is maybe lost in the weeds and not open source.
Without more info, it feels like this is still an open problem.
My interest in this is for video games, since so many are multiple GB in size but are also missing this multidimensional compressor. If we had it, I'd guess a 10:1 to 100:1 reduction in size over what we have now, with no perceptual loss in quality.
Disk space is cheap relative to people-time and data-collection-time.
"Perceptual" is misleading- we're not just humans looking at jpegs. Many people are doing hard-code quantitative analysis for machine learning. better to downscale the images (2X, 4X, 8X or more) than to use lossy compression if you are truly low on disk budget.
To play devil's advocate, I think there might be an opportunity for lossy compression of images below the Nyquist Rate of the microscope:
https://svi.nl/Nyquist-Calculator
I only bring this up because I worked with industrial cameras at a previous job and even the best lenses only got us to maybe 1/2 or 1/4 of the resolution that the sensors were capable of. Microscopes should probably include an effective resolution in each image file's metadata. If they don't, then we have no data on how noisy the image is, so this is all a moot point.
Microscopists are savvy enough to match sensors and lenses, it's typically called "diffraction limited" (https://en.wikipedia.org/wiki/Diffraction-limited_system) based on the Rayleigh criterion (https://en.wikipedia.org/wiki/Angular_resolution#The_Rayleig...). I rarely if ever hear microscopists refer to nyquist limits. They would not overimage something at a higher magnification, or more correctly, they would be aware it was happening and either downscale or not worry.
Next, when you make a point like "we can get a 1TB SSD for under $50". It shows you're not understanding the domain microscopists work in. Their scopes typically cost $100K (with a camera for another $50K, and a stage that costs $10K). The cost of the storage system (or parking the data in S3) for 10 years is usually far less than the capital cost of the scope, or the operational expense of maintaining staff on site.
I believe the effective resolution can already be computed from the captured metadata, which includes the objective's numerical aperture, as well as the frequency used for imaging (many microscopes today use lasers to image with a very specific frequency to cause fluorenscence, which itself has a fairly tight output spectrum).
TBH, I think that the scientific community has huge blind spots, mostly due to its own gatekeeping. Some of the things like this that seem to be a struggle remind me of the design-by-committee descent of web development. The brilliant academic work of the 1990s has been mostly replaced by nondeterministic async soup and build processes of such complexity that we have to be full stack developers just to get anything done. All the fault of private industry hoarding the wealth and avoiding any R&D that might risk its meal ticket. Starving the public of grants, much less reforms that are actively blocked like UBI. Now nobody ever seems to step back and examine problems from first principles anymore. That's all I was trying to do.
Edit: I found the term for calculating the effective resolution of a camera with a high-contrast diagonal line, it's "slanted edge modulation transfer function (MTF)":
https://harvestimaging.com/blog/?p=1328
https://www.strollswithmydog.com/the-slanted-edge-method/
https://www.eckop.com/resources/optical-testing/resolution-t...
Except for the domains where nobody uses it because it ruins the data.
No, not in this case. This might be great for game files or whatever you are familiar with but lossy compression would basically mean to be mangling the data after going through all the effort of collecting it (I suspect you have only little experience with lab work?).
When doing experiments, everything is documented in full detail. Settings on machines and equipment, preparation steps, version numbers of software used for processing, …
You really don‘t want to lose information on your experiment.