Lepton image compression: saving 22% losslessly from images at 15MB/s
blogs.dropbox.com
blogs.dropbox.com
For a quick test, I run it over ~1.3GB JPEG pictures I had locally, the finally result is 810M, that's 66% of the original size, very impressive considering it's lossless. It only deals with jpg file though, no png, no iso, no zip, no any formats other than JPG.
If someone can do this over video files that will be PiedPiper comes into real life.
WebP is very promising (also based on VP8) for lossless and lossy compression. Have you considered using it to compress PNGs in the same way Lepton is compressing JPGs? Odds are it wouldn't be bit-perfect though (despite being pixel-perfect).
Also interested in hearing about the tradeoffs between server-side decoding and client-side. Not to keep focusing on it, but WebP has native support in Chrome and javascript decoders for everything else.
I think it is a very exciting time for image formats with several promising new ones on the way (WebP, FLIF, maybe BPG but possible legal issues).
Can also recompress GIF and some RAW formats, but is much, much slower
Lepton is cool because it helps make existing technology a whole lot better, but what we actually need is a better image format.
You wouldn't see the same leap for videos because people have been working hard to make those great for a while while JPEG has been left to rot.
Sure, it's an inexact analogy because PNG is not a block-based DCT coder, but all lossless H.264 does is set the quantizer as absurdly high as it needs to go to losslessly encode a particular frame. Staring from a compressed source, this is not a good recipe for achieving a space-saving result.
It's mostly to let you fast forward, but there is a technical issue there. MPEG2 decoders aren't all mathematically identical, so what happens is the picture tends to drift away from the real thing after a while, and there's hacks like frequent keyframes and flipping the smallest DCT coefficient to get around it…
Maybe you could even go further and actively remove keyframes (in a fully reversible way, just keep a record of where they were)
I'm not sure how much you would save, and for losslessly archiving DVDs you might be better off creating a special format like Lepton for mpeg2
Not sure what you mean by the keyframe removal, though. Such an act would be lossy, significantly impair any P-frames or B-frames (unless you majorly modify them), and, it frankly doesn't sound very reversible. Mind elaborating?
There is no restriction that P/B-frames only reference I-frames, so you don't even need to touch those frames.
For conversion back to MPEG2 it would be ideal to detect or mark the original I-Frames, so you can convert them back (a simpler transformation). But you could also pick any random P/B-frame and convert it to an I-frame with little issue.
In practice you probably shouldn't do that without changing the file extension and the magic bytes to avoid confusion among users and poorly written software.
TIFF is showing its age at this point, with a some high-end scientific applications switching to HDF5. But PNG simply isn't featureful enough for this type of data.
edit: This [3] is a follow-up from 2014 to the article/study from 2014.
TL;DR: "We consider this study to be inconclusive when it comes to the question of whether WebP and/or JPEG XR outperform JPEG by any significant margin. We are not rejecting the possibility of including support for any format in this study on the basis of the study’s results. We will continue to evaluate the formats by other means and will take any feedback we receive from these results into account."
[1]https://blog.mozilla.org/research/2013/10/17/studying-lossy-...
[2]https://en.wikipedia.org/wiki/WebP#Criticism
[3]https://blog.mozilla.org/research/2014/07/15/mozilla-advance...
I could re-encode all of my JPEG photos with a better codec, but the problem is that then they have gone through a lossy compression scheme twice: once with JPEG, and once with the new codec. This could ruin the image quality on some photos, giving them nasty artefacts.
The newer codecs might even be hindered further by the fact that they are being asked to compress an image that's been JPEG'd. I don't think any of the codecs you mentioned are tuned to work their best with image data that already has been butchered by JPEG. They expect to be given the 'pure' original image. This may well cause the codec to perform worse than expected.
Using Lepton gives none of these risks, since jpeg<->lepton is lossless. I could throw away all the JPEGs afterwards, safe in the knowledge that if I want to go back, I can.
Also, it's not really true that JPEG has been left to rot. There are lots of programs available that greatly improve upon the basic JPEG compression, while still outputting a JPEG file that any compliant decoder can read. And there's an even simpler way to shrink your photo file sizes - just drop the JPEG quality level slightly. Most cameras and programs default to a very high quality value. Dropping the default by even a tiny amount can produce remarkably smaller files, and you'll probably never notice the miniscule image quality loss.
I suspect that JPEG is so popular because it is 'good enough'. Most people and programs don't care so much about squeezing out smaller files, so they'll keep using JPEG regardless.
> 74% smaller than lossless JPEG XR compression.
> Works on any kind of image
PNG was not designed to be used for photograph-like images. The rest of those were not designed to be lossless formats, the lossless version is just a tacked-on afterthought.
Very unsurprising to find a codec that can beat those.
It's also worth noting that there aren't many other lossless formats, so it's still a valid comparison. I'm sure neither TIFF nor RAW outperform FLIF either.
The lack of processing comparisons raises that question pretty loudly, and it's definitely important in a mobile world. There's more to performance than size
https://people.xiph.org/~jm/daala/revisiting/subset1_psnrhvs...
so when you use a better encoder like mozjpeg, JPEG is actually competitive with the newest formats, given enough quality
it gets slightly grainier results with blocking artifacts, but BPG and WebM look like they had a painting algorithm blur out all of the details
EDIT: 17GB now (my server is kinda slow) and it's still holding at 0.78x.
Saving almost a quarter of space for most images stored is something that truly gives a competitive edge. (I say most because people probably primarily have JPEG images).
Especially considering how many images are probably stored on services like Dropbox.
And they just gave it away to their competitors.
It's more likely that they have released it because of some profit-seeking interest. They are not charity.
Pretty much every modern company safes tremendous amounts of money from open source (Linux and upwards in the stack), and so the OSS community should rightly be considered a STAKEholder.
The share vs stakeholder obsession in the space of large companies and corporations represents a lot that's wrong with our current markets.
--
Also, the only upside to open sourcing this is getting other involved in development. I just tested on 10k images, and the promise both on compression rate and bit parity after decompression holds true.
Seems to be a pretty stable product, so the that motivation is probably only miniscule.
http://www.nytimes.com/roomfordebate/2015/04/16/what-are-cor...
Can you cite the law which you believe makes that obligation?
The reason you can't is because it's not actually a legal requirement and the reason is obvious: it's hard to say what the best interest is over all but the shortest time frame:
https://en.wikipedia.org/wiki/Business_judgment_rule
Dropbox's management might argue that they benefit more from open-source than they're giving away, that this kind of favorable attention will help them hire the top engineers who make far larger contributions to their bottom-line, that the pricing models are complex enough that this just doesn't matter very much, etc. Absent evidence that they're acting in bad faith, it's almost impossible to say in advance whether those arguments are right or wrong.
Releasing it freely allows it to work its way into browsers, allowing Dropbox to serve up the smaller images directly to users, saving money on bandwidth. It allows other users to contribute improvements. It markets Dropbox as a desirable place to work for engineers.
Not obviously less profitable that the opposite.
It's interesting that they don't consider it a large enough competitive advantage.
Either that or they are using this to attract engineers.
In google there is a regret that they didn't open source a lot of stuff, because it ends up being reproduced outside of google in some form. The open source companies have an advantage in hiring and overall advancement of their product by using the open source version whatever their internal version was.
You also see it in facebook's open source initiatives, like react, buck, haxe and so on.
What regret? A lot of the open-source clones of Google's internal projects come from Google themselves.
On a related note (I can't speak for Dropbox specifically) there are many engineers who desire their work to be open source for their own motivations, and when it's not a significant business risk to do so, nice companies will allow it.
So it works on two fronts, as a "hey, we'll let you open source your stuff" and a "hey, we've got people here who care about and contribute to open source."
[The configuration] "jpg_test2" by Jan Ondrus compresses JPEG images (which are already compressed) by an additional 15%. It uses a preprocessor that expands Huffman codes to whole bytes, followed by context modeling. http://mattmahoney.net/dc/zpaqutil.html
It's saved "multiple petabytes" of space.
Backblaze storage is $0.005/GB/Month = $5k/PB/Month.
The GitHub repo has 7 authors, perhaps costing Dropbox $200k/year each and taking most of a year ~ $1M to develop this system.
So this might pay for itself after 200PB*Months, assuming Dropbox's storage costs are the same as Backblaze's prices, and assuming CPU time is free. (TODO: estimate CPU costs...)
Of course, advancing the state of the art has intrinsic advantages, but again, it's interesting to look at the purely financial point.
You're correct ofc, download costs = 10 months of storage.
The speed quotes made it sound like client-side was a concern. Why would you go to all the effort of devising a new image compression format saving 20%+ storage and on the wire, and not have it decompressed client-side, especially when you control the client?
Lepton can decompress significantly faster than line-speed for typical consumer and business connections. Lepton is a fully streamable format, meaning the decompression can be applied to any file as that file is being transferred over the network. Hence, streaming overlaps the computational work of the decompression with the file transfer itself, hiding latency from the user.
In my humble opinion, dropbox's image gallery webapp is considerably faster than any other I've seen, especially when compared to imgur.
Dropbox is valued at $10b (regardless of what you think of that number, someone was willing to buy at that price). But they only have 1800 employees. If by some miracle 80% of them are engineers that is $7M per engineer. All this for a "pure software" business. No inventory, no manufacturing, they only recently started doing their own operations.
Investors would be telling them "more engineers" => "more software" => "better Dropbox". So now you have to go out and snap up engineers in Silicon Valley, which is very very hard. But you don't offer the perk of a Google/Apple line on a resume nor the compensation of e.g. Microsoft. What do you do?
Answer: you offer people the chance to work on cutting edge breakthrough technology. Not only that but it's all open source! This is a dream job for some people. They will turn down every other gig for this one.
It doesn't matter that PB/mo is chump change because an investor doesn't know that, they only know the engineering headcount. If by some miracle an investor does know this is a waste of time, then you just point to this HN post and observe how many developers are interested in this technology and how it is attracting developer mindshare that can be exploited for additional hires down the road.
I'm not saying it's rational–it's not. But I think there were strong incentives to greenlight a project like this, even if the actual cost savings were zero.
> For those familiar with Season 1 of Silicon Valley, this is essentially a “middle-out” algorithm.
I really wish that every single compression-related blog post would stop referencing Silicon Valley.
Did it really?
The wording almost implies that this is novel, but it is actually Gamma coding [1], which in the signal compression community is often called Exp-Golomb coding [2]. I wonder why this is not acknowledged, considering that they mention the VP8 arithcoder instead.
Eg. just tested on a small 10MP Sony ARW raw image - the raw file is 7MB, the camera jpeg is 2.3MB, and a tiff compressed with LZW is 20MB (uncompressed 29MB). The raw tiff run through lzma is 9.3MB. But either way, if ~7MB is the likely lossless size, if the JPEG is good enough, at ~2.3MB it's a pretty big difference, if we're talking petabytes, not megabytes.
(I'll get around to testing lepton on the jpegs shortly)
Apparently Sony uses some kind of lossy compression for it's files - I just tested with a jpeg2000 encoder on the same file above, and the size of the j2k file is approximately the same as the ARW: 7MB. Btw, the lep-file is 1.7MB.
Note that the uncompressed (flat) PPM file is 29MB as is the uncompressed TIFF - but simply running the TIFF through lzma reduces the size to 9.3MB. So ~7MB isn't that far off.
[ed: And while the lep-file was 1.7MB, shaving a bit off the original jpg, mozjpeg with defaults+baseline created a jpeg (at q=75, per default) 472k in size. Lep managed to shave a bit off that too - ending up with PPM->mozjpeg->lepton resulting in a 359K file. The (standard) progressive mozjpeg ended up at 464K.
This is not quite apples to apples, though, I think the comparable quality setting for mozjpeg would probably be 90 to 95 or so -- ending up around 1.6MB. But for this particular (rather crappy) image - I couldn't readily tell any difference.
Which I suppose is where https://github.com/danielgtaylor/jpeg-archive comes in.]
On top of that, most sensible raw formats only store 12 or 14 bits per pixel, instead of 16.
And then most are compressed, some losslessly and some lossily (like the infamous Sony format that packs it down to an average of 8 bits per pixel but does exhibit artifacts).
We tried JPEG2000, which was better quality per a file size, but the web worker decoder was slower than the JPEG one adding seconds to the total download/decode time.
EDIT: We're currently doing 256x256x256 (equivalent to a 4k image) on eyewire.org. We're speeding things up to handle bigger 3D images.
EDIT2: If you check out Eyewire right now, you might notice some slowdown when you load cubes, that's because we're decoding on the main thread. We'll be changing that up next week.
Most microscopy images are stored uncompressed or with lossless compression. But unfortunately this doesn't scale with newer imaging modalities. Here's two examples:
Digital histopathology. Whole-slide scanners can create huge images e.g. 200000x200000 and larger. These are stored using e.g. JPEG or J2K in a tiled BigTIFF container, with multiple resolution levels. Or JPEG-XR. When each image is multiple gigabytes, lossless compression doesn't scale.
SPIM involves imaging a 3D volume by rotating the sample and imaging it from multiple angles and directions. The raw data can be multiple terabytes per image and is both sparse and full of redundant information. The viewable post-processed 3D image volume is vastly smaller, but also still sparse.
For more standard images such as confocal or brightfield or epifluorescence CCD, lossless storage is certainly the norm. You don't really want to perform precise quantitative measurements with poor quality data.
jpeg-archive [^1] is designed for long term storage and you can still serve the images over the web. imageflow [^2] has just been kickstarted and looks really promising for use with ASP.NET Core.
mozjpeg is also showing progress and if FLIF takes off then that will be great. Scalable images would be fantastic. No more resizing and all the security issues that brings [^3].
[^1]: https://github.com/danielgtaylor/jpeg-archive
[^2]: https://www.imageflow.io
[^3]: https://imagetragick.com
* Implemented in C++ (-std=c++0x and -std=c++11 work, -std=c++98 doesn't work).
* Needs a recent g++ to compile (g++-4.8 works, g++-4.4 doesn't work).
* Runs on Linux and Windows.
* Runs on i386 (-m32) and amd64 (-m64) architectures. Doesn't work on other architectures, because it uses SSE4.1 instructions.
* Can be compiled without autotools (http://ptspts.blogspot.ch/2016/07/how-to-compile-lepton-jpeg...).
A picture should of course be worth 1000 words.
The idea was to train a neural net and build up a database of features (maybe on the order of 1-10 GB, or whatever is just small enough to ship) to estimate the missing details from downscaled and extremely over-compressed JPEGs. If it worked, I think it would also improve the quality of all the 10-20 year old images out there where the uncompressed source is long gone. Sort of a Blade Runner-style "enhance" tool, but of course it would only be filling in aesthetically plausible details.
[0] http://www.cs.northwestern.edu/~agupta/_projects/image_proce...
Given the recent usage of Rust for the implementation of Brotli compression (https://blogs.dropbox.com/tech/2016/06/lossless-compression-...) and that it's used for data storage (http://www.wired.com/2016/03/epic-story-dropboxs-exodus-amaz...) this somewhat surprises me.
The reasons given for Rust as stated in https://blogs.dropbox.com/tech/2016/06/lossless-compression-... would seem valid here too: > For Dropbox, any decompressor must exhibit three properties: > > 1. it must be safe and secure, even against bytes crafted by modified or hostile clients, > 2. it must be deterministic—the same bytes must result in the same output, > 3. it must be fast.
Now I do see http://huonw.github.io/simd/simd/ was being developed in August 2015, but it seems to be gathering dust of late.
I really do wish that Rust would provide nice alignment guarantees (eg 32 byte) without depending on customizing the allocator, and builtin, safe, SIMD instructions
Same could be said about C/C++. I'm guessing the answer is much simpler: the author(s) are comfortable with C++. And they probably don't deploy much rust code yet.
In Dropbox's case they control both client and server and can just compress/decompress in the dropbox client. Using lepton on websites wasn't really Dropbox's goal (but it would be cool if a sufficiently fast JavaScript library existed).
The backup tool bup (https://github.com/bup/bup) does this.
Here is a superfast implementation of rANS: https://github.com/jkbonfield/rans_static
Lepton uses the arithmetic coder [1] from VP8. Using arithmetic coding instead of Huffman encoding to get better compression was always an option in JPEG, but it has been historically avoided due to patents [2].
Compared to VP8-Intra, the compression used in lossy WebP, JPEG is missing the prediction step, usually called 'filtering' [3], which is the single largest contributor of WebP's compression outperforming JPEG.
Reading through the Lepton blog post, it seems they're using a different method of prediction, based on observations about typical gradients and correlations between AC and DC coefficients. VP8 uses a more 'traditional' approach of predicting your neighboring pixels, which was borne out of run-length encoding, but also very applicable to video's moving macroblocks. A comparison would indeed be enlightening.
[1] https://tools.ietf.org/html/rfc6386#section-7
[2] https://en.wikipedia.org/wiki/Arithmetic_coding#US_patents
[3] https://medium.com/@duhroach/how-webp-works-lossly-mode-33bd...
I presume someone (or likely many people) are working on exactly this.
Also there is a free book that contains descriptions of various common and exotic compression formats http://www.mattmahoney.net/dc/dce.html
Overall I think we are yet to see the full potential of deep learning unleashed on data compression. For example the neural network in cmix compressor is quite primitive compared to modern architectures. Someone will certainly find a way to do better than that!
However I am wondering about the two different EXEs available on the page (one has an avx prefix that I need to try and figure out). If anyone has info on that that'd be useful.
I could see this potentially being useful for schools and businesses doing quite a bit of digital scanning so I'm going to try running some tests using it and some images I think we have available somewhere.
The AVX version will be faster, but it requires a recent-ish CPU: https://en.wikipedia.org/wiki/Advanced_Vector_Extensions#CPU...
When running Lepton through JPEG photos downloaded from Flickr, I've got about 23.37% of size savings.
When running Lepton through JPEGs generated by mozjpeg (default settings: -quality 75, progressive) from JPEG photos download from Flickr, I've got 22.63% of size savings.
The mozjpeg output is about 3.83 times smaller than the original JPEG photo, on average.
The FLIR Lepton is thermal imaging sensor (http://www.flir.com/cores/content/?id=66257), and has been for a few years now. They really should have used a different name.
It's not illegal, it's just dumb. Sure, you could implement your entire C++ application inside the `std` namespace, and I'm sure it'd work fine, but you /shouldn't/. If you're going to start a project, at least google the name first.
We live in a world where it is common for blue sky to occupy a portion of the frame and green grass to occupy another portion of the frame. Since images captured of our world exhibit repetition and patterns, there are opportunities for lossless compression that focuses on serializing deviations from the patterns.
I can't read the article due to technical constraints, but understand that e.g. JPEG has a lossy quantisation pass followed by a lossless encoding/compression pass of the result of the first stage. If they're reproducing bit-identical result to the input JPEG, it must be a (very good) optimisation of the latter stage. [How'd I do?]
If that's the case, then not only could we rely on the results of everything we've decompressed so far to use for prediction (which is like one-sided image in-painting), but we also could store a few bits of semantic information (e.g. from an image-net-based CNN, from face detection) about the content of the original image before re-compression, and use that semantic information for prediction as well via some generative model. All of this would obviously be trading computation for storage/bandwidth, but it this seems like an exciting direction to me. Again, nice work.
As for having the mega-model that predicts all images better: well it turns out with the lepton model out you only lose a few tenths of a percent by training the model from scratch on each images individually. We have a test case for training a global model in the archive (it's https://github.com/dropbox/lepton/blob/master/src/lepton/tes... ) That trains the "perfect" lepton model on the current image then uses that same model to compress the image (It's not meant to be a fair test, but it gives us a best-case scenario for potential gains from a model that has been trained from a lot of images) and in this case it doesn't gain much, even in a controlled situation like the test suite.
However the idea you mention here may still be a good idea for a hypothetical model--but we haven't identified that model yet.
That would actually compress rather well. A hard to compress random image would not look like TV snow (white dots, with space between them), but rather randomly colored dots that are continuous in the image.
https://dsp.stackexchange.com/questions/2010/what-is-the-lea...
Consider images of slides for a presentation that are text on a flat background. If you know the value of the pixel just to the left of the current pixel, then if you guess that the current pixel will be the same, you will be right most of the time. This is obviously not true for random noise. Consider a really simple compression scheme where a pixel is stored as a single bit of a 1 if it is the same color as the previous pixel, and the color is stored as usual, but with an additional 0 bit prepended to the color. When you guess wrong, you pay a tax of 1 bit, but when you guess right you save N-1 bits where N is the number of bits per pixel.
For random noise, this will grow the input quite a bit, but for simple flat-shaded graphics it will shrink the input quite a bit.
In this case they are making use of the fact that JPEGs store a certain mathematical formulation of a picture. It turns out JPEG doesn't store that mathematical formulation very well, so you can squash it loselessly into a better formulation, then later turn it back into the original JPEG.
It's always good to see new techniques out in the open.
https://github.com/tjko/jpegoptim
I'm hitting the limits in my OneDrive account and jpegoptim seemed to reduce my photos quite a bit.
I'd say 22% for Lepton and 5% for jpegoptim, based on fading past memories of mine.
>For the standard test image the new LenPEG 3 compresses the image so efficiently that data storage space is actually freed on the computer right up to the entire capacity of the storage devices
> LenPEG compresses the image into a file of minimal size: one bit.
Is the image Lenna? If yes, delete all other data on the computer's storage devices. If no, proceed to the next step..."
Part of ffmpeg open-source library