BlurHash: Algorithm to generate compact representation of an image placeholder
blurha.sh
blurha.sh
It requires minimal javascript support-- just enough to build an img tag with the appropriate data url, and renders just as fast as any other tiny image. Upscaling blurs away the imperfections.
That 607B of fixed header includes some compression settings that don't need to change between images:
- 132B of quantization tables, saying how much to divide every coefficient by-- for subtle details the difference between "128" and "129" doesn't matter, and you can just send "2 * 64".
- a 418B Huffman table, which stores the frequencies of different coefficient values after quantization-- small values are much more common than large ones, given the divisors. You can have smaller image data with properly tuned tables, but that only matters for large images. For 250B of image data, it's a waste.
- Miscellaneous other details that are fixed, like the file type magic headers, the color space information, and start/end markers.
You can dig in more by putting a thumbnail like https://i.imgur.com/M0EsJH1.jpg into this site: https://cyber.meme.tips/jpdump/
And I figured, yeah, that's still a lot of data, and it's massively misusing several different tools. Why not create something that is actually designed for the task?
This is really disingenuous-- a DBA should not be a gatekeeper of what gets put into a database. Storing 200 bytes of extra data per row costs you absolutely nothing. Even if you assume 100 million photos, that's 20GB of data. That's a little over $2 of EBS cost per month. Who cares.
Don't think that's the correct adjective here.
But parent is right that the choice of what shouldn't be coming from your DBA; they should be more interested in how.
That isn't the end of the conversation; if the response to "we shouldn't do Y" is "it'll be a pain if we can't do Y", then you can explore solution space together, where that solution space includes both trade-offs to support doing Y and alternatives to doing Y. Use experts for their expertise, rather than the equivalent of "I don't care, make it happen".
Or, in short: yes, of course the DBA doesn't unilaterally dictate the architecture, but maybe listen to their expertise anyway.
Gotta love organic db schemas with long history...
This is horrifying, but at least not as horrifying as the public sector database I once had to work with that predated proper support for foreign keys, so there was a relationship table that was about 4 columns wide and must have had more rows than the rest of the database combined.
Even the database they had moved this schema to struggled with that many join operations.
[edit: and that almost all network traffic is now L7 routed because of gatekeeping firewall administrators in the late 90's, so everything had to go over :80 or :443]
Seems kinda rude to ignore the motivations and incentives of the people who are responsible when problems happen.
Yes devs and infrastructure are often at odds professionally by design. Making it personal ignores the reason that different parts of the app have different owners that must agree to make changes.
“Those damn developers won’t just let me push my hotfix to master and make me have to do a code review when our severs are getting hammered.”
"Why are we creating new table and foreign key relationships for this single bit of data?"
"Oh, because we aren't allowed to add any columns to <the appropriate table>"
"Wait. None?"
"None."
(Only the last conversation I had of this sort involved MySQL, which was notoriously bad at adding columns to live tables)
The cost of 5 extra joins for all business activities is going to outweigh whatever you think you have going on with that original table.
I think in some ways the situation was improved by expanding developer responsibilities to include these tasks rather than having a dedicated role. When you have to deal with the consequences of your own decisions instead of someone else being responsible, some activities won't be done at all, while others will go more smoothly. Without the database being personified, it's more of a doing something for somebody (else) instead of doing it to somebody.
To be fair, isn't it the DBA's job to gatekeep the database? The application programmers and business managers don't have the in-depth understanding to know how every individual change will effect the database. Of course they should follow best practices, but it's not really their job to know it inside and out - that's what the DBA does.
You don't know that 200 bytes per row would cost nothing - that's the DBA's job to understand it's cost. And we aren't talking about the amount of data, but without an in-depth understanding of what tables they were adding it to, how they're indexed, cached, etc, then you don't know the real cost. That 200 bytes might mean less can be stored in memory, if this is a frequently fetched table then that could be bad.
That's not to say the DBA made the right call, but that we don't have enough information to know if it was.
That is not my experience at all. On images less than 100kb, webp adds an overhead, and doesn't reduce the file size at all. A fixed header is also harder to achieve.
Besides this very small bug, very useful piece of technology!
You could save a step by just sending which image to download and use the filename as the hash you use to render the result, but this algorithm requires you store the relationship between the hash and the image somewhere (unless I am missing something obvious).
[0] source code: https://github.com/woltapp/blurhash-python/blob/7469c813ea64...
2. Which disallowed characters? The dictionary appears to be "0123456789ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz#$%*+,-.:;=?@[]^_{|}~". No slashes or null bytes, which are the common disallowed characters on server platforms. It does have characters your shell cares about, but that's only a problem if you're not quoting.
Point of order though: I don't think that's right for MacOS, though you're certainly right for Windows. HFS+ lets you use any Unicode symbol including NUL (because the filenames are encoded Pascal-style). I don't know the details how how APFS does it, but it also appears to support any Unicode code point, though it additionally mandates UTF-8. [0]
[0] https://developer.apple.com/library/archive/documentation/Fi... grep "filenames"
The name “:.:” can’t be used.
Try using a name with fewer characters, or with no punctuation marks.This is because classic MacOS used : as a directory separator instead of /, and this behaviour is preserved in the UI, but obviously not in the underlying layers.
Anyway, chances are you are already storing a filename in your database, so you would need a second field for the blurhash string.
assets/example.png/LEHV6nWB2yk8pyoJadR*.7kCMdnj
assets/example.jpg/LEHV6nWB2yk8pyoJadR*.7kCMdnj assets/example.png#LEHV6nWB2yk8pyoJadR*.7kCMdnj
assets/example.jpg#LEHV6nWB2yk8pyoJadR*.7kCMdnj
No need to send the extra data to the server for each request.With this charset: 0123456789ABCDEFGHIJKLMNOPQ RSTUVWXYZabcdefghijklmnopqr stuvwxyz#$%*+,-.:;=@[]^_{|}~.
I imagine you could tweak that to replace problematic chars with safe substitutes.
\ / : * ? " < > |
which means "*:|" would need to be removed if you want to support that use caseCoincidentally, I've just finished some work on a project[1] that is in the same space (identifiers for images).
For the reasons you pointed out, I found Douglas Crockford's base32[2] encoding to be a good fit.
(83 is about as many safe characters as you can reasonably find, and it allows some nice ways of packing values together.)
I wonder how the efficiency compares to just encoding on the bit level.
https://github.com/woltapp/blurhash/blob/master/Algorithm.md
All but the first component of the Fourier transform. (The first component is the average of the data.) The term comes from electrical engineering, but Fourier transform has lots of applications also outside of electrical engineering.
You can't really use them as filenames because they won't be collision free, too. They're not meant to be an identifier, but instead a compact representation of the image that can be stored in a database.
Maybe that shouldn't be allowed, but it is :P
The JS implementation of Gradify [1] did something similar but by using CSS linear-gradient which is even better in terms of performance. I wonder if the same could be implemented here.
[1] http://gradifycss.com/ [2] https://github.com/QueraTeam/gradify
If code helps, here's how it looks as a React Hook in Typescript:
https://gist.github.com/WorldMaker/a3cbe0059acd827edee568198...
(I offered the code to the react-blurhash repository in its Issues.)
But yes the biggest performance gains aren't in pure, static IMG tag usage scenarios, the biggest performance gains I saw were in combo with CSS animations, and that was something important to what I was studying. As Blurhases are useful loading states this seems a common use case to me of having a Blurhash shown/involved in things like navigation transitions, and it seems pretty clear browsers have a lot more tricks for optimizing static images involved in CSS Animations than they do for canvas surfaces.
I was looking for a way to better amortize or skip the initial render as well. It should be possible to take the TypedArray `decode` buffer and directly construct a Blob from it, but I couldn't find a MIME Type that matched the Bitmap format Blurhash produces (and Canvas setImageData reads) in the time I've had to poke at the project so far. As I mentioned, memory pressure was a performance concern in my testing, so I'd also be curious about the performance trade-off of paying for an initial canvas render and getting what seems to be a "nearly free" compression to JPG or PNG from the GPU in converting that to blob, versus using a larger bitmap blob but no canvas render step.
I'm guessing that it's approximately the size of a single image, possibly quite a bit more.
Edit: Nevermind, derp, https://blurha.sh/blurhash.01ae00eea611ada5e9c7.js contains all of the JS of the entire website.
The algorithm itself fits into a few k.
Of course, my first clue should have been that it was also the only Javascript file.
(I also appreciate that this is probably me getting overexcited about this, since odds are the JS is cached, and in the case of mobile apps it doesn't matter much.)
https://pkg.go.dev/mod/github.com/fogleman/primitive https://primitive.lol/
I've used it in a project to generate vector thumbnails that are under 1KB that I can just base64 encode and send along with the general text of a page - so the page renders with the text and low-fi images, and then I can load the hi-fi ones as necessary.
The lo-fi images are art in themselves, sometimes I preferred them to the actual images :-P
You could probably devise a much smaller data representation too, to get the size down.
Like cinemagraphs, but unique and alive in a nice way.
<canvas width="100px" height="100px" imgsrc="hirez.jpg" blurhash="abcd..." />
Also, it's really simple! Go ahead and create a library for it! The algorithm is tiny, and easily ported, so it's quite fun.
Via https://news.ycombinator.com/item?id=20339442 (1 comment).
That's the right design decision, because rendering a progressive jpeg repeatedly is rather expensive - as well as redecoding the whole jpeg, you also need to re-render any text or effects overlaying the jpeg, and recomposite the frame. And then you're gonna have to do all that work again when a few more bytes of progressive jpeg arrive from the network...
IIRC the limit is most commonly 6. So if you have 7 or more images the 7th will not even partially download until the 1st has downloaded in full.
With this method, the small blurred images will load before any of the larger ones so you should see them all relatively immediately, absolutely immediately if encoded in the HTTP stream as data rather than being retrieved with an extra HTTP request, before any of the larger images start to transfer.
- Upload any image
- Now upload a transparent png
- It does not clear the previous image, layering them together
- It hashes the combination picture. Super neat!
That's not a "hash", that's an encoding. Hashes are one-way.
Edit: to the downvoter, please have a look through https://en.wikipedia.org/wiki/Hash_function