Minify and Gzip (2022)
blog.wesleyac.com
blog.wesleyac.com
I don't have a site that sees a ton of traffic, and I know what some of my direct competitors spend on IAAS (and how long their pages take to load). I'm orders of magnitude below it - hence, I do not spend much time investigating optimizations. I can name a list as long as my leg as to where I bet I can win quite a bit on optimizing for IAAS cost or UX speed. Business-wise I decided it's not (currently!) worth investigating.
One of the more intriguing ones on my list is that I __do not minify anything__. My reasoning is that post-gzip, who cares. It's such a tiny difference. Yes, turning e.g. class names or whatnot in the CSS, or local identifiers in javascript, into short letter sequences means even post-gzip it'll be smaller, but by so little. The reduction in file size when e.g. eliminating all unnecessary whitespace is really, really tiny.
For any site that isn't trying to optimize for 'millions of hits a day!' levels of traffic, why _do_ you minimize all that stuff? It makes the developer experience considerably shittier and the gain seems a few orders of magnitude too small to even bother with it.
* assuming your website doesn't get instantly evicted from cache.
** modern browsers have mechanisms for caching the parsed/compiled code for a given JS file hash, I think, but I don't know how it works
Identifiers are often turned into atoms (which makes table lookup performance not really depend on how long they are) but for cases where identifiers are being looked up in dictionaries as raw strings, making them shorter could also improve runtime performance for your JS. I'm not sure how much this actually matters in practice though - I've seen dictionary key lookups show up in profiles here and there but it's usually not the main bottleneck.
In a Java minifier ages ago, I modified it to suffix sort the constants pool and saw a little over 1% further size reduction. Given that the pool is half of the file, that’s >2% on the pool.
The reason for suffix sorting was that entities in bytecode, JS and CSS often have end of statement markers, so suffix sorting gets you slightly longer runs than prefix sorting.
> It makes the developer experience considerably shittier and the gain seems a few orders of magnitude too small to even bother with it.
Does it? It's very rare that I interact with the minimised code
I think the OP may be referring to situations where production needs to be debugged, in which case navigating production code usually becomes a massive pain in the rear.
That said, if one can get away with not needing to do more work and still have a performing site then how little the difference is remains irrelevant.
That's a problem in its own, though. While I agree that source maps can and often do solve the minification problem/confusion, it can also introduce another "vector" for something to go wrong and make production debugging a pain. Like everything, it's about tradeoffs. I've worked on some apps where, yes, minification really did matter, but minification has also become a default even when it's questionable whether it's even called for. Whereas just shipping the code as-is is less likely to introduce a pain point when debugging, source maps are more likely to introduce them when the source map somehow doesn't work or, in rare circumstances, can't accurately represent the runtime code when transpilation is emulating the original code by doing something almost completely different somewhere down the chain.
Although this is an "it depends" thing, like everything else, I lean on the side of not minifying things in 2023 when both gzip is adequate and the time it takes to parse and run the main code is not the source of performance problems. But the last time I brought a similar topic, a bunch of HNers got mad because "starving people in Africa with flip phones and 3G tho", so maybe I'm still objectively wrong years later.
Especially things like removing as much whitespace as possible or replacing true/false with !0/!1 just don't really save a lot of space. The only two I found to make a significant difference were removing comments, and shortening identifiers (and usually with most of the savings being in "removing comments", depending on the amount of comments and code style).
At the scale of Wikipedia (~10 billion pageviews a month), even a single byte saving can mean terabytes worth of bandwidth saved.
The problem here is that you've switched from proportional thinking to arithmetic thinking in the middle of the thought.
In most cases, if you started with proportional thinking, you should stick to it.
Example of the fallacy: 0.1% of my salary goes to avocado toast, which might not seem that much, but over ten years, that's a thousand dollars!
The fallacy here lies in the fact that a thousand dollars is insignificant over 10 years, but the statement makes it feel significant by switching from proportional to absolute.
"The avocado toast fallacy"?
For example, suppose an engineer who is paid $200k per year can spend a week to implement an optimization that reduces costs by 0.1%. If that is 0.1% of something that costs $50 million per year, then that's not a bad use of time.
Is it not? I don't know. I don't run a multi-million dollar business. So this is an honest question. 0.1% of $50 million is a meagre $50000. If someone is running a business that is spending $50 million per year on something, are they really going to care about a small $50k saving per year?
Spin it around: is it worth it for a sales person who makes $200k per year to spend a week to land a $50,000 per year contract? I think it potentially is, for the same reason, even if the organization as whole makes hundreds of millions of dollars a year.
The thing is, say it takes me an hour to shave off ten bytes in the final build. Value my time per hour at 200$, and the saving of that ten bytes over a year at Wikipedia scale at a thousand dollars, and even investing five times that amount into yak shaving would still be net profitable.
But does it really move the needle anywhere where it matters? If you are serving a million bytes in every response but you manage to trim a single byte from it, I don't think it is going to move the needle anywhere.
Sure it will save a terabyte when you have served a trillion responses. But in a trillion responses you are saving only one terabyte out of an exabyte of responses. That's not much!
In many cases wikipedia probably isn't even paying for bandwidth due to peering agreement.
As far as performance goes, i suspect whether a single byte matters depends a lot more on if the extra byte results in the same number of packets or not.
Remember kids, if someone tries to use unusual time interval (like months for bandwidth-related stuff), they just want to make a small number look big
In either case, they said a single byte saved. Any amount of minification is going to save more than just a byte.
It’s also incredibly low effort to do so.
const index = "Hello HN".indexOf("H")
console.log(index);
>> 0
index == false
>> TRUE
index === false
>> FALSEWe used to write things like "LET A=COS PI" instead of "LET A=1", or "SIN PI" instead of 0, because keywords were represented as a single byte and numeric constants as five bytes so two keywords represented a useful saving of three bytes.
Back then it made a huge difference but now that modern computers have hundreds or even thousands of kilobytes of RAM free it's a bit pointless.
Jokes on you, now I'm running IntelliJ, DataGrip, Firefox, and Slack and thousands of kilobytes of RAM free is a rare sight to see.
Eight Meg And Constantly Swapping.
It might be better to have slightly bigger payload on the wire to save time after decompress.
func("hello world");
func("hello world");
Then naively it seems that can be optimized to this:
let s = "hello world";
func(s);
func(s);
However once compressed, the second result is usually larger! It's basically because the second example adds the `let s =` part which is new unique content it has to store.
So if you want to minify JavaScript to a shorter version uncompressed, it's good to transform the first to the second. However if you want to minify JavaScript to the smallest compressed size, it's actually better to do the reverse transform and turn the second case in to the first! But then you get in to tradeoffs with parse time with a longer input, especially with very long strings - so I'm not sure any of today's minifiers actually do that. (Edit - turns out Closure Compiler does: https://github.com/google/closure-compiler/wiki/FAQ#closure-...)
fwiw I don't bother minifying anything on my personal site. I realise that complex sites can have large enough payloads to make it worth it, especially for Javascript, but for the typical developer blog it's likely overengineering in the extreme. If I were to apply any postprocessor it would probably be a prettifier, for the HTML, since the template substitution results in mismatched indents. But then again anyone looking at the DOM is likely to be doing it in the dev tools, which already renders a projectional view, rather then the original source
Instead of trying to make the code as small as possible, they will try to make it as redundant as possible. For example by reusing names. For example "message="hello";print(message)" can be minified into "hello="hello";print(hello)", which may compress better than "a="hello";print(a)".
and the second version compresses better, as my intuition already thought, without knowing how gzip works, because you still need a kind of identifier for repeated strings, and how much shorter than 1 byte can they be.
Also, there seems to be a 40-byte minimum size on zstd; Making the original example significantly smaller (or larger with repeated patterns) all yielded 40 byte files. I had to change it to e.g. print(hello);foo(hello); bar(hello) to get any difference between using "x" and "hello" as an identifier.
Firefox has also seemed to wander back and forth between View Source just reuses the full Dev Tools and View Source is a disconnected view from Dev Tools for whatever reason. At the moment the big difference flag seems to be if you install/use Firefox Developer or not.
(Allegedly, one of the reasons is that some non-technical users have strange reasons that they View Source from time to time and they get upset if it opens Dev Tools because Dev Tools is just too confusing. Mac Safari has that interesting dance to ask it to opt in to showing Dev Tools. I seem to recall IE11 did something similar that if you ever opened Dev Tools for more than a few minutes it then took over for View Source, but otherwise used the older and simpler View Source from ancient IE history. I'm surprised that if the reason is "it confuses non-technical reasons" there isn't an opt-in flag in Chrome or Chromium Edge. The "Firefox Developer Edition" is a separate SKU is also an interesting way to approach that.)
https://github.com/google/closure-compiler/wiki/FAQ#closure-...
https://closure-compiler.appspot.com/
Should be set to:
// @language_in UNSTABLE
// @language_out ECMASCRIPT_NEXT
for it to work as expected. $ wget https://repo1.maven.org/maven2/com/google/javascript/closure-compiler/v20230502/closure-compiler-v20230502.jar
[…]
$ java -jar closure-compiler-v20230502.jar
The compiler is waiting for input via stdin.
The Closure Compiler still is a Java project. It just gets primarily distributed via NPM these days. location /js/small {
gzip off;
}
location /js/ {
gzip on;
# Use the .gz file if it exists
gzip_static on;
}https://blog.cloudflare.com/this-is-brotli-from-origin/#:~:t....
The usual reason to stick to gzip level 4 is that compression gets extremely slow at marginal savings.
I would expect that cloudflare caches compressed data so they pay the compression cost only once.
The CDN is still useful in this case of course, both for ddos mitigation and because it can cache the common stuff like your website's stylesheets.
However it happens, I think if you're being real, it would take A LOT of these micro-optimizations multiplied by A LOT of cache misses to add up to anything important.
What I got out of this is, unless you're looking at the specific bytes in the network tab every time you edit your code, just minify/gzip and forget about it.
This website has real life data on the matter for popular libraries:
* https://github.com/privatenumber/minification-benchmarks
Compare the trophies indicating smallest size for Minified versus Minzipped (gzip). Generally the smallest minified size yields the smallest minified+gzip size, but there are some notable anomolies outside the range of statistical noise. It is not practical for a javascript minifier to take a compression algorithm into account - it would blow up the minify timings exponentially.
I go through a list of examples that I found where it does or it does not matter whether you "optimize" it for size or not. I also theorize (I do not have enough knowledge to verify my claims though) why this is happening:
https://medium.com/@fpresencia/understanding-gzip-size-836c7...
In your CSS use a variable for the property that changes, and put the fallback in it too.
In that way the next person to work on the code can see what changes.
In your media queries only have CSS variables, as locally scoped as possible, and name them after the property with -- at the front.
On this way it is self evident what the media queries are about.
For example:
p { color: var(--color, black); }
@query {
p { color: chocolate; }
}
And have the queries after the CSS for legibility.By writing CSS this way, your source will be compact and easy to maintain.
Because of this you won't need to mangle it with CSS mangling tools or bother with CSS minifiers. Your production time will speed up immensely.
Also, style the elements, then add classes and IDs if you have to. Stay away from div's, use all the semantic elements with CSS grid and you will get much better results.
And as the brotli people will insist on telling you, it's good even without that dictionary: https://news.ycombinator.com/item?id=27163981