JavaScript-Is-Weird as a compressor
github.com
github.com
% npx google-closure-compiler -O ADVANCED --js output.js
(()=>{})["co"+(1/0+[])[4]+"structor"]("co"+(1/0+[])[4]+...
Not pasting the full thing. But it reduces the output.js file from ~118 KiB to ~9.92 KiB, which is pretty good!There is technically not much stopping the compiler from inferring that 1/0 === Infinity, recognizing (1/0+[])[4] is free of side-effects, and eventually concluding its safe to substitute the whole expression with "n". Google Closure already has optimizations for string concatenation, so if it were able to perform an optimization pass with Infinity, then it would also be able to emit the string "constructor" instead of "co"+(1/0+[])[4]+"structor"
117708 output.js
$ uglify-js output.js -c unsafe | nodeHello world!
$ uglify-js output.js -c unsafe | wc -c
2740
$ uglify-js output.js -c unsafe | cut -c 1-60(()=>{}).constructor("conso"+"".constructor["from"+(()=>{}).
$ uglify-js output.js -c unsafe | gzip | wc -c
192Do the various JIT engines use an intermediate representation (IR) and what does it look like?
That way it can start executing quickly and not waste time on compiling/optimising things that only get run a few times
Had no idea about this command!
Looking at the table at the end, I'm not surprised at all that the "weird" obfuscated code is ~2000x the size of the original source.
But I am surprised that that the gzipped weird code is still ~25x the size of the original source, as opposed to ~0.25x for gzipping the original source.
After all, the amount information in the weird code should still ultimately be approximately the same as the original source code, right? Or maybe double or something like that. I'm very surprised it's twenty-five times as much.
The only reason I can guess is that the "weird" process results in information structures that are represented in an extremely hierarchical way, and gzip is built for stream compression, and is unable to find/represent/compress hierarchical structures?
And if that's the case, it makes me wonder if there are any compression algorithms which are able to handle that better? That might not be based on "dictionary words/sequences" as much, but rather attempting to find "nestable/repeatable syntax patterns"?
If the "weird" version is ~2000× the size of the original source, gzip would need a sliding window of ~2000× the size (about 64MB) to obtain equivalent compression.
I'm surprised people keep reaching for gzip for this sort of comparison. LZ4 / zstd compress better at the same compression speed, and zstd and brotli compresses much better. Brotli is also supported in all modern browsers.
There's basically no reason to use gzip in new software. (With the exception of compressing for the web - at least until zstd makes it into browsers.)
Taking a very high level view of this: the compressor can't know about the rules that are the difference between a .weird file that is valid .js and a byte sequence that is not and neither can it know the difference between bytes that are valid .js and look like .weird and are part of the subset that are encodings of a .js file and those that are not.
The compressor builds a more or less accurate model of observed probabilities, but even the most optimal encoding would still have to keep escape hatches around for all other bit sequences. Those escape hatches are not free.
New, fancy, and really slow compressors with their contexts, arithmetic coders and neural networks will give a much better result, maybe close to ideal, but it will take orders of magnitude more time and use up GBs of memory doing so.
I don't think I've ever seen such a dramatic difference in compression rates before. It's fascinating to see.
2) Thanks for leading with an example of a negative result! That's what any researcher faces every day, unlike what gets published, after all
Asking ChatGPT is just like looking up documentation, just faster ...
ChatGPT? None of these. Its a giant probability calculator that wants my prompts and its own answers to foster itself. This also nurtures its inaccuracies and makes it easier to confidently lie. I hardly find any reasonable use for ChatGPT other than quick completion suggestions for at a maximum of 2-3 lines, because thats what its designed for.
If you don't like ChatGPT, that's fine. But it's clearly useful to many people, and you aren't entitled to feel superior for not using it.
I will break the HN spirit but you kind of wanted to say that there. I didnt imply that at all, nothing near that. The dataset is just, imo, too big to give extremely accurate (or more accurate than human-thought) answers for specific questions.
I went that route since reading the CLI parameters was something I needed to do, but wasn't actually part of the issue I was trying to solve. I could have spent a few cycles of googling+testing+fixing/refactoring, but it was a trivial task (that I didn't remember how to do off the top of my head) and would have distracted me from the main task.
Of course at that point you're probably more interested in a common binary format, and should start thinking about wasm instead.
The compression algorithm knows that 165427-165427 is really 165427-same, but does not know that it's really 0 and the same as all the other things that resolve to 0.
There must be a lot of similar things like that in this particular case that relies on knowledge of the rules of js.
I guess it's tempting to think next about adding ai to a compressor so that it could know the actual rules of js and refactor it, but I just think of the tragedy that is jbig that does OCR as part of the compression, except, gets it wrong, and the original data is lost forever without a trace. And it's built right in to some scanners and happens before the compressed output even leaves the device. The user never sees anything else. There is no better reference uncompressed version anywhere if the user did not know about the problem and override some default settings.
Obfuscation is a serious threat to the open web, and things like fingerprinting can be incredibly invasive.
Web browsers typically only support static “prettifying” (ie auto indent). I’ve seen websites probe for chrome extensions, canvas and all kinds of APIs. Deobfuscators are often not enough to restore legibility. (I assume disassemblers are similar, but I’ve never tried.)
I would love to have a trace/intercept/breakpoints of any external APIs called in order to restore a sense of control over what code websites run. Ideally, integrated in the browser. With WASM gaining popularity, this will become (much) more important, imo.
gzip expgz zstd expzstd original
263120 421278 252449 389004 985084
So unlike in the article, expansion + compression is at least better than the original. The ratios are (smaller # better; the opposite of how most compression algorithms advertise their performance, but what the original article used): 0.27, 0.42, 0.25, 0.39. gzip and Zstandard aren't a lot different in either case. Whatever patterns that the Javascript weirdifier uses is less obvious to compression algorithms in general than just bitwise substitution.Here's the program if you want to look for bugs: https://go.dev/play/p/wwNXVzO2TO-
19648 dommy-2.0.js
28092 dommy-2.0.weird.js.xz
115023 lodash-4.17.15.js
138808 lodash-4.17.15.weird.js.xz
7114 modernizr-custom.js
11148 modernizr-custom.weird.js.xzIf your results disagree with your premise in unsurprising and obvious ways - please start your article with that so I can stop reading it.