This is conceptually similar to what OP does by storing the (numerical) difference between the words. Also, if you have a list of numbers that aren't random, they generally compress better if you turn it into a list of the differences between the numbers.
A simple compression algorithm (miniLZO is apparently 6KB compiled) might be small enough and save enough bytes with compression to make it worth it for OP.
If you want to skim, check out the EXAMPLE section toward the bottom.
29.44% - Original sorted list just compressed with zstd
22.45% - Matching prefix characters from previous word replaced with space
19.38% - Matching prefix characters from previous word removed
Years ago I worked on a J2ME (Java2 Mobile Edition) application that had no business being attempted given the very small archive files allowed. We did it anyway and it actually worked pretty well. We very quickly hit the max file size however, and every feature request meant first shrinking the existing code base to make space. First we had an intern fixing bugs in the code minifier we were using, especially around deleting unused (usually debug) methods. For some reason they rejected on archive size, not payload size, so while I started out doing 'honest' work with shrinking the binary, I had spent a lot of time in college noodling with compression algorithms so my eye was eventually drawn there.
Those were in the days when I could still read JVM assembly code, and shortly after I started thinking about the compression, I realized that the constant pool entries start with the type and then the size of the entry. So while our minifier made the reasonable assumption of sorting the constant pool by type and then alphabetically within it, because most of the constant pool was strings, and strings are variable length, it was hit or miss whether the header would be treated as a run or just Huffman encoded (the fallback). If I suffix sorted, then all but the last string in the pool would be followed immediately by the header for the next string, increasing the average run length.
This ended up knocking almost a kilobyte off of the archive size. Depending on your perspective that sounds like a little or a lot, but in our case each feature cost about 500 bytes, so that change pushed the cliff I was walking toward out almost a month (and slowing the growth rate), just by changing a sort algorithm.
I filed a ticket with Sun about this, but as it turns out they already had the dense archive format in flight, and within a couple months my observation was moot because the dense format can compress constant pools across and entire archive, not just a singe file. That was at least an order of magnitude better than what I had.
It's quite likely a lot of the files we use have similar problems in them. Off the top of my head, JSON compression probably would be much higher if we treated it as unordered, and did more aggressive minification particularly for JSON-at-rest. Sorting sibling keys by value instead of by name for instance.
Similarly, having to contain the decompression code in the measured result size and it being a relevant contribution is something that only applies in some use cases of compression.
That's why people still write for the Z80: it's a fun toy.
For genetic data, HapZipper beats general-purpose compression. https://www.ncbi.nlm.nih.gov/pmc/articles/PMC3488212/
So, yes, actively researched, but you've got to pick a specific task that makes sense. Even small niches are viable; I made a task-specific compressor to strip the essential numbers out of a remote sensor report to make it small enough to squirt to a satellite.
It kinda didn't matter from 2010 on. But one area I've written my own specific "compression" methods in, for the last few years, has been in shipping data in and out of webworkers (in-browser or in Node). This is where there's still enough of a performance penalty on a lot of devices for sending 1MB that in use cases where you're spawning lots of workers to run long tasks, it makes sense to trade time to compression for a smaller transfer size.
Eg. You work for a doorbell company and the boss says "yo, can we make our doorbell have 6 tunes instead of one, because our competitors are doing that. No, we don't want to change microcontroller".
https://blog.transitapp.com/how-we-shrank-our-trip-planner-t...
Working on GPUs, I see many, and work on some task specific compression ideas as part of my job. The compiler has it’s own ways of compressing code & debug info. The hardware has it’s own ways of compressing textures. A recent feature we built on my team is a compressed encoding for adaptively subdividing curves. All of these things have the primary goal of reducing memory bandwidth, which in turn increases the speed of computation because memory is so frequently the main bottleneck.
I had to do this in the early '80s. The alternative was scrapping the boards and redesigning them to allow double the EPROM size but that would have been a lot more costly than writing the decompression routine and manually compressing the strings. It would also have delayed delivery.
Anyway, thinking about the transposing idea some more: this would effectively split the word in to 26² = 625 "buckets" of three-letter suffixes. What we could do to make those still compress decently after transposing is look for shared suffixes in multiple buckets, and ensure they get grouped together in the same order before transposing. This would result in short runs in those suffixes, squeezing some more compression out of it.
... which should also work really well for implicit delta coding.
Hmm... you know, the basic concept here shouldn't be too difficult to implement and try out out, thanks for the ideas! :)
If you have a means of doing RLE that performs otherwise, I'd love to understand how it works.
FYI, turning it into 12972 by 5 and Brotli compressing achieves 15,093 bytes, which is less than if you first turn the data into an ASCII trie then Brotli compress that (14,180 bytes) (Source: https://github.com/adamcw/wordle-trie-packing#all-words).
My only claim in this post is that anything can be pre-sorted if you want to achieve some better DS compression but of course, that means you have to have some map to undo the sorting after you decompress it (and obviously the utility only exceeds the computation time if the documents are longer than the associated dictionaries). The person who wrote the gameboy wordle compression did it by necessity, which is beautiful, and the way people used to do things when you had to fit them into tiny structures like that, so, huzzah! to that person.
But yeah, the observation that you could handle 40% reduction on the first two characters is a good clue.
So much has been missed in lossless image compression, along the same creative lines. We humans can look at an image of a red-black-yellow Cardinal bird sitting on a green-gray stem in the middle of a forest, and basically compress it in our mind in a way you'd have to throw thousands of CPU hours against. If you knew you only had to consider a red bird on a green background, you would have a whole different domain-specific strategy for compression; the amazing thing about our minds is that we can devise that compression strategy from the inputs, remember the strategy for that specific set, and then recall our own compression strategy well enough to decompress the data later.
There was, actually, an attempt in the 90s to do something they labeled "fractal compression" which was more or less an attempt to come up with a lambda function for a particular image; extremely CPU expensive to compress, and might or might not be lossy depending on the goal, but the salient thing was that the compression strategy was unique for each image. That didn't really work out as a commercial concept for a whole host of reasons. But it's one of those corners of extremely clever pre-modern code that might be worth a bundle to revisit now.
So if you wanna code something really, really fun -- consider a fully adaptive compressor that comes up with a specific strategy for each general sub-batch of use cases.
If you want to start a company doing that, let me know because I literally just came up with this idea 12 seconds ago.