Could anyone explain why LZ77 is preferred by implementors versus LZ78?
It seems important for compressibility to prepare the data for maximum self-similarity, in addition to the LZ algorithms (as evidenced by the sort in this article). Could someone point towards a good modern summary of the approaches or heuristics?