1. They're overdone
2. They're uncreative
3. They're unoriginal
4. They're a crutch
5. They're assembly-line writing
6. They're used by tabloids to appeal to supermarket zombies
7. They've helped turn respectable mags into tabloids (or are a symptom of it, not sure which came first)
8. They're psychologically manipulative (I don't know how, I just know it)
9. They work
:)
I do not mind lists, but I cannot stand lists that are seperated to pages for inflated page counts!
If the original title begins with a number or number + gratuitous adjective, we'd appreciate it if you'd crop it. E.g. translate "10 Ways To Do X" to "How To Do X," and "14 Amazing Ys" to "Ys." Exception: when the number is meaningful, e.g. "The 5 Platonic Solids."
Using 50MiB from a tarball I had lying around, best of 3:
prog time(s) size(%)
gzip -1 1.703 38.7
gzip -9 14.670 35.1 (worse than xz -1)
bzip2 -1 7.714 34.0
bzip2 -9 8.035 31.4 (worse than xz -1)
xz -1 6.278 31.3
xz -9 42.445 22.7 (best compression)
lzop -1 0.670 46.1 (fastest)
lzop -9 18.144 38.0http://partiallystapled.com/~gxti/misc/2010/06/25-shootout.p...
http://partiallystapled.com/~gxti/misc/2010/06/25-linux-2.6....
I'm insisting because PAQ is a really incredible collection of compressors. Far more intelligence in it than in the LZ* bunch.
This guy did a not so rigorous analysis but it is mostly OK: http://changelog.complete.org/archives/931-how-to-think-abou...
A little bird told me there's a better algorithm in speed/compression rate coming soon ;)
xz format seems to have this feature, but: $ xz --list xz: --list is not implemented yet.
EDIT: also, like gzip but unlike bzip2, you can stream data through xz (at an insignificant penalty in compression ratio).
Right now, CPUs are fast enough that LZMA is a realistic prospect. It boils down to what the extra CPUs will cost vs the cost of the storage saved (incl. the "cost" of space in the datacentre, the administrative overhead of more hardware, etc).
This could be achieved by using Javascript and fetching data as JSON, but it seems to be something that would be very useful and beneficial if it were standardized.
The server supplies a pre-created dictionary (template) which gets cached, and further requests get a diff to that dictionary as a response.
I've only found Google search using it on the server side, and the browser penetration is fairly low, but it seems promising. It makes requests very fast.
I've also noticed that Google search serves image thumbnails as data:// URLs directly in the original response. They are going a long way to ensuring each page loads entirely in one request.
Really expensive algorithms aren't that much better to offset cost of sending template logic to the client.
You could probably save some processing time by integrating compression directly into server-side templating - take advantage of the fact that some parts never change and keep them pre-compressed or at least cache some statistics about them to aid compression.
However, that page does not give the same picture of widespread adoption that the Wikipedia article does. That use was what interested me: I'm surprised that it only came to my attention when I was trying to recall the name of the 7z command line utility p7zip.
Second, I am reading Of Mice or Men for the first time this summer. It is new to me!