Windows file system compression had to be dumbed down
blogs.msdn.microsoft.com
blogs.msdn.microsoft.com
You'd think so, but allow me to point at my steamapps folder saving 20% disk space even with the not-very-good compression that NTFS offers. If it could use a better algorithm and a block size of 2MB instead of 64KB, it could nearly double that.
Hard drives have always been growing in size. We have always been in a 'post-file-system-compression' world. But people want to do things like fit on an SSD, so compression continues to be useful.
I just wish it didn't ultra-fragment files on purpose.
Note though that some controllers do compression to reduce write amplification (e.g. the SandForce controllers used to do this). So, by using filesystem compression, you might be increasing write amplification and shortening the SSD lifetime.
Now let's compare the two scenarios:
1. With filesystem compression: you could store twice the amount of data (~512GB). However, since the controller is not able to compress the data, all cells are completely full.
2. Without filesystem compression: you can store 256GB data. However, since the controller can compress the data, only half of the actual cells are used.
Suppose now that in both setups, the SSD is nearly full and we rewrite or perhaps extend some data. In the former case, the controller has to scavenge for partially-used cells to combine and erase. In the latter case, the SSD controller still has 128GB of pristine cells to write the changed data immediately.
Of course, compression was always somewhat of a cheapskate solution (and one of the reasons the SandForce controllers are not liked much). Higher-end SSDs would just have more storage space than logically addressable to have some leeway when the SSD is almost full.
You compress it to 1MB, SandForce controller fails to do anything because already compressed, writes it to flash, and you wear out 1MB of physical flash memory.
You don’t compress it, SandForce controller compresses it to 1MB, writes it to flash, and you wear out same 1MB of physical flash memory.
How (1) might be increasing write amplification and shortening the SSD lifetime compared to (2)? Aren’t they nearly equivalent?
Or if Microsoft could sit and define a new version of NTFS in which they change LZNT to LZNT2 with better compression ratios and different requirements for modern systems :)
I think it's a pure anachronism at this point: it makes more sense to me in the era where small drive sizes meant a lot of professional users relied on external drives and “just upload it” wasn't going to fly with a 9600bps modem, not to mention hardware costing more and failing increasing the need to move drives between machines simply to get back in service.
The article was pretty clear that the context they had in mind for this requirement was servers in a data center, not your home machine:
> Without that requirement, a hard drive might be usable only on the system that created it, which would create a major obstacle for data centers (not to mention data recovery).
Keep in mind they still thought they'd be targeting Alpha processors as late as the Win2K RC's.
That said, it would still be one hell of a weird edge case to need to take a drive out of an x86 Win2k server, drop it into an Alpha Win2k server, and still care about its contents (vs wiping it for a newly provisioned host). But when you are writing OS filesystems, you have to care about edge cases... especially edge cases that may apply to thousands of racks worth of machines.
I can imagine the reverse being more common though - I worked at a shop around that time where we had a handful of very expensive Alpha NT 4 (and VMS and...) boxes, and a lot of x86 NT boxes. I could imagine the magic smoke being let out of an Alpha and having to drop the drive into an x86 box for data recovery.
Or until the host machine dies, in which case they have to be readable in another machine.
(Happened to me recently: the laptop's motherboard stopped working, I used its disk with an external case and an older operating system version while I waited for the warranty replacement.)
Even fastest compression algorithms like LZ4 need Core i5-4300U @1.9GHz in order to get 385MB/s [1]. You'd need a pretty powerful setup to have it keep up with SSD speed and you also need to be mindful of heat generated. Also, it would be pretty useless if the volume's encrypted.
Now we're living with SSDs that can do 2GB/s and no CPU can decompress that quickly.
I'd totally buy a machine with a state-of-the-art CPU paired with one or two FPGAs that can be programmed as accelerators for crypto, compression, etc.
Looking at these numbers you linked to (which seem to be in megabytes, it seems to me that decompression speeds could keep up. And I know write speeds on SSDs are a lot slower than the spec'd numbers, so the compression write speeds look plausible to me too.
The general idea is that modern CPUs tend to be instruction starved and sit idle because they are waiting on the slow memory buses that connect everything.
Am I missing something?
Filesystem implemented deduplication is probably a much better long-term strategy than filesystem compression, but both online deduplication and online compression can impose hefty costs on frequently written data.
ZFS and BTRFS can "get away" with implementing better compression because they're copy-on-write by default. In that model, compression can actually reduce the cost of the read-modify-write, and clever use of journaling and caching can reduce the cost of fragmentation.
* - that I'm aware of, I won't pretend to know the state of the art.
The OS/file system still has to keep metadata about where stuff is stored, and many files fragmented into many tiny slices will bloat that metadata, meaning you may spend a little more time needing to read that metadata from disk (and space storing it in the first place) and spend a little more time to "re-assemble" these slices into something virtually continuous the user land expects. More metadata may also mean disk caches become full more quickly, leading to metadata being evicted more often.
This may all not matter on a beefy laptop/desktop/server, but your underpowered, memory-challenged ARM SoC NAS may notice a bit.
With RAM, you don't just say "Give me the 4 bytes starting at 0x10A6E780". Instead, a whole block of memory around that address is loaded into the cache (pausing execution in the meantime), and then the value is pulled from there.
If the next value you read is in that same block of memory, you don't need to spend time loading a new block of memory into the cache.
In practice, on hardware I have tested on, HDD's have a 10x or more slowdown from random IO relative to sequential, while SSD's have a 3x slowdown. Your hardware is definitely not mine, take this with a Hummer-sized grain of salt.
Sequential access is faster even on random access devices because the hardware itself can cooperate to predict what-is-next, and prepare the data. Read-ahead caching is common in enterprise devices and software (see: SQL products) and CPUs, SSDs, and RAM can all participate, but only if the next locations are well-known.
There are translation layers between virtual memory mappings and RAM or storage devices, sure, but to look at two examples in detail:
1. CPUs have instructions for retrieving the next bytes, and cache lines are often 8-16 words (32-64 bytes) in modern processors. The optimization of putting data sequentially is significant enough that it's a common optimization for game developers, for which the latency of repeated cache misses can blow the CPU budget on a frame. (One source: http://gameprogrammingpatterns.com/data-locality.html)
2. For solid state disks, the SSD writes in large page-sized increments, and though I understand it's possible to read smaller chunks, if the SSD has RAM on-board (common) why not read the whole page into a LRU cache? Moreover, as SSDs themselves have a virtual mapping between LBAs and the actual data locations, why not retrieve the next LBA too? Reads aren't destructive, SSDs are implemented as essentially RAID devices over multiple flash chips for which multiple reads incurs only a small additional cost, so now sequential reads are automatically cached and the SSD improves performance over random LBA access.
An interesting note: even before SSDs became the standard, Microsoft changed the defragmentation algorithm to ignore chunks larger than 64MiB. I can't say for sure they chose an "empirical best" option, but apparently even for spinning disk, that provided a large enough runway to get most of the benefits of sequential access. For SSDs, that size is almost certainly smaller - I would guess between 512KiB and 2MiB - but still relevant for performance.
Causing fragmentation to avoid fragmentation is a bit silly, especially when it means that none of the space saved is contiguous, leading to other files fragmenting too.
Disk space is cheap. Player time isn't.
And, no, my time as someone who plays video games is much cheaper than today's SSD storage. I wondered a few days ago why my Windows disk was so full. Video games. Literally 50% of it was video games. If I could save a couple GB on that, I'd be willing to give up a couple more seconds on every level reload.
Flash storage is cheaper than it's ever been, and my time means more to me than it ever has. If I could spend a few dollars to have more gigabytes of fast storage, that sounds like a bargain.
I don't know about you, but I have better things to spend money on than storage that could otherwise be slightly further compressed. I'd rather wait a couple seconds longer and burn my money on upgrading the graphical fidelity of my machine.
Better yet, I could buy 3-10 games on Steam on sale for that money. Do you still think that's a worthwhile upgrade? Or would you rather wait another 3 extra seconds for every level reload in your game?
If you have more money than time, you could always disable compression. My problem is that I can't enable any sort of good compression and I'm broke.
3 seconds extra between levels, plus 20 seconds during the initial engine load, plus 5 seconds closing down the app, etc...Time adds up. If I've got half an hour available to play a game, I want to get in+out fast. My SSD can stream data faster than my CPU can decompress it. Why would I want to handicap the hardware that I paid good money for?
On the other hand, 15 years ago, I remember shifting data around to fit on my 20GB hard drive, doing the minimum install and running games from the CD to save disk space. My priorities would've been a lot different, then.
You can stream music. You cannot stream effects; the timing of effects is such that you can't afford a frame's delay or people will notice it's off-kilter. And, as it happens, effects are precisely what are loaded uncompressed; almost every game engine's streaming-music feature uses MP3, OGG, or whatever.
But it's entirely possible that you could decompress some of that audio into RAM beforehand if you're not confident every device can play that audio compressed immediately. Or store back a cache file temporarily. Or just not even worry about the CPU because the vast majority of users are GPU-limited rather than CPU-limited.
Also cpu time to decompress audio is microscopic it might actually be faster especially if the consumer has a hard drive to read 1/10 of the data and decompress.
Additionally there is probably a smart middle ground between keeping the data as small as possible and 50GB of raw audio.
To me it is. I delete games when I'm done with them. The disk is reusable.
Load time is normally IO bound, not CPU. If your game has the typical IO bound loading, compressing stuff will make it load faster not slower.
Given this is happening while a game is running, it means that the loading code has to compete for CPU and memory with all the rest that is going on in the soft realtime system that is a game.
Not a reason, but I haven't payed attention to sound cards since my Creative SoundBlaster 16. I assumed that some amount of progress/feature-creep had been made on them.
If you told me I would've believed that every sound card these days had an embedded mp3 decoder as well as some 3d audio components.
Going back about 15 years ago: https://en.wikipedia.org/wiki/Sound_Blaster_Audigy
I've got some version of that series of cards sitting in my closet. I feel like that's about when sound cards really peaked. The hardware was mature, CPUs weren't fast enough to generate all the effects games wanted on the fly, surround sound was popular, etc. It made a lot of sense to offload that stuff to an external card.
Compare that to now: I've got some Chrome tab doing audio decoding, but my task manager reports 0-1% CPU use for each of my tabs. Audio decoding is fast enough to be done easily on the CPU. Ditto for the effects and such that are in use now. The equation shifted.
While it might make sense for consoles where you need every last ounce of CPU power and have the audio stored on 50GB Blu-Ray, on PC it's usually just a leftover from the console world.
https://www.reddit.com/r/skyrimmods/comments/59u0iw/the_skyr...
Wait, you run the Steam apps from a compressed folder? Doesn't that kill performance?
Regardless, I have long felt that the right way for Steam to make this work would be to have, in addition to the "Download" and "Delete" tools, an "Archive" button that moves it to a specified path (by default, on the install drive, but configurable to another drive or a NAS) and compresses it - with whatever compression they want.
I want my Steam apps uncompressed on my SSD. I don't have room on my SSD for hundreds of gigabytes of Skyrim textures that I haven't played in 6 months. But if I delete the app, then I have to wait hours (and cause Steam some expense) to download it again.
Not everyone has multiple drives or a NAS, but a local archive would definitely be useful.
Also useful if you don't have fast internet for your gaming PC, but access to it somewhere else.
No. In fact, it may improve performance, by virtue of transferring less data from slower I/O devices. Yes, slower, even if it's a SSD.
I'm skeptical of this actually being the case in any real-world scenarios. There have been a number of tests of running games from a RAM disk vs an SSD, with precious little difference in load times.
The key is the data sent to the cpu and decompressed, makes up for the stall from hitting memory or i/o. Comparing ram vs ssd is the wrong comparison to make, with both you're hitting stalls due to memory. You want to compare reads of uncompressed versus compressed with the note that (and i'm just making numbers up with this analogy as i'm about to sleep), 900KiB of compressed data in, 2MiB of data out. 1.1MiB bonus and yes I'm assuming huge compression but for times your cpu is idle it makes perfect sense.
And yes, lz4 compression on things like movies still helps. I shaved off over 200GiB on my home nas with zfs.
But it seems game developers didn't get the memo...
Except they totally did. You'll find that most game assets are already compressed, and compressed using an algorithm that allows for reasonably quick decompression using the CPUs of the day.
Do you think winzip users consider those points cited by MS when zipping files?
More like a case of MS wanted to suckle the Fed's law enforcement wallet by introducing insecurity through convenience.
Feels to me like nobody talks of nibbles anymore, maybe because we have ample memory and storage and can usually afford to waste some bits for the convenience of byte alignment. (It's half a byte, or 4 bits.)
Disk compression is interesting because Microsoft originally included it already in MS-DOS but lost a lawsuit brought by the company behind a popular utility called Stacker: http://articles.latimes.com/1994-02-24/business/fi-26671_1_s...
That was the first time Microsoft got in hot water for bundling features into DOS/Windows (the web browser would be the straw that broke the camel's back).
Oh if only they knew what would happen
The patent industry always wheels out the small inventor as PR when defending the system. As soon as noone's looking they use the same system to keep small inventors out.
In short, in the early 90s Amiga had configurable per-file compression algorithms. There were CPU-optimized versions of almost all of those codecs, so someone using an ancient 68000 could interact with files compressed by a PPC. I could pull a drive out of my fast system and give it to a buddy with an old, slow CPU, and he could either 1) live with the reading speed penalty, 2) decompress each file one time and then use the unpacked versions, or 3) recompress each file with an algorithm more friendly to his system.
I don't think the Windows OS team is dumb by any stretch. I do think they might have been hampered by NIH syndrome, and weren't aware of (and likely couldn't care less about) how these problems were solved on other OSes.
I think NIH qualifies as a form of stupidity.
Reads, in general, will be "boosted" when the filesystem is compressed.
You get an (almost) free boost by reading and extracting compressed data on the fly into memory/cpu. ie. read 4mb's off the disk but it expands to 8mb's (or more!) in memory, so read performance is elevated.
Writing of course, is slowed.
For some server loads, disk compression still makes a lot of sense - making the claim "We live in a post-file-system-compression world" a little dubious.
Even on the fastest CPU's, writing a file that must be compressed first, will always be slower than the same file on the same hardware, but not being compressed before written to disk.
It may not be magnitudes slower - depending on the data, hardware, and algorithm, of course - but it will incur some write penalty. So, with disk compression, you get a write penalty and a read boost.
I do not agree with your statement that compress + write is always slower, though. The analysis for cost/benefit of compression is the exact same logic for both reads and writes, but with different formulas based on how fast you can compress/decompress and read/write blocks. Let's imagine that your disk takes 10ms to write 64KB and 20ms to write 256KB. Then if compression of a 256KB block down to 64KB takes 5ms, writing is faster with compression done first. On the other hand, if compressing it takes 50ms, then writing it raw is quicker.
Most of Windows 10's boot, when Fast Startup is enabled, doesn't involve the typical system folders at all. Effectively when you shut down, they collapse the userspace, and store the kernel/services/etc into the hibernation file. When you "boot" the computer, the hibernation file is re-mapped into memory sequentially, and the user is prompted to login.
Compacting the OS definitely saves space. But I don't understand how it would reduce boot times from a technical perspective, since IO isn't even reading those folders during a default boot on Windows 10.
You'll get a performance boost by reading compressed data off the (slow) disk and expanding it in (fast) memory.
This can effectively increase read speeds greatly (reading a 2mb file off disk but it expands to 3mb's for example, that 3rd mb came almost "for free").
You are completely right for this case.
tldr: NTFS does a better job of avoiding fragmentation than FAT, but both need defragging.
If you're blasting highly compressible data to disk, compressing it on the fly can, in some circumstances, have a net bandwidth greater than just writing the data to disk (and greater than the disk alone is capable of). Yes, you incur more CPU load, but it's a net win. It's not universally true, YMMV etc.
(And yeah, natch, insert recitation of the Liturgy of the Optimizer as a charm against the appearance of a certain Knuth quote.)
The counter to my counter, of course, is that specialized silicon for compression can easily keep up with even these speeds. In fact, Sandforce SSD controllers build in compression to boost read and write speeds!
> Well, okay, you can compress differently depending on the system, but every system has to be able to decompress every compression algorithm.
But that's not how this works. You don't design an entirely new algorithm for every possible performance, you create one or two algorithms and tweak them. Then make one or two decompressors that can handle any compression setting. Simple example: lz77 (used in deflate/gzip) whose decompressor is extremely fast regardless of your compression settings.
> Now, Windows dropped support for the Alpha AXP quite a long time ago
Why does this sound like "... and they changed nothing"?
> you can buy 5TB hard drive from [brand] for just $120.
Sure but we also have more data to store. If you wanted 100 games in the 80s you'd need what, 1MB storage? I'm just guessing, I'm too young to know that. Now that'd be what, 1TB? It might be cheaper but I'm just saying, storage prices going down does not change that compression is a good idea.
> many (most?) popular file formats are already compressed
Before they were all about data accessibility and recovery on different/damaged systems but compression doesn't help that. Now this is a good thing? Additionally, binaries and libraries are not compressed; many of my documents are just text files; and databases still benefit from this a lot. (But who has a database on their computer? Anyone who uses a web browser and email program.)
> We live in a post-file-system-compression world.
I still think it's a fine idea.
> Tags: History
Lol
I can't remember any other non-niche architecture with 8*2^n word size that shares this (mis-)feature. (I suspect that Cray 1 and it's derivates also share this, but that probably counts as niche architecture)
It's fine for OpenBSD or NetBSD to support it, but Microsoft has not supported it in a long time.
[1] http://superuser.com/questions/135594/what-is-more-important...
Keeping the decompressor robust would easily solve the drive portability concerns, making the data readable (even if not at maximum speed) across machines and architectures.
It's a lot more expensive to install that disk it into an ultrabook/mba.
says a guy working for a company shipping system with >6GB (often 10GB in 100K individual small files!) of redundant NEVER EVER touched data inside WinSxS. Data that cant be moved off the main drive without serious hacks (hardlinking). This fits in the general pattern of indifference to user hardware. Another one (my fav) is non movable hiberfil.sys that MUST be on primary drive, cant even move it with hardlinks. There goes 16GB of your SSD for a file that gets used ONCE per day.
This is what happens when you hire straight out of college programmers and put them on top of the line workstations.
C:\WINDOWS\system32>dism /Online /Cleanup-Image /AnalyzeComponentStore
Windows Explorer Reported Size of Component Store : 8.18 GB
Actual Size of Component Store : 7.78 GB
https://technet.microsoft.com/en-us/library/dn251566.aspx
"This value provides the size of files that are hard linked so that they appear both in the component store and in other locations (for the normal operation of Windows). This is included in the actual size, but shouldn’t be considered part of the component store overhead."
2 still leaves up to 5GB of redundant garbage in WinSXS, things like multiple versions of random DLLs nobody ever uses, 360MB for ~11 versions of 'getting started' package consisting of same movie files, color calibration data for obscure scanners on bare minimal install etc
Turning drive compression omits this directory while happily compressing files in /system.
No, you communicated before with somebody else.
> still leaves up to 5GB of redundant garbage
No, if you read that example, there is less than 1 GB which is stale (the diff between the first and the third number):
>> Windows Explorer Reported Size of Component Store : 4.98 GB
>> Actual Size of Component Store : 4.88 GB
>> Shared with Windows : 4.38 GB
As I've said, you intentionally didn't quote your machine's third line but I'm sure it's not a 5 GB difference.
Moreover, that difference can be purged if necessary. Which you can also do on your machine with just a single command, per link on the same page:
https://technet.microsoft.com/en-us/library/dn251565.aspx
There is a trade-off, the probable case why nobody does it unless really necessary:
"All existing service packs and updates cannot be uninstalled after this command is completed"
> powercfg -h off
It's Microsoft. Would you critical of a pigeon for defecating in flight?