The value difference is not because of invisible buffers! Marketing material and usable OS space are measured in different units, but the usable amount of bits in each value is exactly the same.
So 1TB = 1000*1000*1000*1000 bytes/1024/1024/1024 = 931GiB (or 0.9TiB)
I don't give a shit whether your operating system likes to show you disk usage in binary or decimal units, because I'm not talking about software at all. I'm explaining how the hardware is built. I especially don't need another person to try to explain what the binary and decimal units are, after I've repeatedly used both correctly.
Modern QLC has very low write endurance so those drives need to have spare space to use when it starts to wear out.
Second, in the SMART/drive health data, there's usually a counter tracking how many reserve/spare blocks remain usable as in not worn out and retired. That's a one-way counter; deleting data won't un-retire defective or worn-out blocks.
It's almost unheard-of for a drive to directly expose a realtime counter of how many blocks are currently unallocated, not being used to store data, and in the erased state making them available to accept new writes.
If you have two identical electric cars for the same price, one with a range of 350 miles and one with a range of 280 miles, which one would you buy?
If you have two identical phones, one with a 4300 mAh rated battery and another with a 3400 mAh battery, which one would you buy?
If you have two equally priced SSDs, one with 500GB capacity and one with 480GB, which one would you buy?
To enterprise customers the sales rep will explain "look, this device has a bit less space, but the write endurance number quoted here is much better, over 5 years and 5000 drives this will lower your TCO by X". In the consumer space you build your entire brand around one price-reliability tradeoff and stick to it. Usually the cheaper one with the bigger headline stat sells the most, and everyone else has to make up the lost volume with margin.
On most hard drives, the beginning of the logical block address space corresponds to the outer edges of the platters where the bits are going past heads with the highest linear velocity, so sequential throughput is higher than elsewhere on the disk.
A wise sysadmin buys a drive with more of that hidden internal space, instead of getting a larger drive with less hidden space. The logic inside the drive is much better at zeroing/out and background trim on that internal hidden space, than host-addressable blocks that are currently unused. That's because the drive has no idea if you're about to use them, while the hidden space is guaranteed to be always free.
In fact - a fun fact. Flash has been used as a Write buffer on storage arrays for a couple of decades.
The pricing of enterprise SSDs is... not exactly fair for the performance you get.
when you get fined $1mil/min by the SEC for downtime, and data loss or corruption can cost you in the hundreds of millions, or someone dying at a hospital, enterprise gear is cheap. reliability is the key and for what you are paying. now yes, there may be a specific consumer drive more reliable than a specific enterprise drive. so which do you buy? well the enterprise array vendor tested the crap out of everything in every combination and workload and environment, and picked one for you, and it comes with full support and SLAs that you can use to meet gov regulations.
performance for an enterprise drive is not something you usually consider, at all. in fact, did you know that when I quote a storage array, I can't even specify the drive type of vendor, and who knows what will get shipped? Ionly specify drive size.
your performance comes from all your workloads clumped together, spread over a thousand of these drives connected with infiniband, with dedupe and compression on the backend, and hundreds of terabytes of RAM. the perf of an individual drive is not relevant.
but yes, sticking it into an AMD server you bought on newegg when they had a sale is not its purpose, and is a very bad deal.
now when we talk about wise sysadmins, we're talking about guys who know their stuff, and do "important big stuff." not a guy at a small business ordering from newegg. and that guy - he shouldn't be coming up with any storage policies because he lacks the needed large-scale experience.
I sold a 1PB usable-effective (after 3x dedupe/compression) space storage array last year. It was $2mil, after a 65% discount. If one time in its 5year lifecycle the "enterprisey" stuff on that array prevents about 10 seconds of downtime, it paid for itself.
But for most of us, infrequent downtime is acceptable, and most single-machine downtimes are either automatically mitigated or have very limited impact. In that scenario, getting a prosumer SSD and over-provisioning it can be a sensible choice.
So, write me some software that's going to make all that work and make the storage highly available. make sure it works for everything.
you're exposed to solutions that make that work, daily. when you swipe your credit card buying condoms - you have no idea how much happens so your transaction doesn't get lost, corrupted, or errored out.
Maybe he doesn't, maybe he does - you don't know nor do I.
I'm pretty sure this is how IBM salesmen used to respond when confronted with those newfangled Unix systems which were starting to appear here and there, nibbling first, then taking larger bytes out of their market share. Instead of the litany of diverse systems they'd have thrown LPARs, SYSPlexs and ESMs around but in the end it still came down to the same thing: this stuff is too complicated to be left to amateurs. They were right, in a way... until those amateurs grew their wisdom teeth and took a large part of their market away from them.
Yes, "enterprise" stuff is complicated - often overly so [1] - and it has its place. This does not make it the only viable solution to these problems, something will eventually come up to eat your lunch just like IBM saw its herd of dinosaurs being overtaken by those upstart critters from the undergrowth. Maybe some smart software system which "guarantees" data reliability and availability without the need for "enterprise" storage devices? It wouldn't be the first time after all.
[1] https://github.com/EnterpriseQualityCoding/FizzBuzzEnterpris...
Under sustained long duration (ie: hours to days) of continuous full load performing small block writes, even the highest rated consumer SSD drives will begin to show significantly reduced throughput in comparison to pretty much any enterprise SSD drives.
Additionally, it's totally normal to see >=5 drive writes per day and 5 year warranties on enterprise drives. Consumer drives usually are rated significantly below 1 drive write per day and rarely for more than 3 years of warranty. If you're performing a lot of writes, you're going to need to replace worn out SSDs and that's not free in a business (remote hands, downtime, etc). So buying less durable storage has costs which don't show up on the original purchase invoice but do need to be factored in within a business setting. The warranty period is the best indication of how confident the drive vendor is about drive durability.
Spending 2+X per GB means rather than 5x the same X it’s 5 times 1/2 or less total space. And that’s before you consider write amplification issues with having dramatically less total SSD space.
Enterprise SSD’s have a few benefits, but they shouldn’t be your default choice for all servers.
Looking at my list of consumer SSDs and their warranties:
Samsung EVO 850: 2y
Samsung EVO 960: 3y
Samsung EVO 860: 5y
Samsung EVO 970: 5y
Samsung PRO 970: 5y
For comparison with spinning rust HDDs:
Seagate Ironwolf: 3y
WD Red (both EFAX & EFRX): 3y
Seagate Ironwolf and WD Red are not enterprise spinny hard disk drives, instead look at Seagate X18 or WD HC560.
I have a piece of logic in my indexing code that essentially transposes an ~100 Gb multi-value dictionary on disk. If you do this the naive way with random writes, the write amplification makes it a complete non-starter. All the caching layers and buffers fill up with completely disjointed 8 byte writes and it takes ages to write.
What I've ended up doing is to in an intermediate stage write the data to be written into a series of files, containing pairs of offsets and data (up to like 100Mb each); and then going over the files one by one and essentially evaluating them as assembly instructions.
Both passes have relatively good data locality, and despite essentially writing 2.5X as much shit to disk, it takes hours rather than weeks to do this operation.
I see this sort of issue in a lot of programs. As it is a dead easy problem to make in your program. You need to write something out you just splat it out somewhere. With a dozens 1/4/8/16 byte writes instead of one big write. Basically not thinking about how that data is getting into your files. Most of the time that is just fine and not that big of a deal. But as your data set grows or you want better perf you have to worry about it. I usually use something like filemon and can see what is going on. You can see the pattern where there will be a large stack of I/O with hundreds of very small read and writes. While SSDs are an order of magnitude faster than the older drives. They still have their command structure and kernel context switching you have to deal with. You in some cases want to minimize that as it can become a large portion of writing and reading data. As with most optimizations (in this case the drive is faster) we just moved where the bottleneck is (to the kernel and bus typically).
Actually, many storage arrays use battery backed RAM, not flash as a write buffer. Flash does not have the endurance needed to serve as the write buffer for a large storage array. Some products that use battery backed RAM for this purpose will dump the contents of that RAM onto flash. I worked on a messaging appliance that used that approach, albeit with supercaps to provide the hardware time to dump 4GB of DRAM onto a compact flash card. Supercaps were easier to monitor and maintain. There are also plenty of hardware RAID cards that use batteries for their write buffer as well.
Edit: there are also persistent memories like MRAM that don't need power to retain their contents which are used in this space as well.
smaller cheaper arrays (unity, pure, isilon) and HCI (nutanix, vxrail) that don't have terabytes of RAM - they have gigabytes, and pretty much always use flash as a write cache. In fact, I cannot name one that doesn't.
No one in enterprise storage cares about the indurance of flash. All that means is that twice a year, a vendor engineer comes out to replace a few flash drives under your support contract.
And flash has been used as a write cache behind a smaller RAM cache, for two decades, by all major storage vendors.
Yes and no. Large contiguous IOs are still faster to read than a bunch of random small sectors. You generally want your frequently read files/blobs to be split into as few operations as possible.
This avoids the ugly read/modify/write, as well as much of the behind the scenes magic for minimizing write amplification and shuffling blocks around.
Most controllers handle the block to page mapping problem well. They just mark the old block in the map as dead and write a new block on a new page. GC will later erase pages, hopefully prioritizing pages with the most dead blocks (subject to wear-leveling concerns).
But that same concept can be extended to partial block writes. It complicates the read process but there is no reason the controller can't coalesce multiple partial block writes into special partial update blocks and update the map accordingly. Basically write a record saying byte range (X,Y) was overwritten with just those bytes - no need to read the old block, merge the change, and write the whole block back out. Obviously you need to handle things like read chain thresholds (too many partially overlapping updates).
Let the GC handle merging the partial updates into the full block during idle periods. Again prioritize doing this for blocks on pages that are going to be erased - you had to rewrite those blocks anyway, turning what would be write amplification into a "free" write.
Like I said a lot of controller firmware doesn't bother but that doesn't mean it is impossible or unprofitable to do.
Eventually the solution really just become - get an SSD because you could throw thousand of requests and the results was fast enough.
Sequential reads need to be optimized to produce optimal performance for selected workload. That usually means applying a lower level of inline compression.
In some cases deduplication works better before, in other after, compression. Sometimes post-process dedupe is more suitable than inline.
Then there's erasure coding and data protection methods that are still being optimized for NVMe with sequential workloads, including random workloads which are being sequentialized to work better with latest flash media.
I would even say developments in sequential IO are becoming more important than random IO.