Is sequential IO dead in the era of the NVMe drive?
jack-vanlightly.com
jack-vanlightly.com
Sequential reads need to be optimized to produce optimal performance for selected workload. That usually means applying a lower level of inline compression.
In some cases deduplication works better before, in other after, compression. Sometimes post-process dedupe is more suitable than inline.
Then there's erasure coding and data protection methods that are still being optimized for NVMe with sequential workloads, including random workloads which are being sequentialized to work better with latest flash media.
I would even say developments in sequential IO are becoming more important than random IO.
The value difference is not because of invisible buffers! Marketing material and usable OS space are measured in different units, but the usable amount of bits in each value is exactly the same.
So 1TB = 1000*1000*1000*1000 bytes/1024/1024/1024 = 931GiB (or 0.9TiB)
I don't give a shit whether your operating system likes to show you disk usage in binary or decimal units, because I'm not talking about software at all. I'm explaining how the hardware is built. I especially don't need another person to try to explain what the binary and decimal units are, after I've repeatedly used both correctly.
Modern QLC has very low write endurance so those drives need to have spare space to use when it starts to wear out.
Second, in the SMART/drive health data, there's usually a counter tracking how many reserve/spare blocks remain usable as in not worn out and retired. That's a one-way counter; deleting data won't un-retire defective or worn-out blocks.
It's almost unheard-of for a drive to directly expose a realtime counter of how many blocks are currently unallocated, not being used to store data, and in the erased state making them available to accept new writes.
If you have two identical electric cars for the same price, one with a range of 350 miles and one with a range of 280 miles, which one would you buy?
If you have two identical phones, one with a 4300 mAh rated battery and another with a 3400 mAh battery, which one would you buy?
If you have two equally priced SSDs, one with 500GB capacity and one with 480GB, which one would you buy?
To enterprise customers the sales rep will explain "look, this device has a bit less space, but the write endurance number quoted here is much better, over 5 years and 5000 drives this will lower your TCO by X". In the consumer space you build your entire brand around one price-reliability tradeoff and stick to it. Usually the cheaper one with the bigger headline stat sells the most, and everyone else has to make up the lost volume with margin.
On most hard drives, the beginning of the logical block address space corresponds to the outer edges of the platters where the bits are going past heads with the highest linear velocity, so sequential throughput is higher than elsewhere on the disk.
A wise sysadmin buys a drive with more of that hidden internal space, instead of getting a larger drive with less hidden space. The logic inside the drive is much better at zeroing/out and background trim on that internal hidden space, than host-addressable blocks that are currently unused. That's because the drive has no idea if you're about to use them, while the hidden space is guaranteed to be always free.
In fact - a fun fact. Flash has been used as a Write buffer on storage arrays for a couple of decades.
The pricing of enterprise SSDs is... not exactly fair for the performance you get.
when you get fined $1mil/min by the SEC for downtime, and data loss or corruption can cost you in the hundreds of millions, or someone dying at a hospital, enterprise gear is cheap. reliability is the key and for what you are paying. now yes, there may be a specific consumer drive more reliable than a specific enterprise drive. so which do you buy? well the enterprise array vendor tested the crap out of everything in every combination and workload and environment, and picked one for you, and it comes with full support and SLAs that you can use to meet gov regulations.
performance for an enterprise drive is not something you usually consider, at all. in fact, did you know that when I quote a storage array, I can't even specify the drive type of vendor, and who knows what will get shipped? Ionly specify drive size.
your performance comes from all your workloads clumped together, spread over a thousand of these drives connected with infiniband, with dedupe and compression on the backend, and hundreds of terabytes of RAM. the perf of an individual drive is not relevant.
but yes, sticking it into an AMD server you bought on newegg when they had a sale is not its purpose, and is a very bad deal.
now when we talk about wise sysadmins, we're talking about guys who know their stuff, and do "important big stuff." not a guy at a small business ordering from newegg. and that guy - he shouldn't be coming up with any storage policies because he lacks the needed large-scale experience.
I sold a 1PB usable-effective (after 3x dedupe/compression) space storage array last year. It was $2mil, after a 65% discount. If one time in its 5year lifecycle the "enterprisey" stuff on that array prevents about 10 seconds of downtime, it paid for itself.
But for most of us, infrequent downtime is acceptable, and most single-machine downtimes are either automatically mitigated or have very limited impact. In that scenario, getting a prosumer SSD and over-provisioning it can be a sensible choice.
So, write me some software that's going to make all that work and make the storage highly available. make sure it works for everything.
you're exposed to solutions that make that work, daily. when you swipe your credit card buying condoms - you have no idea how much happens so your transaction doesn't get lost, corrupted, or errored out.
Maybe he doesn't, maybe he does - you don't know nor do I.
I'm pretty sure this is how IBM salesmen used to respond when confronted with those newfangled Unix systems which were starting to appear here and there, nibbling first, then taking larger bytes out of their market share. Instead of the litany of diverse systems they'd have thrown LPARs, SYSPlexs and ESMs around but in the end it still came down to the same thing: this stuff is too complicated to be left to amateurs. They were right, in a way... until those amateurs grew their wisdom teeth and took a large part of their market away from them.
Yes, "enterprise" stuff is complicated - often overly so [1] - and it has its place. This does not make it the only viable solution to these problems, something will eventually come up to eat your lunch just like IBM saw its herd of dinosaurs being overtaken by those upstart critters from the undergrowth. Maybe some smart software system which "guarantees" data reliability and availability without the need for "enterprise" storage devices? It wouldn't be the first time after all.
[1] https://github.com/EnterpriseQualityCoding/FizzBuzzEnterpris...
Under sustained long duration (ie: hours to days) of continuous full load performing small block writes, even the highest rated consumer SSD drives will begin to show significantly reduced throughput in comparison to pretty much any enterprise SSD drives.
Additionally, it's totally normal to see >=5 drive writes per day and 5 year warranties on enterprise drives. Consumer drives usually are rated significantly below 1 drive write per day and rarely for more than 3 years of warranty. If you're performing a lot of writes, you're going to need to replace worn out SSDs and that's not free in a business (remote hands, downtime, etc). So buying less durable storage has costs which don't show up on the original purchase invoice but do need to be factored in within a business setting. The warranty period is the best indication of how confident the drive vendor is about drive durability.
Spending 2+X per GB means rather than 5x the same X it’s 5 times 1/2 or less total space. And that’s before you consider write amplification issues with having dramatically less total SSD space.
Enterprise SSD’s have a few benefits, but they shouldn’t be your default choice for all servers.
Looking at my list of consumer SSDs and their warranties:
Samsung EVO 850: 2y
Samsung EVO 960: 3y
Samsung EVO 860: 5y
Samsung EVO 970: 5y
Samsung PRO 970: 5y
For comparison with spinning rust HDDs:
Seagate Ironwolf: 3y
WD Red (both EFAX & EFRX): 3y
Seagate Ironwolf and WD Red are not enterprise spinny hard disk drives, instead look at Seagate X18 or WD HC560.
I have a piece of logic in my indexing code that essentially transposes an ~100 Gb multi-value dictionary on disk. If you do this the naive way with random writes, the write amplification makes it a complete non-starter. All the caching layers and buffers fill up with completely disjointed 8 byte writes and it takes ages to write.
What I've ended up doing is to in an intermediate stage write the data to be written into a series of files, containing pairs of offsets and data (up to like 100Mb each); and then going over the files one by one and essentially evaluating them as assembly instructions.
Both passes have relatively good data locality, and despite essentially writing 2.5X as much shit to disk, it takes hours rather than weeks to do this operation.
I see this sort of issue in a lot of programs. As it is a dead easy problem to make in your program. You need to write something out you just splat it out somewhere. With a dozens 1/4/8/16 byte writes instead of one big write. Basically not thinking about how that data is getting into your files. Most of the time that is just fine and not that big of a deal. But as your data set grows or you want better perf you have to worry about it. I usually use something like filemon and can see what is going on. You can see the pattern where there will be a large stack of I/O with hundreds of very small read and writes. While SSDs are an order of magnitude faster than the older drives. They still have their command structure and kernel context switching you have to deal with. You in some cases want to minimize that as it can become a large portion of writing and reading data. As with most optimizations (in this case the drive is faster) we just moved where the bottleneck is (to the kernel and bus typically).
Actually, many storage arrays use battery backed RAM, not flash as a write buffer. Flash does not have the endurance needed to serve as the write buffer for a large storage array. Some products that use battery backed RAM for this purpose will dump the contents of that RAM onto flash. I worked on a messaging appliance that used that approach, albeit with supercaps to provide the hardware time to dump 4GB of DRAM onto a compact flash card. Supercaps were easier to monitor and maintain. There are also plenty of hardware RAID cards that use batteries for their write buffer as well.
Edit: there are also persistent memories like MRAM that don't need power to retain their contents which are used in this space as well.
smaller cheaper arrays (unity, pure, isilon) and HCI (nutanix, vxrail) that don't have terabytes of RAM - they have gigabytes, and pretty much always use flash as a write cache. In fact, I cannot name one that doesn't.
No one in enterprise storage cares about the indurance of flash. All that means is that twice a year, a vendor engineer comes out to replace a few flash drives under your support contract.
And flash has been used as a write cache behind a smaller RAM cache, for two decades, by all major storage vendors.
Yes and no. Large contiguous IOs are still faster to read than a bunch of random small sectors. You generally want your frequently read files/blobs to be split into as few operations as possible.
Most controllers handle the block to page mapping problem well. They just mark the old block in the map as dead and write a new block on a new page. GC will later erase pages, hopefully prioritizing pages with the most dead blocks (subject to wear-leveling concerns).
But that same concept can be extended to partial block writes. It complicates the read process but there is no reason the controller can't coalesce multiple partial block writes into special partial update blocks and update the map accordingly. Basically write a record saying byte range (X,Y) was overwritten with just those bytes - no need to read the old block, merge the change, and write the whole block back out. Obviously you need to handle things like read chain thresholds (too many partially overlapping updates).
Let the GC handle merging the partial updates into the full block during idle periods. Again prioritize doing this for blocks on pages that are going to be erased - you had to rewrite those blocks anyway, turning what would be write amplification into a "free" write.
Like I said a lot of controller firmware doesn't bother but that doesn't mean it is impossible or unprofitable to do.
Eventually the solution really just become - get an SSD because you could throw thousand of requests and the results was fast enough.
This avoids the ugly read/modify/write, as well as much of the behind the scenes magic for minimizing write amplification and shuffling blocks around.
CPUs are similar. People started writing programs in C, so CPUs started optimizing for the outputs of C compilers. If someone were to invent CPUs and compilers today, the list of optimizations would probably be different, and the performance characteristics of real-world software would probably be different. (It's not just C; people wanted virtual address spaces, and that was slow in software, so now there is a TLB, etc.)
Actually, it goes even deeper. Half the startups I see on HN related to software operations have a quickstart like "don't worry, you don't have to rewrite your code, we'll magically do everything for you". Why not just refactor the code? It would take about an hour, and then you don't need a crazy Rube Goldberg machine to achieve your desired results. (If you want specifics, think service meshes, and how they now detect what language your app is written in so they can rewrite your HTTP handling functions to pass around a trace ID. Back in the day, you just set outgoing.Headers["x-trace-id"] to incoming.Headers["x-trace-id"]. Now we have layers and layers of OS-level machinery that still don't do as good a job as spending half a day adjusting your codebase. Billions of dollars invested in saving half a developer day! Wow!)
This is the root of attacks that rewrite the return address of a kernel call. If the return address were managed/protected outside normal data operation this would not exist.
Like, think about it, a filesystem is really a database, right? Copy-on-write even makes the "transactionality" explicit. And a lot of high-performance databases will go farther and skip the filesystem and treat the device as block storage... which it is. The filesystem is a leaky abstraction with filesystem blocks and flash pages and flash block erasure, etc.
Well, what if the SSD was just an object store? In that model you slice away a bunch of layers of abstraction and just let the SSD worry about all that. SSD says it's committed? Alright then, guess it is.
Obviously you are very much at the mercy of the SSD to implement ACID correctly however...
I don't think there's really a ton you can do at an OS level. ZFS L2ARC, system swap, etc, but a lot of the benefits would come from tuning DB configs/etc to suit the new hardware. If it's a fast SSD... ok are you the DBA?
Similarly, PDIMM is really best utilized by a ground-up rethink of applications operating in a persistent context (think of like, javacard multicore 1TB or whatever). If you treat it as slow RAM it's gonna be slow RAM. The point would be building systems and applications that exploit the "writing the memory is now writing the disk" idiom. The ideal ecosytem for PDIMM isn't Linux at all, it's JVM, or LISP/Haskell. Or android I suppose lol.
When I worked on FTLs it was always a thing executives and product managers liked to talk about but didn't seem to provide incredible results and people generally didn't like using them because it was just harder / different to manage. Has that changed?
Is this similar to what IBM had since the 60s already on their hard disks (which they call "DASD"), namely "CKD" (count-key-data) storage?
It's not really that so much as contiguous or nearby access is faster on NAND, just like it's faster on hard disks, so software optimized layout and access patterns can look similar for both.
> CPUs are similar. People started writing programs in C, so CPUs started optimizing for the outputs of C compilers.
This didn't start at a one-way street. Actually CPUs were around first, so the first C compiler optimized for the CPU it generated code for. And from that point onward, C compilers optimize for the CPU.
> If someone were to invent CPUs and compilers today, the list of optimizations would probably be different, and the performance characteristics of real-world software would probably be different. (It's not just C; people wanted virtual address spaces, and that was slow in software, so now there is a TLB, etc.)
Unlikely. At the periphery yes, but there are fundamentally difficult things for CPUs and compilers to do which shape the solutions. Before somebody says "those fundamental difficulties might be different if we invented things differently" - the story of CPUs and of compiler optimization is about inventing ways to make those fundamental difficulties less difficult.
There wasn't one day somebody invented the CPU and that was that, millions of people around the world have worked countless hours inventing and improving parts of CPUs and compilers for the past 70 years or so. There is inertia, but new ideas that are good enough can catch on even if it means entirely new models and ecosystems have to be invented. That's how we got multiprocessors, clusters, SIMD, GPUs.
I worry about the pathological cases. Imagine you have an append-only log, and you write and fsync() one byte at a time. Each time you write a byte, the entire flash block (are they still 4KB these days?) has to be erased. So you end up chewing through 4000 durability cycles on your SSD, whereas if you had waited to write an exact 4KB block, then you'd use only 1.
A hard drive is fine with this access pattern, though you'd probably be doing a lot of seeking to update fs metadata and it would be so slow that you'd come up with some other way. So maybe this pathology only exists in my mind.
NAND pages are what, something around 16-64kB these days, and block sizes are 10s of MB. In SSDs the logical size of those things may end up being larger at the FTL level if they are ganged together, but that's about your minimum.
Those are program and erase units respectively. With NAND, you do not erase a block to write. You can program pages in a block incrementally, and then you have to erase the entire block before reprogramming any.
Flash translation layers have to make this look like a disk. To do that they will do something like gather writes into a page size chunk in a small cache that is non-volatile or can persist itself on power failure. Then that chunk is programmed out to a free page. A mapping structure records the new NAND location of the logical block addresses you wrote. And a garbage collector comes along behind and compacts and frees data in block size units, erases them, and puts them on the free list.
It's a log structured filesystem with one file (the block device), if you've read any of those papers.
If you write+fsync to sector 0 of your disk 500 times, that data will get stored at 500 different places on the NAND (ignoring larger persistent caches in front of the NAND that some devices have). Selecting what pages to use and what to garbage collect etc is all part of wear leveling that is intended to prolong live of the drive. That's why endurance ratings tend to be in total writes to the drive, not writes to any particular sector.
To handwave the numbers, if you have a 100GB SSD that might be implemented with 105GB of NAND. Then if you had a program/erase endurance of 1,000 cycles you will be able to write that block 25.6 billion times. Your drive can do about 20-30 thousand QD1 writes per second, so about 12 days of that write+fsync block workload. Again assuming no front end cache on it.
You definitely can wear out NAND drives if you write a lot to them (large streaming writes would be easier to do it with), but regardless of what you do in the software layer, the drive will (should) last for its rated endurance no matter what kind of write patterns your application does. A block is a block as far as the NAND sees.
For consumer stuff you generally have to be pretty extreme or have malfunctioning software to wear them out as far as I've seen.
> A hard drive is fine with this access pattern, though you'd probably be doing a lot of seeking to update fs metadata and it would be so slow that you'd come up with some other way. So maybe this pathology only exists in my mind.
but recently there was an article on hn that talked about an arm-instruction for some js stuff.
Maybe it's not relevant to databases, but in my experience everything written with SSDs in mind has been horrifying. Compare two machines with nice NVMe drives running Windows 10/11 and Windows 7. Your user experience will be virtually identical. Swap the SSDs out for hard drives. The Windows 7 machine will get slower and have the many split-second delays we all remember from back in the day. The Windows 11 machine, however, will completely shit the bed. Massive multi-second delays crop up trying to do the most simple tasks.
So far developers, instead of taking advantage of SSDs, are just using them as a crutch.
And there's now a bunch of fs implementations. f2fs and btfs both have some support.
But there's still no products actually available to buy. We could be getting so much better at NVMe, so systematically making things better. But we've kind of been stalled out for a while, after a bunch of ceremony figuring out what we wanted to do.
SMR might have become more relevant if it had provided a useful increase in density, but that never materialized as was originally promised. Retail prices for SMR drives were never much better than CMR. The largest generally available HDDs are available in CMR flavours because that's what's needed in real world servers. There's no 10x density improvement (or even 2x). Meanwhile flash is marching relentlessly down the cost reduction path afford by Moore's law + 3D layer stacking. Really fast 1TB NVMe SSDs are under $100 now, and it doesn't look like flash cost reduction is going to slow down any time soon.
Let's go back in time almost a decade. A bunch of smart people had figured out that the cumbersome Flash Translation Layer (FTL) on SSDs made it really hard to get expectable & consistent performance. They were building possible specs to try to directly access each flash block, to be able to fill them up as they saw fit & clear them out as they saw fit. They wanted direct access, with far less work juggling complex data-mappings done on the SSD itself. Two examples of Open Channel SSD works: https://openchannelssd.readthedocs.io/en/latest/ http://lightnvm.io/
Note the huge banner on the second: "Zoned Namespace (ZNS) SSDs have replaced the work on OCSSDs, and is now a standardized interface." Everyone realized that the protocols we had for Zoned Storage were basically adequate to get us what we wanted. We'd just make a lot of 128kB or whatever sized zones on the SSD, and let people manage them themselves.
Currently this means opening blocks & then appending to them. There's been outstanding hope, in the future, that we might go further and allow more random write patterns, but still with the same contract of no over-writes (just clears).
That work seemed like it was ready to go in 2020 & 2021. SSDs were supposedly sampling/becoming available. https://blog.westerndigital.com/zns-ssd-ultrastar-dc-zn540-s... https://semiconductor.samsung.com/newsroom/news/samsung-intr... And that was reaffirmed again a year latter. https://www.tomshardware.com/news/samsung-and-western-digita...
But here we are in 2023 & there's still no Zoned Namespace SSDs (ZNS) one can purchase.
And I'm not asking to get one off a shelf but not only can I not buy one on amazon or newegg I can't even find a price when I look for a specific model!
I think it's fair to say that to a first approximation I, as a person, can't get one.
Edit: I found one site with a price on a WD drive, but they only have refurbished drives and the price is 25% off of $2200 for 1TB, so I'm going to ignore that site.
That's true of most enterprise/server components. At most you'll find a laughably inflated list price that hardly anyone actually pays.
The whole story here is that modern SSD controllers tend to embed a bunch of very fast very expensive data-processing cores to quickly do a ton of fancy mapping. We don't hear it quite as clearly these days, but for example for a while higher end Samsung SSD controllers were announced as penta-core ARM Cortex-R[1] systems, and I think those Cortex-Rs are what most folks do. (It was groundbreaking news that WD released an open source "swerv" RISC-V core which they'd experimented with using instead, https://blog.westerndigital.com/risc-v-swerv-core-open-sourc... .)
Ideally, the fantasy is, we can build much cheaper & faster controllers that do far less. Zoned Namespace drives should ideally have enterprise grade, fast, dual-port access, but in many ways, I feel like they should ideally be far cheaper than DRAM-less drives. They should leave even more up to the host. They should be incredibly dumb & simple drives. They should be lower power, by far. They should eskew DRAM. Drop all the legacy baggage & expose what you are, un-intermediated, so we can be fast & use it well. Which you, the drive, with all purpose tasks, could never do.
It's not worth it, to me, to make a drive that still keeps the ancillary old baggage of convention. If you want to make a good ZNS drive, make a good ZNS drive.
Yes, that's because we're in the drive-managed transition phase of zoned storage (not clear when -if ever- we'll leave that phase). The first zoned disks needed to present themselves as a dumb unit to the OS, because none of the operating systems had any support for it.
Now that at least some OS'es have native support for zoned storage, we might finally see host-managed zoned storage on the market, but as you say -- the technology has not delivered on the projected storage benefits, so the value-add of CMR disks is very much an open question.
And there's also the consumer confusion angle: how to market a device that can physically connect to their computer, but might logically not work at all? So I'd suspect to see the first host-managed drives in enterprise datacenters (think Amazon Glacier) -- but FAFAIK they're nowhere to be found.
Argh. SMR, that is.
For an analogy, I only buy LED lightbulbs, love that tech. When they first became widely-available, my house had a few incandescent bulbs on the porch that I probably just left on for 5 years.
The supposedly superior led bulbs in that adoption period often died surprisingly fast, compared to the claims.
To bring it back to storage. I use as much NVMe as I can, but I’m still a bit uncertain on how much a risk I’m taking with my filesystem choice and access patterns.
My answer: buy more SSDs when they go on sale! I don’t have intuition about lifetime even now.
I know, HDDs have equivalent classes of worse problems. I just learned to expect them to reliably fail, like a few days after a build ;)
$ sudo smartctl -A /dev/nvme0n1
Available Spare: 100%
Available Spare Threshold: 10%
Percentage Used: 0%
Data Units Read: 626,936 [320 GB]
Data Units Written: 22,262,139 [11.3 TB]
This is on a few months old PC which is why percentage used is still 0%. The number does slowly go up over time especially if you do heavy writing.As with LED lightbulbs, the consumer industry is focused on costcutting to deliver shitty products at low prices. Any high quality SSD is more than adequate for consumer use. Main thing to look at is how many layers the flash has. SLC is pretty much not available in consumer products. Triple layer is very adequate. All the cheap products have four layer, which you don't want unless you know you're not going to be writing often.
Firmware bugs or bad controllers are still a potential issue. I have a 1.5 year old drive which is at something like 65% percentage used, it's a big drive with a high TBW rating and only a moderate amount of data written to it, but the excessive wear is known issue with that model. It'll get replaced under warranty but that's only kicking the can down the road since the replacement will have the same issue.
The strange thing to me that gives me weird feels despite the science is that I’m never had an ssd die that I installed. And I’m no Linus of LTT here.
Plenty of apple machines and etc have needed a catastrophic replacement. That just throws off my ability to make rational decisions about it at times.
“Dude! That SSD was premium priced, what the heck? Thermal failure?”
(Seems likely, but ugh. Note: Opinion was formed a few years back, so it’s not something I’m actively diagnosing.)
So when SMR got released into the open market, it was exclusively drive managed, and mainly on the low-end, as a cost reduction strategy. With all the performance downsides you'd expect.
It is a real shame that host managed SMR isn't available outside the mega-scale corps. It would be nice to have access to an intelligent storage stack mixing flash and SMR drives without having to go work for one of the big N companies.
The IO write speed is a bit of a secondary concern. And since the data is immutable, you don't actually overwrite it all the time. And of course, SSDs have been around for quite long and spinning disk was already on the way out over a decade ago on high end servers. So, things like Kafka and event sourcing mostly became popular after that started happening; not before. This was never really about SSDs vs. spinning disks. Instead it is about the guarantees that come with having immutable data in a distributed setup. You see that in the Elasticsearch world as well (lucene uses append only data structures). Elasticsearch emerged around 2010. Almost from the beginning the consensus was that ssd was vastly preferable to spinning disks for scaling. Not because of the writes but for reads. Of course lots of people still used spinning disk around the time as well.
So, sequential reads and writes are two things. Elasticsearch writes sequentially; but reading is random access. Which is why SSDs are nice to have for Elasticsearch clusters even though it uses append only storage.
I know I went pretty much all SSD ages ago. I do run backups to a few 14TB externals but nothing outside of that is spinning rust.
The point is, DRAM-less 2TB M.2 NVMe disks are at $80 right now in select non-US markets. This is mere ~5x HDD, and slow as 4Gbps on writes at worst conditions.
Personally it really felt a bit rotten to me that Intel had created drives like the 670p that would simply stop working after a designated amount of writes. On the one hand, predictable life-span is kind of of a good thing, but sending drives to the landfill before they're actually done seems monstrous. Anyways though, yeah, DRAM-less drives are available for incredible prices.
We got a bunch of enterprise grade Samsung SSDs and there'a large difference in sustained I/O. It's not "instant" by any means, but it there are other things slowing I/O.
I am all ears... how are log writes random?
The twain are similar, but aren't exactly the same. Optimize for the cases where you are writing multiple blocks in a chunk, with contiguous free block allocation policies and write behind and whatnot, and you haven't in reality quite optimized for the case where what you're actually doing is overwriting the tail block in the file repeatedly, oftentimes changing as few as mere 10s of bytes in that block with each read-modify-write request.
(And yes, logging mechanisms often do like to ensure that logs are synchronously flushed to disc line by line, effectively and intentionally defeating write caching and write behind optimizations.)
> At base, most logfile output is a read-modify-write
It's worse, actually, as you never can tell if this would be RMW to the same block or to the other free block. In the former you waste the writes for the whole blocks. In the latter you can hit a write amplification despite never have been writing any meaningful amount of data.
Now you just need to nail the peak of $/GB, I'm looking at both 870 EVO / QVO, they have 4TB versions for $300.
For "write once" or append only data these could work well, think user registry or sequential log.
But you need to offload the active data on SLC drives or take the risk of the wearing driver firmware in the exact devices you buy/patch.
The OS needs to be on SLC anyhow.
This doesn't sound right. Does this mean that the drive decides to arbitrarily move blocks between your partition and unpartioned spaces to manage GC?
An SSD will still provide a virtual bitmap so file systems can still order it to write a bit at a given location like ye olde days, but the SSD controller ultimately decides where to actually write that bit which more than likely has no relation to the specified location.
Incidentally, yes: This also means defragmenting an SSD is not only harmful, it is outright misleading because the state of the file system is not related to the state of the bits on the drive.
You can just treat it as memory, and be able to mmap() it from userspace?
The SSD controller will handle all the bad block remapping, ECC, etc. transparently in that case.
It should be possible to make a PCIe x16 card for this, so the CPU will not be tied up as much waiting for the flash.
Now, as for a write-optimized log, it's important from a durability perspective that data is persisted to media as soon as possible after the OS orders a flush. In the storage hierarchy, even the increasingly rare battery-backed DRAM drive has its place.
If that means sequential writes are significantly faster than random writes then no amount of OS abstraction will make an application that is structured to require random writes as fast as one that uses sequential ones.
If you don't care about performance, then the abstraction layers we have let you ignore this level of detail already.
I really care more about correctness, portability, and long-term maintainability. Optimizing for the idiosyncrasies of a given OS at a given point in time go against them all. Performance can often be increased by adding hardware (or waiting for the next Moore iteration).
There are a few cases where performance is critical, in the sense that, if you don't achieve it, the system is not fit for purpose, but they are vanishingly rare.
Yeah, it's better compared to (1 MiB/s writes;8 MiB/s reads) on HDDs, but it's not that good.
Is that simply a typo or am I missing something? Shouldn't it be mostly random for long-term storage?
In the worst case though, the long-term storage can still have random writes though no (for btrees anyway)? I figure LSM tree writes can always be sequential, even during compaction. But btrees can only do solely sequential writes if the entire tree needs to be rewritten. And you don't want to always do that if only parts of the tree need to be updated?
With all the different layers in modern storage, I think you can get asymptotically close to the sequential rate once you start hitting some of these other multi-block granularities, i.e. around the block size for flash erasure, encryption, redundancy coding, etc. These don't have to be that big, i.e. several megabytes.
Of course it's workload specific but that type of a performance hit is going to effect a lot of workloads.
> Betteridge's law of headlines "Any headline that ends in a question mark can be answered by the word no."
https://en.wikipedia.org/wiki/Betteridge%27s_law_of_headline...
(Betteridge's law of headlines strikes again.)
PCIE is serial.