Samsung Announces Key-Value SSD Prototype
anandtech.com
anandtech.com
A lot more primitive than this, of course, but funny how old ideas eventually become new again.
The original indexed file format under MVS, ISAM, heavily relied on the physical hardware keys on (E)CKD disks. In the 1970s, IBM replaced ISAM with VSAM, which didn't use that physical key feature and handled all the keys in software – and it was actually faster. (I don't know if the reason for it being faster was due to some fundamental flaw with the underlying idea, or just due to limitations of the particular implementation.) (Despite being deprecated, IBM continued to support ISAM until the mid-2000s, when they finally removed it from the operating system.)
IIRC, the legacy PDS (partitioned data set) format also uses it, but not the newer PDSE 1.0 or PDSE 2.0. (PDSE = partitioned data set extended). (PDS is like an archive file format, for object code it functions similar to .a files on Unix, for source code it was used to make up for the fact that originally MVS didn't really have the concept of directories, so keeping all your program's source modules in a single dataset was more convenient.)
I think the mainframe filesystem (VTOC) also uses it a bit.
There's probably a few other random things in z/OS that still need it, but the newer stuff (VSAM, PDSE, HFS/zFS, etc) doesn't really use it. I think if IBM really wanted to, they could add support to z/OS for running on industry-standard FBA disks instead of ECKD. (They did the same thing to VSE all the way back in the 1980s.) I think the main reason they don't do it, is not technical, but commercial – mainframes needing special SANs helps keep their storage business alive, if mainframes could work on industry standard SANs, they'd face more competition in that area.
I dislike this meme - things are generational, not cyclical.
You have a comparatively weak CPU (think ARM Cortex M series) on an SSD and drive companies have been notoriously bad at firmware development.
Frankly, I don't trust Samsung to implement a file system. I prefer to not even use their enterprise flash based on past trauma, but those are usually harder fails than the subtle fuckups they can pull off in an opaque FS.
Is is a real pity that so few filesystems systems have integrated data checksumming -- I'd love to use something other than ZFS for this.
The computer also gets a mention in his 2011 book "Alan Turing and his contemporaries".
On most magnetic drives there is no inherent relationship between the sector number and its physical location on the track and the drive simply waits for the sector with correct sector number in its header to come under the head.
This is false. Sectors are approximately in block order. You can write a predictor which can predict the latency to read a block based on the last block read using that information.
You can also take the lid off a drive and watch it with a high speed camera to see where each bit of data gets stored.
A read was in 3 parts parts: first do a seek over the right cylinder, then select a head, then read a sector. That least part read everything spinning by until it saw a sector with the correct sector ID.
Sectors were not in block order because if you read sector 1 then tried to read sector 2, it would have already spun past the head, meaning sequential sector reads would require a complete revolution of the disk.
So instead, sectors were interleaved with a skip. There were 9 sectors on a track, so the sector IDs might be written as: 0 3 6 1 4 7 2 5 8. You could read 3 sectors per disk revolution instead of 1.
Having enough logic in all your storage and peripherals that the hardware itself presented a nice, directly usable API was the thing that differentiated mainframes in the era when minicomputer were cheaper and started getting faster cpus. (That and "no-one's ever gotten fired for buying IBM".)
When a fundamental technology is evolving rapidly, not much effort is put into optimization. It instead goes into keeping the momentum of evolving the fundamental technology.
When it hits a physical limitation, everyone scrambles to now optimize.
Until the next generation of fundamental tech shows up and evolves rapidly again.
Think about air travel and how every airline is now all about fuel efficiency.
A single new technology eg: a new battery chemistry that is say, 2x as energy dense as the best right, now comes along, electric air propulsion will become a major thing. And airlines will be scrambling to offer faster, quieter options that may not necessarily be the most efficient.
It's not all Electron apps out there.
Also, when I call something "accumulated cruft", it's not just vague "bloat" concerns I'm referring to; you have the problems of several layers of leaky abstractions, and those leaky abstractions may well date from an era of radically different hardware. It's true this can cause "bloat", but it causes other problems, too.
There is too much cruft in the TCP/IP + HTTP stacks, so the solution is to implement a new protocol on top of the lowest abstraction that can be practically used (IP+UDP) in user space; at least until kernels implement it.
I will delegate judging if this is a good idea to more knowledgeable parties though.
Running everything that used to be done in the network layer up to the application layer and putting everything on top of HTTPS is the opposite of what the parent is advocating for.
My software engineering instincts tell me this is all a bad idea and those layers should be kept separate as they were, if indeed they weren't already intermingled more than they should be (e.g., SSL & SNI; "why should SSL have to know anything at all about how HTTP deals with host names?" say my instincts). However, I intellectually realize that my software engineering instincts are not calibrated for the size of the HTTP world, nor the amount of resources that can be thrown at this problem. It's a great example of slicing through many layers that make sense independently, but just never quite went together properly, to make something that's going to perform better without 30 years of mismatching layers stacked on top of each other.
(I'm not saying HTTP3 is perfect, only that it is a good example of someone at least trying to do what I was talking about.)
It circumvents the established network and software stack and builds a new protocol on the lowest-level abstraction that can be used practically (IP + UDP) because everything else would require hardware and kernel support, which would take many years to manifest.
Seagate had a product called Kinetic that was an Ethernet-connected key-value spinning disk. We had a 1U box with a number of 4000GB drives in our lab for an evaluation.
It was interesting but we didn't find a use for it. Product is dead now as far as I know.
Sun Microsystems' motto was "the network is the computer" but over the last fifteen or more years there's been a steady trend in the opposite direction.
The computer is a network. I sometimes wish this would be embraced more thoroughly, not just in the way you mention, of offloading work to specialized processing units, but of building a network fabric inside the box. It's one of the things we never copied down from classical mainframe design.
Just a note, That trademarks now belongs to Cloudflare [1]
[1] https://blog.cloudflare.com/the-network-is-the-computer/
I think some network chips have associative operations like this in hardware (find route matching ...), it would be interesting to have a generally available solution.
But clearly I'm missing something because everybody loves them and uses them. Is this a technology I would only use when building websites at scale or something?
Redis is very fast, dependable and relatively easy to use. The alternative, other than a competing product, is pretty much that you have to write your own custom replacement Redis in some manner. That might be fine at a small scale for fun or in an experimental product that isn't meant to go into production; otherwise you want something well proven, that many other engineers you can hire are likely to have experience with (it's hard to over-emphasize that last point).
https://meta.stackexchange.com/questions/69164/does-stack-ex...
Key-value is great when you have a key-value pair that you'd like to cache. Specifically in web workloads, maybe you'd like to cache the API response (value) to a certain API query (key) such that its returned without processing power from your backend. Or maybe user permissions per action, where they don't really change as they're reliant on the user's role.
I'm not entirely sure what a valid use-case would be for an embedded device, but if there's any kind of SQL query you would make on an embedded device that returns a response that hardly changes, you might also spin up some kind of key-value store such that it's returned without (well, faster than) SQL query time.
It probably doesn't make sense to completely replace a SQL database if you're trying to represent relational data.
Suppose I have some function, get_members_of_group(dt, groupname). It returns a set of strings. There are arbitrarily many groups, so I can't just cache all of them in one key. Someone might make a new group at any time, and groups are never deleted or changed, they just age out.
So far, this seems great. I make get_members_of_group_from_redis() first check redis for a key group_members_{dt}_{name}. If it's there, return it. If it's not, get it, cache it, return it.
But groups can also be empty.
Redis doesn't let me store an empty set. If a group comes back empty, the fact that it's empty can't be cached because you can't tell the difference between "not cached yet" and "cached but empty".
I've googled it a bit and all of the workarounds (sentinel key/value pair, storing "empty" at that key so the SMEMBERS fails with an error, etc) are hacky and make me not want to use Redis.
Is this a me problem and everyone is OK with slamming their db with hits if a cached function returns empty, or is everyone just ignoring all non-str datatypes and storing everything as JSON? Or is it just that weird to have a function that might return an empty set that nobody else has this problem?
> or is everyone just ignoring all non-str datatypes and storing everything as JSON
I doubt it. I've been surprised by how many of the datatypes I have found uses for.
Sometimes speed is worth it.
If you care about typing and complex querying though, you might be better off using SQLite or similar.
Is this a technology I would only use when building websites at scale or something?
Basically, though I'd amend that to read "websites that might need to scale at some point."For a lot of projects, shared memory is a perfect choice. But if you might need to scale up to multiple instances at some point, shared memory obviously won't do. And since it's not like there's a huge penalty for using Redis, a lot of times it's just convenient to use it from the beginning.
There is compared to shared memory.
If you actually need to scale and don't have an in-memory cache layer over Redis you're likely going to get hurt.
Redis is useful because it can do so goddamn much. Caching, task queuing, hyperloglog, geospatial indexing. Set it up once, and it can replace a very large number of cloud services. And it has a lot of options for replication, backups, clustering, etc. It is like a Swiss army knife.
https://martin.kleppmann.com/2016/02/08/how-to-do-distribute...
Kleppman is the author of Designing Data-Intensive Applications and generally someone whose opinion I trust in such things.
I love Redis and use it extensively at $dayjob, but I stick to Consul for managing distributed locks. Consul was built from the ground up to handle such things; Redis handles it as a bit more of an afterthought / consequence of other features.
Regarding the article, I realized I used the lock for efficiency and not correctness, and it might be why I never encountered issues.
You should use Zookeeper, etcd or Consul
Personally I'd strongly suggest using a proper consensus implementation such as etcd, consul, zk.
Key-value can also be just part of a database. Like having a key-value index.
In Java it gives teeth to the "write once, read many" pattern. By avoiding side effects you can scale your team larger. But it's a constant drain on resources to police this convention. There's always someone who thinks their reason to violate the rules is valid. Pushing the data out of process increases the cost of modifying data that should be read-only, and violations advertise loudly. You can't achieve it by obfuscation, intentional or otherwise.
Practically though, at the time KV stores were coming into existence, you could buy hardware with so much memory that Java couldn't keep up. They hit a GC wall that was causing serious problems. And running multiprocess was just anathema. If you push all of the long-lived objects out to the KV store not only does memory drop like a rock but what's left over is 'young' objects, and GC often optimizes for short-lived objects. In a way it becomes a self-fulfilling prophecy. By making in-process caching expensive they made it unwelcome.
In this same timeframe, latency on network cards became lower than latency to your drive controller. Reading something out of memory on another machine in the same data center was faster than getting it off of your own disk.
Meanwhile in an environment where all tasks are pretty isolated from each other, in-process caching is ungainly. Either you suffer with a low hit ratio, you do some sort of bizarre traffic shaping, or you do caching at the ingress point so you don't make the requests at all. But that informs your engineering priorities to such an extent most people don't like. But I see it as one of those "uncomfortable" things that Continuous Delivery advises you to face head-on. Cache invalidation is hard, yes. Wear a helmet.
So the shared-nothing people also liked these tools, which means you have two groups that probably normally wouldn't talk to each other pulling for the same team.
Was this before SSDs?
I think it makes more sense to think of KV as a minimalistic distributed filesystem, rather than as a limited kind of memory or a stripped SQL database.
Here, we don't care what gets stored there, because our API server isn't really "using" the data that gets returned; we just throw it back at the client. We also don't care about complex querying. That's where KV stores shine.
Its just a different pattern for different use-cases, at least in DB land. If you need complex querying, then KV probably isn't for you. But, if you need ridiculously fast lookups and perfect horizontal scalability, then you could equally say that SQL isn't for you, and KV is.
He said (paraphrasing) "a lot of people repeat this mantra that DynamoDB (and by extension, NoSQL in general) is flexible--which couldn't be further than the truth--it's not flexible, but it's extremely efficient."
He went on to elaborate that relational DBs are not going away, and if 1) your data will never be huge and 2) you just can't predict the query patterns in the future of your app--a relational DB is still the way to go.
However, if your data is both huge and you can spend the up-front careful planning of the queries you'll need to support and can afford the risk of future migration efforts that could be needed if your query needs outgrow your model--then DynamoDB is a slam dunk.
If you fall into that category, then you'll need to spend the time learning the "advanced patterns" of DynamoDB which include key overloading, adjacency-list pattern, materialized graph pattern, and hierarchical key pattern--and then compose those patterns into a custom "schema" (in the parlance of relational DBs).
I'm building an app right now and it took me about 30 hours to collectively enumerate the 40-odd queries I'll need to support and the 10 iterations of overhauls of my design. But boy, was it worth it, because the DB is going to easily be the cheapest m'fu* component of my app. Compare that to (taking an extreme alternative) paying for something like MS SQL server which is colossal money sink. Even compared to open-source like Postgres, my setup here will be probably < 1/10th the cost and if, for whatever reason, I see a drop in traffic or I need to shut down for a period, my DB isn't metering for compute--only storage.
In a drawn-out answer to your question of "would (I) only use (this) when building websites at scale or something?" the answer is generally "yep, probably".
"We have perfect support for both kinds of type: strings and and JSON blobs."
- A cross-platform filesystem that you could read/write from Windows, macOS, Linux, iOS, Android etc. Imagine having a single disk that could boot any computer operating system without having to manage partitions and boot records!
- Significantly improved filesystem performance as it's implemented in hardware.
- Better guarantees of write flushing (as SSD can include RAM + tiny battery) that translate into higher level filesystem objects. You could say, writeFile(key, data, flush_full, completion) and receive a callback when the file is on disk. All independent of the OS or kernel version you're running on.
- Native async support is a huge win
Already the performance is looking insane. Would love to get away from the OS dictating filesystem choice and performance.
Not wishing to be a wet blanket, but do you really believe having disk manufacturers dictating filesystem choice and performance will actually be an improvement?
- Apple made APFS, which seems to show roughly the same performance as HFS+ on SSD hardware. Some benchmarks show faster, some show slower, but it's certainly no 5X improvement across the board.
- btrfs and ZFS have some amazing features, but remain niche and don't seem to yet have wide deployment. Linux distros seem locked on EXT4.
- Windows has NTFS (and ReFS for the server)
At a consumer level, I care about: a) reliability b) performance c) power. SSD manufacturers so far have delivered hugely on all three. I'd love to see how much more they could deliver by moving a lot of filesystem functionality onto the device. Sure the first few iterations might suck like early SSDs did, but in a few years?
BTRFS, NTFS, ZFS, EXT4, and APFS/HSF+ don't really matter for general users sure, but each has certain advantages and limitations that make the flexibility should you need it very useful.
I want my storage hardware to have a higher level API than "block device". I don't want to format a drive in some OS-specific format, I want a built-in format that I can access from anything. I want this format to be durable, well specced so that I can easily access historical data. I want embedded devices to support richer storage functionality than afforded by FAT32+.
Consider that before S3, we stored files on servers in bespoke ways where we had to worry about permissions and filesystem differences (path too long!) and encryption and durability and a hundred things that we now... don't. Instead we have a simple API with dozens of conforming backend providers, including the ability to do it yourself. Imagine if computer storage experienced a similar revolution?
Sure, someone can make super-limited drive which only does 3 basic ops (read, "upsert", list). Would it be useful for general purpose computing? I doubt it.
I'm not sure I'd agree with that:
- "no attribute": S3 does actually have quite sophisticated ACLs and permissions that can be set at the object level
- "no way to say "create file but do not overwrite it": there are a few ways you can do that. Versioning, bucket policies, etc.
In fact I disagree with both yourself and the GP with your assumption that S3 is simple. If you want to do anything "enterprisey" with S3 then you absolutely do need to be aware of the maximum size of allowed objects (these days it's 5TB but it wasn't so long ago when it was only around 5GB). Then there's bandwidth costs (which means you need to also take into account VPC traffic and thus S3 endpoints) and ACLs (soooo many different places you can set permissions on objects - it's easy for someone untrained to get totally lost). There's different storage options in S3 too - depending on the frequency you want to request the data.
S3 is only superficially simple - but then the same can be said for any file system. If you just want somewhere to dump files then any fs will work. But when you start needing to performance tune and do all the other considerations that the GP highlighted with local server storage, then S3 quickly becomes as complicated as every other option out there.
This is the most "apples to oranges" comparison I've ever seen!
Filesystem do something totally different than hard disks or SSDs, and you can't replace one with another. One stores data in hardware form, the others defines the format, reliability (e.g. journaling), querying, metadata, attributes, security, etc of the data.
- Currently, SSD controller firmware goes to incredible lengths to present something that is not a contiguous block device (due to the way flash works) as a contiguous block device to the OS. Likely this includes a huge amount of logic for storing and updating the mapping between logical blocks and actual hardware storage location. Since the control APIs currently are very low level, the controller basically has to buffer everything, transform that into what the flash needs to do, and then execute, not unlike how a CPU decodes instructions and converts to uops or reorders. Presumably this is quite difficult as early controllers had all sorts of problems such as performance tanking when the disk got close to full.
- The filesystem then pretends that the very non-uniform storage provided by the SSD is actually a uniform block and introduces a layer of metadata in a bespoke format to determine the actual disk locations of files, directories, and their associated metadata. Along with that, it can include lots of very nice extra features such as encryption, journaling, hashing (for integrity and de-duplication), CoW, symlinks, hardlinks, directory hardlinks, etc.
- It seems likely then that were the storage API to be raised to a higher level (files and metadata, rather than block storage), the SSD controller can implement every one of those features with superior reliability, performance, and power usage. We've already seen where software encryption went to hardware encryption -- the result is that the encryption is mostly transparent to performance. What would happen if a file's extents mapping was simply the physical location mapping the SSD controller is already doing? If if could transparently perform hashing on read/write or keep a proper journal of file operations? Handle CoW transparently?
And all of that implemented in a way that those just become standard storage features rather than requiring a certain OS to run.
Right now, if my filesystem dies, I have a ton of tools and manuals for recovery. This is the format is well known, and diagnostics tools are easily available.
If my SSD dies, it's just gone. In two cases I had, the drive just was not allowing any file reads at all -- I had zero chance to recover the data. The drive internals are totally opaque.
So no, I would prefer my SSDs to be super dumb -- I do not trust the manufacturers to get recovery tools right.
I'd love to see something like Open Channel SSD's ( http://lightnvm.io/ , https://openchannelssd.readthedocs.io/en/latest/ ) becoming mainstream.
Spec at http://lightnvm.io/docs/OCSSD-2_0-20180129.pdf
> This specification defines a class of SSDs named Open-Channel SSDs. They are different from traditional SSDs in that they expose the internal parallelism of an SSD and allow the host to manage them. Maintaining the parallelism on the host enables benefits such as I/O isolation, predictable latencies, and software-defined non-volatile memory management, that scales to hundreds or thousands of workers per terabyte.
Yeah but are you really sure that hardware encryption is safer? Are you sure that it's not secretly saving the key somewhere where it can be extracted? Are you sure that the hardware encryption works at all and that it's not just writing the data unencrypted?
It just that instead of running on main cpu, it would run on SSD controller's microchip. And will probably be binary only, that people wont be able to inspect or fix.
Samsung has good hardware, but not that great software.
I'd rather have the opposite. SSD hardware exposing their internals, to os developers.
Some of us care about not losing metadata randomly too... or suddenly having hardlinks get duplicated...
It's insane that there still isn't a standard file system. They all do 99% the same thing now. Making them incompatible is just a pissing contest, and the users lose.
That’s not even including facts of how operating systems use filesystems that aren’t technically enforced in the filesystems itself (e.g. case on Windows). There is a lot more going on here than a pissing contest at user’s expense. In fact, if e.g. Windows threw in the towel and started using UTF-8 file names with case sensitivity tomorrow, the users would be the first to yell that half their software no longer works!
I wonder if that is possible to implement right now... It’s simple to adapt something like a raspberry pi to become a usb device, pluggable into another computer. It could emulate a USB drive, but depending upon the host computer’s OS, it could present the raw drive data as a different file system, fat/ntfs/ext4/whatever, and translate the host’s reads and writes back into the appropriate reads and writes of an internal ‘universal’ filesystem.
You’d have to make compromises to cope with FS-specific features (like how to translate users/groups onto a more basic FAT representation of the storage, etc) but in principle it could work.
...except I’m not sure if a usb device can detect the OS of the system it is plugged into?
Great idea, and if done inside an FPGA or just using some microcontroller without involving a full OS into the stack, it would be absolutely perfect. But hey, isn't that essentially what they are doing but they are changing the "API".
When it comes to fingerprinting the host OS from a USB device: https://media.blackhat.com/us-13/US-13-Davis-Deriving-Intell...
You would still have to manage namespaces, so that would be partitions in disguise, and you would still have to manage some special object name for "the thing you load and run to bootstrap the operating system", so that would be boot records in disguise, and you would still have OS specific ways to store their OS specific meta data that represents OS specific semantics, so that would be more or less filesystems in disguise. The only remaining thing would be basic interoperability ... or in other words: What FAT provides now.
> - Significantly improved filesystem performance as it's implemented in hardware.
Obviously, no such thing woud be implemented in hardware, but rather just in software on a processor on the disk device. So, it's still software, just with the guarantee that you can not possibly fix it, improve it, or adapt it to your needs.
> - Better guarantees of write flushing (as SSD can include RAM + tiny battery) that translate into higher level filesystem objects. You could say, writeFile(key, data, flush_full, completion) and receive a callback when the file is on disk. All independent of the OS or kernel version you're running on.
Which has exactly nothing to do with the topic at hand. Any "normal" SSD could do all of that. The fact that storage devices more often than not don't care about correctness has nothing to do with whether that is possible, and everything with whether you can build a cheaper product that gets better benchmark results if you don't care. Having an even more complex interface, if anything, is likely to make manufacturers cheat even more on this.
> - Native async support is a huge win
Wut?
> Already the performance is looking insane. Would love to get away from the OS dictating filesystem choice and performance.
So, you would like to get away from being able to implement any filesystem you like, or choose one of a dozen or so that fits your needs best, because that free choice in your mind somehow translates to "dictating your choice", to a situation where the manufacturer of the hardware dictates the filsystem you use (that is, the filesystem that is implemented by the filesystem driver running on the processor on the disk hardware), because ... having no choice of filesystem gives you the freedom to choose, I suppose?
Furthermore, I highly doubt this would be the reality. I imagine SeagateFS, Western Digital HyperFS, paying more for a better FS.
I wouldn't be surprised if we get Filesystem DRM, maybe Filesystem subscriptions, you can only backup Seagate drives to genuine Seagate backup systems, and that costs extra.
But maybe I'm just cynical....
From a high level users perspective file systems are already key value stores so this isn't going to change anything. There may be some gains from moving to hardware, but given hardware manufacturers track record with buggy, broken, closed source software I wouldn't be too keen.
> A cross-platform filesystem that you could read/write from Windows, macOS, Linux, iOS, Android etc.
Fat32? There's plenty of newer file systems, for the most part if they aren't cross-platform by default then either the vendor has their own solution (Apple) or your system is fairly modular by design (linux).
> Imagine having a single disk that could boot any computer operating system without having to manage partitions and boot records!
The BIOS/UEFI Still needs to know what to boot, so the drive must be partitioned somehow. You'll want to store that data somewhere, and now you have a boot record.
---
I'm not these drives solve either of those issues. Now the other ones you mentioned, those I agree with.
I suppose it’s only fitting that the top-rated HN comment is someone suggesting that a k/v SSD should be used to reimplement...filesystems.
"Future possibilities" ... in 2019.
It frustrates me to no end that there are some areas where programmers are willing to go to extreme lengths to be perfectly compatible (C/C++, Unicode, IEEE FP, ethernet, TCP/IP/HTTP, HTML/CSS/JS, PNG/JPEG, ...), and in other areas it's just accepted that everybody is completely incompatible and users have to deal with the mess (newline characters, SQL dialects, filesystems, syscalls, opcodes, graphics APIs, driver models, some multimedia codecs still, ...).
I can put a "Pile of Poo" character in a text file and have it work correctly everywhere from my telephone to a server in Finland -- but I can't put that text file on a disk and expect to see the file on two PCs that happen to be running different operating systems.
I wish we were far enough along that I could complain we can't (efficiently) run one compiled program on any operating system -- there's really no technical reason we can't do that -- but we're not even close. We can't even look at a list of data files on a disk.
What is everybody working on?! We don't need any more web 2.0. We need to fix 50 years of historical accidents and incompatibilities. In 1969, Richard Hamming said "Today we stand on each other's feet", and AFAICT nothing since then has changed. Nobody cooperates with anybody else. As Alan Kay once said, "The real computer revolution hasn't happened yet". Adding more JavaScript trackers is not going to help us get there.
OK, rant over. Sorry.
there is no hardware, its just someone else's embedded firmware
I don't see how this is related to have hardware key/value. The only thing affecting filesystem support is os developers supporting filesystems.
> having a single disk that could boot any computer operating system without having to manage partitions and boot records!
Again, not related to object stores at all. You still need your UEFI to be able to find your boot images on the disk.
> receive a callback when the file is on disk. All independent of the OS or kernel version you're running on.
Hardware interupts are always going to have to be handled by the OS.
> Native async support is a huge win.
All I/O is already asynch on the hardware level. And again, I don't see the relation to the technology actually discussed in the article.
I really don't see huge wins here. This will just increase the surface area for bugs, with less ability to fix them, and less ability to recover data when things go wrong.
I definitely do not trust drive manufacturers to write high quality software. It's just not one of their core competencies.
[1]: https://www.theregister.co.uk/2015/08/17/seagate_kinect_open...
The transfer rate of an HDD is well over 1Gbps.
I’m not sure that’s accurate. There were no intermediate steps to get to 10GbE, it was a straight jump from 1GbE and it was relatively easy.
To get to 40GbE, it involved running four lanes of 10GbE. To get to 100GbE, it involved running four lanes of 25GbE. 50GbE is just two lanes of 25GbE.
We’ve gone from a single serialised stream to multiple parallel streams in order to reach next order speeds. This is magnified when you start looking at 400GbE and 800GbE services that operate on 8 lanes (QSFP-DD or the confusingly named OSFP).
It wasn't relatively easy, because you can't run 10GBASE-T over Cat5e, and because a 10GbE NIC was as expensive as the rest of the drive put together.
Today you can run 2.5GBASE-T over Cat5e, but the standard was too late, and there was very little demand between 1GbE and 10GbE, so there are virtually no controllers. What controllers you can find are usually 10GbE controllers that can also do 2.5GbE. Until 2Gbps home internet and LANs become popular, there's no reason to expect change.
>There were no intermediate steps to get to 10GbE
That's the problem. Disk transfer rates can approach 1.5Gbps, so a 1GbE is a serious bottleneck. But 10 GbE hardware was significantly more expensive. What am I going to do, build a JBOD with a dozen disks and 1x10GbE, or put a dozen disks onto the network with 10GbE interfaces?
So instead you need 2x1GbE on every disk, which complicates management and doubles cabling, switches, and cost.
>We’ve gone from a single serialised stream to multiple parallel streams in order to reach next order speeds.
Which is hilarious when you consider PATA.
https://ceph.com/geen-categorie/500-osd-ceph-cluster/
One thing to note, in that test they were using a chassis that limited the drives to two 1 Gbe connections but that's not a limitation of the drives. Now that WDLabs has fizzled out who knows what became of that platform, it was actually quite interesting but it wasn't really directly using the drive to natively support object storage, internally it was still presenting itself as a block device to the Ceph daemon running on it.
But making it 2x1GbE is a huge complicating factor. And 2.5GbE is way out of the cost curve per bps.
The interface is for random access. Quite a few database optimizations depend on sequential access, I.e. accessing records in key order following a random access. This is why B-trees are so important. Sequential access in key order does not appear to be a possibility with this technology.
I'm not sure where you are getting this idea. How keys are organized would be up to the device, and devices could support multiple schemes. Clustering keys in a sorted order and fast in-order iteration seems like an essential requirement.
Just scanning over the specification [1], I see an iterator interface for key groups (6.4) and a setting for ordering (5.4.3)
Edit: just to clarify, I agree that this would in no way obsolete databases. But software DBs could utilize key/value disks internally.
[1] https://www.snia.org/sites/default/files/technical_work/KVSA...
Slide 43 is for write performance though. RocksDB does a lot of clever things with writes, like delayed and batched flushing with a WAL for recovery. Would be interesting how that benchmark was written.
Other slides with read benchmarks show very significant performance improvements.
But comparing directly to a software DB seems inappropriate to me anyway. We can't expect (and definitely wouldn't want) a hard disk to offer something as complex as RocksDB/LLDB/etc.
I would rather imagine that the software DBs start taking advantage of key/value disks internally.
You might have many, many sequentially related records in one 4k block of an index retrieved from a SSD. Maybe 200. Then in turn you can retrieve those index blocks sequentially and get a performance improvement.
In turn, when you're doing a merge or bitmap scan over the index, this can make a really big difference.
However those who say it obsoletes anything beyond the most trivial of KV stores are way off the mark. It's a possible optimization for all sorts of database systems, certainly including RDBMS systems (many of which are layered over a KV store or sorts). There are disk systems for Oracle DBs that can run a subset of SQL right at the storage later, filtering by predicates, doing index searches, etc.
In theory, there should be no performance advantages to embedding the storage engine this way. But there is -- back to that in a moment. Not only will a properly optimized storage engine run just as fast on the CPU but there are performance advantages to doing so for applications like databases. If you are designing a state-of-the-art storage engine for complex and high-performance storage work, these devices are not for you, you can always do better with bare storage.
The key phrase is "properly optimized". The I/O scheduling and management underlying a typical popular KV storage engine is actually quite far from properly optimized for modern storage hardware, with significantly adverse consequences for performance and durability. The extent to which this is true is significantly underestimated by many developers. From the perspective of the storage manufacturers, more and more theoretical performance of their product is being wasted because the most popular open source storage engines are incapable of taking advantage of it as a matter of architecture. The kind of architectural surgery required to address this is correctly seen as something you can't upstream to the main open source code base.
The software in these devices is typically an open source storage engine where they ripped out the I/O scheduler, storage management, and whatnot, replacing it with one properly optimized to take advantage of the hardware. This could be done in software but storage companies aren't in that business. Their hope is that people will use these devices instead of LevelDB etc, with the promise of superior performance that justifies higher cost.
In practice, these devices never seem to do well in the market. People that are using the KV stores these are intended to replace are the kind of people that do not have particularly performance-sensitive applications, and therefore won't pay a premium. And it adds no value, and has some significant disadvantages, for companies with serious storage engine implementation chops or software storage-engines that are well-optimized for this kind of hardware.
tl;dr: These are like in-memory databases. A simple way to improve the performance of applications instead of investing in hardcore software design and implementation but providing no other value.
Why is this?
Or perhaps the h.264/h.265 codecs built into modern CPUs and GPUs?
And this isn't at all a new phenomenon, either. We've been using accelerators and coprocessors (as they were often originally known) for decades.
The error rate of raw flash is incredibly high -- it takes loads of error correction to make it reliable, especially with modern MLC and QLC memories -- and the structure of flash erase blocks means that it doesn't support random-access writes. A proper flash translation layer, implemented in hardware, means that software can forget about all the strange features of flash and use it as general-purpose block storage.
And SSDs already abstract that without adding a key-value store on top of it.
Managing 16MB of SLC NOR flash is trivial for a 580MHz MIPS core with 32MB of DRAM. Managing 16TB of TLC NAND is completely different. There's a reason that SSDs almost always have 1GB of DRAM for every 1TB of NAND.
It's telling that high-end embedded Linux devices, like Android phones, typically use storage devices which implement their own translation layer, like eMMC or NVMe devices, not raw flash.
You may need a basic background in solid state physics for some parts.
Disk I/O is always the bottleneck.
Source: My day job has been as a DBA for ~15 years.
In web apps
See:
-CockroachDB (a spinoff of Google's Spanner)
-MySQL's MyRocks (Facebook)
-YugaByte (Postgres compatible sharded SQL)
-Ceph (Opensource alternative to Amazon S3)
I am guessing this is looking into improving the performance even further.Your filesystem is then layering a bunch of work on top of that to map the things you care about - files - into a bunch of fragments and metadata into those 4K chunks. You would gain the ability to do things like:
1. Throw away all the spinning disk emulation code in the SSD.
2. Align the FS level primitives with the storage: if you've set your RAID chunks to 64K per array member (for example), store a 64K object in one write, not break it into 16 x 4 K blocks. If your ZFS filesystem is set to 1 MB records, write 1 MB objects to disk, not many 4 K chunks.
3. Variable sized objects mean your filesystem could simply dispatch whole files as objects: if the FS knows your photo is a 20 MB file and your source code file is 1K, it no longer has to break the photo into many blocks, or waste a whole 4K block on a 1K file, it writes a 20 MB object and a 1K object.
4. Applications could access the storage even more directly where it makes sense: Postgres, for example, stores large records via the toast mechanism, where very large column in a row will be stored as a separately to the rest of the table (so as not to blow out the table files). You could extend that special case to simply address the storage directly, and not bother with filesystem overhead at all.
With this KV drive, I could potentially 'hardware accelerate' my database backend and all of the complexity of log-structured merge trees falls away. It would just be a handful of calls into a KV SSD library that handles the magic of getting a KV pair to disk. All I would need to worry about are the higher-order primitives I can build on top of a KV store. Using hashing schemes and clever metadata structures, you can encode virtually unlimited amounts of information into a 256 bit key. One simple flat store of KV items can contain all of your indexes, objects, scripts, jobs, settings, etc. Additionally, using hashed metadata keys lends itself well to sharing the store across unlimited drives and nodes. 256 bits comprises an unimaginably large key space, and SHA512 is actually even faster to execute on most x86 64-bit hardware if you can eat the 2x overhead on keys (which can be safely assumed to be inexhaustible through the heat death of the universe at this point).
Several vendors are also supporting the Zoned Storage concept of making SSDs that have similar IO constraints to shingled magnetic recording (SMR) hard drives. Those constraints aren't a perfect match for the true characteristics of NAND flash memory, but it does handle the problem of large erase blocks.
But I'd like to compare this to hypothetical key-value software that stores its data directly on a partition (instead of in files). Isn't this essentially the same thing? The only difference that I can see is that the software would be much harder to update, and you can offload some CPU on to the processor on the drive.
Am I looking at this correctly? I don't get why you would want this to be a hardware device.
You already need to upgrade SSD firmware these days. Linux fwupmgr has support for few of them.
1: https://www.tomshardware.com/reviews/hp-ex920-ssd,5527.html
Chances are that this Samsung device isn't even using different hardware from their standard SSDs -- just different firmware.
Some of the companies making enterprise SSD controller ASICs have decided to throw in extra ARM Cortex-A53 or similar cores for the customer to put software on, giving the application access to the data without a PCIe bottleneck. Some companies are putting machine learning accelerators on the SSD. Some are adding dedicated compression or crypto engines to transform data at wire speed. Some are just putting an FPGA on the drive, or implementing the controller on an FPGA and leaving leftover LUTs for the customer's use.
Almost any idea you can come up with about how to move storage hardware past the hard drive-like block storage paradigm has been at least prototyped and demoed at Flash Memory Summit.
https://www.donhopkins.com/home/nfs3_0.pdf
Network Extensible File System Protocol Specification (2/12/90)
Comments to: sun!nfs3 nfs3@SUN.COM
Sun Microsystems, Inc. 2550 Garcia Ave. Mountain View, CA 94043
1.0 Introduction
The Network Extensible File System protocol (NeFS) provides transparent remote access to shared file systems over networks. The NeFS protocol is designed to be machine, operating system, network architecture, and transport protocol independent. This document is the draft specification for the protocol. It will remain in draft form during a period of public review. Italicized comments in the document are intended to present the rationale behind elements of the design and to raise questions where there are doubts. Comments and suggestions on this draft specification are most welcome.
[...]
Although it has features in common with NFS, NeFS is a radical departure from NFS. The NFS protocol is built according to a Remote Procedure Call model (RPC) where filesystem operations are mapped across the network as remote procedure calls. The NeFS protocol abandons this model in favor of an interpretive model in which the filesystem operations become operators in an interpreted language. Clients send their requests to the server as programs to be interpreted. Execution of the request by the server’s interpreter results in the filesystem operations being invoked and results returned to the client. Using the interpretive model, filesystem operations can be defined more simply. Clients can build arbitrarily complex requests from these simple operations.
Raw flash has very onerous constraints on how it can be written: (numbers are from recent QLC chips)
The smallest unit that can be accessed is the individual page (64kB). Those pages belong to blocks (18MB). Each page can be either dirty or clean. You can read any page, but you can only write to clean pages (turning them dirty). To turn dirty pages back to clean, you have to erase them, and you can only erase the entire block containing the page at once. Also, each block can only be erased a certain amount of times (500) until it can no longer be written to.
As you can see, the API this provides is just awful, especially considering what kind of operations (random read/write of 4kB) operating systems offer to programs. So, in order to provide an api operating systems can actually use, there is a massive garbage-collected translation layer on top of the raw flash that makes it usable.
The idea of key-value SSDs is that given that the block layer api is entirely artificial, it might not be the best choice to present to the OS.
I suppose these wouldn't be as generally useful?
The absolutely last thing that any fs designer would want is to outsource to hardware vendor implementation of anything other than a basic operation. We already saw this in hardware RAID which sucks donkey balls compared to software RAID.
This is a pure value add play in a commodity market.
I was really hoping they were going to finally be selling an SSD that was an order of magnitude less costly than spinning rust. That's about what it would take for the tradeoffs in reliability to be worth it.
Note that the reliability concerns about SSDs have seemed mostly overblown, after years of abuse and testing.
Within that context I still see spinning rust as the longer term archive / bulk storage media. The QLC storage might provide a good precaching layer in more complex systems, and might eventually reach it's seemingly intended price point (rather than being a small discount to TLC drives that are more mature and higher performing).
The 'order of magnitude' I'm hoping for might be as small as binary (half) for me to consider it... but at 10X the storage per price it's where I hope something with no moving parts and less drive housing should be.
Which also just happened in this submission.
It is incredibly tedious to see people arguing about titles on every single topic instead of actual discussion. On some topics with less discussion it is literally the entirety of the discussion.
Unless they're putting some expensive, power-hungry CAM in these devices, they're doing exactly what I described above.
The only real advantages could be variably-sized values, and offloading CPU cycles. But a map lookup isn't exactly expensive, especially for an I/O heavy system. And variably-sized blocks aren't particularly relevant, since DBs are pretty good at packing in data.
There's not much there to justify an unusual, expensive device that's far more complicated than the alternative.
Some people care about write performance, too.
You write a dirty mapping in memory. You persist the value, and you persist the mapping. Then you flip the dirty bit.
Co-processors are cheap when they're ubiquitous, like DMA. Co-processors are expensive when they're in custom hardware, like K-V stores. I'll get more performance/$ by using standard SSDs than you will by buying specialized hardware. And because the workload is IO dominated, we'll probably both get the same absolute performance from the same server (that differs only in storage devices).
You'll only recoup those CPU cycles back if you bin pack CPU heavy workloads next to IO heavy workloads, which is rarely desirable for storage services, because it adds a great deal of variance. But you just spent shit loads of money eliminating variance by going to SSD.