I/O Is Faster Than CPU – Let’s Partition Resources and Eliminate OS Abstractions [pdf]
penberg.org
penberg.org
The peripheral never has access to main memory. There is no peripheral-controlled "direct memory access" (DMA). So it's possible to give control of a channel to a userland program without a memory security risk.
Minicomputers of the 1970s had low transistor counts and slow CPUs. So peripherals were usually put directly on the memory bus, with full access to memory. I/O operations were performed by storing into memory addresses, which caused bus transactions detected by the peripheral device. There were no CPU I/O instructions.
Microprocessors copied the minicomputer model. IBM's people knew this was a bad idea, and in the IBM PS/2, they introduced the "microchannel". Peripheral vendors, facing a new architecture that required more transistors, screamed. IBM backed down and went back to bus-oriented peripherals.
That model persists today, even though the few thousand transistors required for a channel controller are nothing today. Even though most modern CPUs have I/O channel like machinery, it's exposed to the program as registers the program stores into and memory accesses by the peripheral device.
So there's no standardization on how to talk to devices at the hardware level. Some CPUs have protection systems, an "I/O MMU", and there have been various channel-like interfaces, especially from Intel, but they have never caught on.
Instead, we mostly have heavy kernel mediation between the hardware and the user program. And way too many "drivers". This has become a problem with "solid state disk", which is really a random access memory device that doesn't write very fast. Mostly, it's used to emulate rotating disks.
Samsung makes a key/value store device which uses SSD-type memory devices but manages the key/value store itself. But you need a kernel between the device and the user program. You can't just open a channel to it and let the user program access it.
So the channel controller system has DMA; just not the peripheral.
> Minicomputers of the 1970s had low transistor counts and slow CPUs. So peripherals were usually put directly on the memory bus, with full access to memory.
That's "bus mastering" DMA. There is such a thing as "third party" DMA, which uses generic DMA controller, which is programmed to push data back and forth between memory and relatively dumb peripherals.
That approximates the channel concept.
PC's have had DMA controllers since 1980 something. The IBM PC had it in the form of the Intel 8237 chip. This is documented as having four "channels", wouldn't you know it.
Once you have two, and you want portable applications, using channel I/O instructions directly inline goes out the window; you need an API.
Porting code is easier when there are portable APIs. But that doesn't diminish from the quality of the abstraction.
Applications will typically specify transfers using virtual addresses, but DMA controllers will understand physical memory (usually, but maybe there are virtually mapped ones out there). Plus there are other issues like restricted ranges: DMA that cannot acces all physical memory but just a certain range. The API will have to convert ranges of virtual addresses to one or more physical ranges, which are then queued as one or more DMA operations and somehow deal with the problem that not just any old memory allocation in the application is DMA-able.
Applications talking to channel hardware directly using dedicated CPU instructions is only possible in some vendor-locked IBM mainframe world. It's not otherwise feasible simply for business/market/economic reasons having nothing to do with technical feasibility.
An operating system API is basically a set of extensions to the instruction set available in the application's virtual machine; it's no different from some dedicated I/O instructions, just perhaps less efficient (which may or may not matter).
Yup, this is what I’ve found on many embedded SoCs. It’s unfortunate. The fact that they need physical addresses makes them not worth using in many cases (mainly when trying to move user buffers around).
How does this not describe a PCIe bus master?
Frankly I don't understand any of this. High-performance I/O works by accessing main memory directly, just like it always has. The CPU then has to wait on main memory, just like it always has. Saying that the CPU is somehow the bottleneck seems to fall into the "not even wrong" area. There is no danger of I/O bandwidth approaching L1 or even L2 speeds anytime soon.
In the real world, if you need to deal with 400 Gb/s, you aren't going to send it to a single general-purpose CPU using any type of bus or channel. The CPU won't see it until another ASIC (or FPGA, I suppose) crunches it first.
Yes, that may impose a limit on the speed of an external network that a CPU can deal with, but that's the way it goes. MCA was never going to save us at this end of the Moore's Law curve.
So you need the hardware to divvy traffic up to multiple queues -- rings, really -- and a core for each ring. The packets just show up in memory, and the cores had just better keep up. If that looks to you like the same old POSIX architecture, I don't know what to say.
It was around 2010 that the network pipes got faster than the SAN file servers, and suddenly file output capacity had to be incorporated into network flow control.
It was actually very elegant. There were 10 copies of PP state (20 in the 7600) and only 1 actual execution unit (2 in 7600). Hardware multi-threading in 1959! So there were 10 PP's executing PP overlay code (drivers) at 1/10 the instruction rate of the main CPU. Each PP had 4K of 12-bit words, which served for both PP code and I/O buffer space. The main memory was 60 bits wide (12*5) and the addresses were 18 bits, so the PP's had 18 bit accumulators for computing addresses.
Since PP's ran only trusted code, they were allowed to scribble anywhere in main memory that they wanted to. At the end of the I/O operation, the PP computed an address for the main CPU and that directly became the interrupt vector address. This meant that the CPU never had to deal with low level interrupts, only the much less frequent I/O operation completion interrupt at the end of a long operation.
(In a past life, I did system software at CDC, and CPU logic design at Sperry-Univac and Amdahl.)
Now I need to ask him about DMA architecture...
So regardless of MC's attractiveness from a systems standpoint, it was a nonstarter in an industry that was headed towards an ecology of clones and racing to the bottom in manufacturing costs. Many PC manufacturers ditched parity memory, too.
You can use VFIO, yes it needs an IOMMU but that basically what you ask for: hardware support - and yeah, I would like IOMMU to be generalized, also because it is dangerous to have some external ports without one, nowadays (even without thinking about efficient access from userspace). And because of that you can not delegate the core of the thing to "peripheral" anymore, basically you need IOMMU and once you have it, well what else do you need?
On the storage level, you can mmap on Optane on DDR4 without intermediate software copy I think. But that's merely an implementation detail anyway. When you think about the complexities of modern controllers and the diversity of options to implement them - and at the same time the most important latencies; well the only thing that matter is the end result, and obviously it can be better to not burn some CPU cycles in some cases (for some very specific workloads), but it is not trivial what to put in HW or SW to achieve that: RAM is already crazy slow, and the faster SSD is still even slower: so you will need to be smart in SW anyway to not be bitten by latencies. So in a sense; no! modern CPU are certainly not slow!
I also think you misrepresent the common modern interfaces availables for (even interoperable) SSDs; they are certainly not "emulat[ing] rotating disks".
I'd like that you give the following some thought to finish: 1/ HW backward compat is not a prime requirement anymore (because a/ smartphone dev model b/ even PC model has shifted enough so that it actually does not support complete backward compat anymore, mostly due to SoC integration). 2/ There is "sufficient" competition, where sufficient can be defined as effective to drive innovations because they can bring market advantages, and there are a non trivial number of competitors on the market (I'll mix a little but this is still relevant to what we are talking about: Intel, AMD, Arm, nVidia, Apple, Samsung, etc.) 3/ The designers working on those chips and systems are not stupid. What we have now is of course affected by technical history, but what we have now is already immensely different from what we had during the PS/2 era. Some of the historical impacts are possibly confined in less than a square millimeter on chips, and not only your PCI was not your ISA but your modern PCIe is also already quite different from your PCI... So the impact of peripheral vendors in the PS/2 era today for all of that: absolutely null, without a doubt.
Not sure where this exploit[1] stands now, but you may need a clean slate implementation that doesn't have to boot up with IOMMU disabled for legacy reasons.
[1] https://journal-bcs.springeropen.com/articles/10.1186/s13173...
From the paper:
> NVMe SSDs perform I/O faster than the OS can accept new I/O requests and notify their completion. Furthermore, the current POSIX AIO implementations in Linux are ugly and adhoc, and they have limited support from file systems. This has lead to new AIO interfaces that eliminate system calls and leverage polling. However, polling for completion is not suitable for large I/O transfers.
AWS Nitro SR-IOV I/O virtualization [1] uses Intel VT-d Posted Interrupts [2] to avoid polling, but this CPU feature is price-segmented to Xeon E5 and higher. For other CPUs, polling [3] is necessary to achieve high IOPS with NVME. Intel and AMD should consider making this feature available on all CPUs, to support NVME and NVDIMM.
[1] SR-IOV: https://www.snia.org/sites/default/files/RonEmerick_PCI_Expr... & https://www.twosixlabs.com/running-thousands-of-kvm-guests-o...
[2] IOMMU: https://www.linux-kvm.org/images/7/70/2012-forum-nakajima_ap...
[3] Polling: https://events.static.linuxfound.org/sites/events/files/slid... & https://www.snia.org/sites/default/files/SDC/2018/presentati...
The Nitro system does use SR-IOV, and it can take advantage of hardware virtualization features like VT-d posted interrupts to lower interrupt virtualization overhead. But VT-d posted interrupts isn't a factor in avoiding polling. I'd go further to say that the Two Six Labs blog post has some inaccuracies, so I wouldn't depend on the information it contains much.
Polling is often more efficient, and lower latency, than interrupt driven operation. If a CPU is mostly doing IO, polling for work can be a much better approach. Jens Axboe has been working on adding polling to the I/O stack for a while [1], and has recently added a new API for submitting I/O that is very promising [2] (25-33% performance improvements on I/O heavy workloads).
[1] https://lwn.net/Articles/705315/
[2] https://git.kernel.org/pub/scm/linux/kernel/git/torvalds/lin...
What are some examples of these processors?
I would be interested in reading more about these mainframe processors architecture. Might you or anyone else have some links?
I need these! Not for communication between the program and peripherals. I'd like to have these for communication between processes, threads, and software Actors. I'd like to be able to map such hardware channels to channels in Golang. What we have today are software channels implemented with, "heavy kernel mediation between the hardware and the user program," which means that one may be required to think a bit too much about how processes communicate with each other, and the performance implications and how these mechanisms can break down.
Closest to complex "IO Channel Program" on PCs was I2O which largely tanked despite leaving huge mark on SCSI controller design.
If you cleanroom a database kernel design based on the assumption that I/O performance is not the bottleneck, you end up with an architecture that looks very different than the classic model you learn at university. It is always a tradeoff of burning a resource you have in abundance to optimize utilization of a resource that is scarce, and older database architectures are quite wasteful of resources that have become relatively scarce on newer hardware.
https://serverfault.com/questions/190451/what-is-the-through...
https://www.anandtech.com/show/9185/intel-xeon-d-review-perf...
Now, that 45.8 GB/sec can't usually be fully utilized by the application, but it's a lot higher than the 200MB/sec of the server hard drive.
There's also the complication of RAID, max memory fetch rate of a single core, etc. And the fact that the database has to do some processing besides moving data around.
But it seems to me that server CPU bandwidth of the last generation is significantly higher than spinning disks.
Meanwhile, a PCIe 3.0 lane has a theoretical throughput limit of about 985MB/s; 48 lanes makes 47 GB/s.
Both are theoretical numbers and in practice will be lower, and you're right that it doesn't give you a lot of free CPU to do any kind of processing on the data, but it seems like the CPU can keep up (as long as Intel continues to have an anemic number of PCIe lanes and until PCIe 4 comes along).
(FWIW, on a dual-socket Skylake-SP, Anandtech only achieved 122 GB/s (of theoretical 256 GB/s DRAM maximum). But Intel claims the dual-socket can reach 199 GB/s with an AVX-512 workload.[1] Scaling both of these numbers down by a factor of two gives a very crude estimate for single-socket DRAM bandwidth on a conventional workload (61 GB/s) and AVX-512 (99 GB/s).)
Meanwhile, AMD's Zen server parts have more than double the PCIe lanes but only a few more memory channels, so I expect you're correct there — huge/fast PCIe arrays could be bottlenecked on DRAM on that platform.
[1]: https://www.anandtech.com/show/11544/intel-skylake-ep-vs-amd...
CockroachDB gives an SQL interface over the top of a distributed data store with better consistency and relations. Cassandra has no real relations over a BigTable/Column-Store solution. It really just depends.
I think with what's coming out of SSD/NVMe and even Optain DIMMS, that there will be databases directly tuned to control/set their own data storage in these environments.
I think this is already happening. I think it was Aerospike that used the FTL of of NANd/NVMe drives for a direct key value store and I think another vendor maybe Fusion had an SDK for this as well. The Optane stuff looks really interesting, are there any server vendors shipping with those?
[1] https://lenovopress.com/lp1066-intel-optane-dc-persistent-me... [2] https://www.storagereview.com/supermicro_superserver_with_in... [3] https://cloud.google.com/blog/topics/partners/available-firs... [4] https://docs.google.com/forms/d/e/1FAIpQLSeX1tN6Qt-aQUK2iVVi...
One could probably argue, quite convincingly, that the point of truly early NoSQL was that SQL hadn't been invented yet.
Moreover, ALL existing relational databases do denormalize data behind the scenes (think about caches and indexes) for the sake of IO optimization, which means that normalization does not achieve the IO optimization goal by itself (otherwise, why bother?)
Originally, schema normalization was just about logical consistency. But that was before query optimizers were invented, and long before they grew to be so very optimized for working with normalized schemata.
Nowadays, even someone who cares 100% about performance and 0% about data integrity (And is somehow still using a relational database, yeah. So what? It's MY hypothetical.) still has good reason to default to something like 3NF. Insert famous Knuth quote here.
Caches and indexes and all that stuff may technically count as denormalizations, but, insofar as the definition of 3NF doesn't mention them, that's a tangential issue.
That's just... wrong.
Best practices with database design are to start with a totally normalized schema because it reduces scope for errors/inconsistencies.
And then to intentionally denormalize where required to increase performance, in order to reduce lookups from joins or computing aggregate functions -- at the cost of having to manually ensure duplicated data remains consistent. Query optimizers can only do so much.
(If it weren't to increase performance, I can't really think of any reason why you ever would denormalize on a single database server, unless you have some extremely complicated calculations that SQL isn't capable of expressing?)
Hopefully under profiler guidance, like for any other optimization. For my part, the frequency that I find that trying a denormalization helps performance has been on a downward trend for a while now. Mostly down to just straight-up duplicating data at this point. Databases really are getting very good at what they do.
MongoDB's insight was that the overwhelming majority of data stores don't matter enough to justify worry about correctness. It doesn't go faster than other DBs if you turn on the "do it right" flags, but their customers mostly don't.
Would you be willing to wait an extra five seconds for page load to be sure the banner at the bottom of "people who looked at this also looked at these" list is fully up to date and correct? Or is filling the boxes with any old crap good enough?
This is a separate question — do you do a live query or pre-compute it? — and it’s not really significantly easier to do that with Mongo’s model.
The pitch I saw was usually based on it being easier than having to think about your data model in advance, which is relevant to your example: everyone I know who picked it did something like that, thought it was less work and that was great, and then had some problem which came down to copies of data getting out of sync and so e.g. the “also looked at” box had the wrong price or a typo which had been fixed elsewhere. Fixing that usually cancelled out what claimed performance advantages and then some.
Interestingly, the paper does say:
[eBPF] enables applications to perform up to L7 protocol (e.g., HTTP or Memcached) processing.
Furthermore, some programmable NICs are able to execute eBPF programs directly on the NIC hardware.Michael Stonebreaker, a Turing awardee, called for complete destruction of old database order in 2008: https://www.dbms2.com/2008/02/18/mike-stonebraker-calls-for-... post which, I think, they created VoltDB (H-Store) and SciDB.
Interesting to note that Andy Pavlo worked on Stonebreaker's H-Store.
> The Peloton project is dead. We have abandoned this repository and moved on to build a new DBMS. There are a several engineering techniques and designs that we learned from this first system on how to support autonomous operations that we are doing a much better job at implementing in the second system.
Here's their alleged unannounced second-take: https://github.com/cmu-db/terrier
The archetypical differentiating features, relative to classic designs, of these future-looking (for lack of a better term) kernel designs is the prodigious exploitation of user space I/O, dearth of multithreaded synchronization, and the paucity of classic tree-like data structures (not necessarily hash tables, just not trees). The industry will get here eventually, some emerging tech sectors absolutely require it for their data models.
Database rearchitecture at the single node level can help optimize this, and maybe improve the megamachine SQL node to scale up to fatter sizes, but really big data is still a network IO and reliability challenge.
And it may be true we can pump a single fat node up to really huge throughput once some db rearchitecture fully leverages SSDs and I/O, but that just exposes more data to downtime if there is a network partition.
But even if SSDs do blow HDDs out the water on sequential IO, it's on random IO that the difference is most stark, the cost difference between sequential and random is much lower on SSD than on HDDs, and both random IO throughput and random IOPS shoot through the roof relative to spinning rust.
An SSD might be an order of magnitude better than an HDD on sequential IO, but will often be 2 or 3 on random IO. That means the loss of throughput from sequential IO to random IO (while still there due to commands overhead and the like) is much, much, much smaller on an SSD than on an HDD. And the story is similar on the latency front, an HDD might have a command latency in the low tens milliseconds, an SSD in the low-mid tens microseconds.
https://www.postgresql.org/docs/11/runtime-config-query.html...
Some reads can be sequential too if you scan through the clustering keys.
From a modern CPU point of view, that's an eternity.
If the bottleneck is not storage bandwidth, and it isn't for many recent systems and modern software architectures, then there is no practical performance advantage to keeping your data purely in-memory. This has been demonstrable ever since PCIe HBAs with arrays of cheap SSDs became practical.
SAP HANA is widely used for very large data sets.
And to my point you quoted, SAP HANA is extremely expensive to operate compared to alternatives. The licensing costs alone will kill you, never mind the hardware requirements.
A single i3.16xlarge, with 15.2TB NVMe SSD goes for $4.992/hr, or $3,594.24 USD a month. 16.33% of the price.
I'm not going to argue whether an in-memory database (measured in ns) wouldn't give you performance improvements over fetching data from an SSD (measured in ms). But not everyone needs that speed or can afford it for that price.
Sources: https://aws.amazon.com/blogs/aws/now-available-amazon-ec2-hi...
Not with storage but I/O, that's what the makers of the Cell processor (PlayStation 3) had in mind originally (it was culled before the final design though): huge IO, non-stop feeding the beast. The arch-goal was to be able to link Cells together to make a "network of CPUs" (network becomes sockets link) able to parallel executions. A hard (wild?) computer science dream/problem. (the final Cell CPU has none of that iirc). I think NVLink is a decent comparison of that purpose/design.
Now running a whole infra on RAM, not just a DB, that makes more sense on paper, and I guess that's what devs do daily with e.g. test environments, like a bunch of containers over a ton of RAM. At a small enough scale, upgrading storage tier makes sense as cost becomes negligible.
For a business / at scale, short of extremely specific applications where you'd indeed have not just a DB but whatever calls it also on RAM --- and remembering that this is not economical to serve "more" users or "faster", since you'd scale horizontally for that --- the use-case or endgame of whatever this DB serves should have that speed as a hard requirement. Likely to be 'one' monster itself, like, a supercomputer? Assembling deep learning datasets from real-time feeds on-the-fly? Skynet? Big brother? :) Jokes aside, one needs a beast to feed that'll take no less to justify a 7x more expensive RAM-based anything at scale.
Normal folks, I think we'll do with caching the hell out of our data for cheap, for now at least. Until RAM becomes abundant and CPU/storage/IO extremely expensive by comparison. (was RAM ever abundant? I can't seem to remember a time when I could just buy without counting, unlike storage or FLOPS relatively to everything else).
Naive solutions tend to either summarize the data, store as logs and then run batch processes to index in some form (or leave unindexed and just brute force the computation), or limit the incoming data rate to whatever could be indexed.
These can work for some use cases, but make it very difficult to operationalize these data sources (i.e use them to make real-timeish decisions).
Even human generated data sources (fb / twitter etc.) can generate something close to that data rate.
Just compare what it takes to draw a triangle in OpenGL 1.x[1] vs Vulkan[2].
[1]: http://nehe.gamedev.net/tutorial/your_first_polygon/13002/
I could see that for some very small niches, but in general I think it would be a terrible development for the industry.
Hardware vendors don't like to share. They don't share code, they don't share common interfaces, they don't even share documentation. As it is now, these are all problems which most userland developers don't have to care about -- those problems get dealt with in the kernel, by developers who specialize in building support for uncooperative hardware.
The average application developer doesn't want to have to figure out how many queues are supported by a NIC just to open a connection on the network. Further: the average application developer isn't experienced enough to do this correctly.
Given the niche where these tradeoffs make sense, I'm not sure why the paper bothers to emphasize security at all.
Is there any benefit to having those specialized developers create frameworks or libraries in Userland which other developers can leverage? This way they remain the interface to uncooperative hardware, but the code is in Userland so the bold folks can try their own approach.
Getting a better OS is not always accessible, even with deep pockets.
The tricky part wouldn’t be sharing code between applications. We know how to do that. The hard part would be figuring out a clean way to share the hardware between all running applications, given that any app could be terminated at any time, apps might be mutually untrustworthy and apps would have to play nice to share resources. I can imagine a hybrid approach where the kernel allocates network queues to applications and suggests userland device drivers. While running, the apps would have direct access to the hardware. And when the app is terminated the kernel would reclaim the assigned hardware for reuse by other applications.
"If we wanted, we could even replicate a kernel device driver style API in userland that network device drivers program against, done as a set of userland library files loaded dynamically based on detected hardware."
This is how many microkernel based embedded RTOS's work. Common API's written for a hardware device, say a NOR flash chip, and the driver for the hardware onboard implements the interface specific to that hardware. You can then dynamically link the driver specific to your hardware. That's what I was doing as an embedded C developer some 10 years ago, and that solution had already been implemented for at least a decade (or 2) before that.
What you don't get however is sharing the disk between multiple processes.
Part of it is the economics at the fringe can pursue speed at any cost. And part is the heady appeal of doing things differently for researchers and advanced practitioners. But, in the long view, I think you are right that it is a bad idea. If you care about maintenance and sustainability, you usually find that these bypass solutions get abandoned as soon as the more conventional approaches can approximate their speed on newer commodity hardware. So there is huge churn in these specialist devices with specialist APIs and tooling.
There is a recurring theme in high performance networking where crazy things are tried and all sorts of fancy protocol offloading written, then eventually deprecated because it is seen as a support burden and a source of bugs. Because each of these specialized stacks has a smaller user base, they are have less economy of scale to invest in maintenance and stabilization.
Truly this is what DC hardware been about for 20+ years.
Look at HPC space, and than look for "uncle Joe hosting." The moment Joe starts looking for HPC and invest heavily for some super duper RDMA powered disk, a mainstream hardware maker comes and spoils everything by releasing a mass market product doing just that if not better, thus ruining the grand scheme of gaining competitive advantage through some "secret sauce hardware"
Look at video transcoding or that new "AI" neural network thing. After few years of entertaining themselves with idea of purpose made hardware, big boys simply went the route of just using a lot of mainstream hardware.
Thus more money for us :) I think this fact almost begs to be taken advantage of. See, how things work in corporate storage products, with EMC being the prime example
OS’s would simply ship with user land drivers instead of kernel space drivers
All of the complexity of 3D rendering is implemented in userspace, and that has been the case universally in production for more than 5 years now -- closer to 10, really. If you replace "all" by "almost all", you can go back much further than that, really all the way to the beginning of GPUs' existence. And yet normal application developers don't have to care, because the driver does it for them.
The point is, drivers don't have to live in kernel space, they can live in user space as well. Networking folks may start being more serious about this nowadays, but it has been the reality in GPUs for a long time.
I asked a friend who works in a quant firm and he was like yes it’s true, and it is pretty insane.
I think there’s research Microsoft and Google are doing for RDMA over 100G Ethernet for intra data center communication as well. Pretty neat.
I'm paying $60/month for 100Mb, and if the trend of the last two decades continues, we should see another 50x increase in bandwidth per dollar over the next decade.
he seems to have more confidence in his bandwidth provider than I have in mine!
The original idea was to let purpose made hardware be distributed across DC rather than every server having to carry it: video codecs on FPGAs, hardware wire speed crypto/compression, databases and k/v stores exposed over RDMA, and remote block storage on SSDs.
I was invited to the opening ceremony for the DC. When company's bosses were showed sfx infused 3D graphs allegedly representing their AI things running on it, I was unable to restrain myself from ruining the atmosphere by asking how it is running when all servers in the DC were shut down :D
Also, there may be ongoing research, but it isn't theory at all. HPCs, HFTs, and the could providers have been leveraging RDMA for a long time - e.g. Infiniband. Doing it over Ethernet (RRoCE) is relatively new, and it isn't necessarily any big leap that is happens over 100G instead of 40 or 1.
However, an interesting point as network links go to 100G+ (esp. for RDMA) is again on the storage/processing side. E.g. a wireshark capture on a 100G connection? ~12.5 GB/second, near max bandwidth for DDR3, and can fill 64GB of RAM in about 5 seconds at full fire-hose. So again the hot-potato of bottleneck will be passed, at least for maximum sustained performance situations.
Side note, AFAIK RoCE exists mostly due to non-technical arguments, particularly the inertia created by existing familiarity and deployment of Ethernet in data centers. I think Microsoft was the one flexing on a standards-body to push it through. It is somewhat of a kludge as Ethernet wasn't designed with RDMA in mind - no guaranteed predictable latency, frames can and will disappear if switch buffers overflow, etc. So IMO "research" into the topic isn't super profound - akin to studying how your sedan might be heavily modified to go off-roading almost (but not quite) as well as a pick-up truck.
Even now many that have the luxury are just going Infiniband from the get-go if RDMA/latency are the key priorities rather than tacked on later.
If memory speed is 100ns then you would notice the memory bottleneck around the time when your processor speed is 10Mhz. This point was reached in the mid 1980s with the 286 processor. Yet through the addition of cache memory this bottleneck was hidden from most software. They continued to operate in a bubble as if they were still running on the hardware of the 1980s.
It’s a bit like life itself...we land mammals carry around bags of water under our skin and our cells are still batched in fluids as if we are still living in the environment of the oceans hundreds of millions of years ago.
Many programming languages have been invented since the 90s but as far as I know none of them explicitly model memory latency and make reference to memory hierarchies. It’s as if they still need to maintain the illusion that they are running on the hardware of the past.
(Note: I once read about a language called Sequioa developed at Stanford that explicitly modelled the memory hierarchy. I don’t know what happened to it).
Sorry I'm not following the math there, whats the relation between 100ns and 10Mhz? Why is that the tipping point?
Netflix implemented similar extensions in freebsd[3]
[0] https://www.kernel.org/doc/Documentation/networking/tls.txt [1] https://lwn.net/Articles/734030/ [2] https://lwn.net/Articles/767281/ [3] https://people.freebsd.org/~rrs/asiabsd_2015_tls.pdf
Netflix's CDN operating system is based on FreeBSD, and they did add a kind of kTLS implementation, but they did not add it to FreeBSD upstream for a host of reasons.
The above-mentioned improvements aim to address the first four without completely bypassing the kernel, instead changing APIs so they step out of the way most of the time, limiting them to coordination tasks and then either offloading to the hardware or directly piping the data to/from userspace mappings without additional context switches where needed.
VMs using SR-IOV address the security aspect.
It's basically the difference between a green field design and tacking all those innovations onto the glueball that is linux. The latter may be ugly and complex, but it has the advantage of being backwards-compatible.
I'm definitely not trying to argue in favor of the paper's opinion. :-)
I just believe the author of the article would disagree that the glueball your earlier remarks describe addresses the author's concerns. They specifically say that Linux cannot be fixed to their satisfaction, or at least argue that idea.
Look for the section on page 5 headed, "Why not use kernel-bypass techniques on Linux?"
This stuff is very new of course, so we have to wait for something to actually integrate all those pieces and then for benchmarks.
Now that we have to deal with multiple vendors of inline hardware TLS offload solutions, it is more critical to get it upstream, and so it is being upstreamed as we speak. The first piece of it (fixing send tags so they are reliable and can be used for inline hw tls in addition to hardware pacing) is up for review right now: https://reviews.freebsd.org/D20117
I like how most of the statements are supported by examples, which makes it easier to understand (after some Googling ofc), especially for someone like me who is a million miles away from academia and a programmer who rarely has to think about kernel/CPU/memory intricacies, mostly due to working with higher level languages and abstractions on top of the OS itself.
My uneducated and naive thoughts on this paper: Instead of replacing the kernel with `parakernel`, is it possible to implement a POSIX compatible kernel layer over the parakernel itself ? So that drivers, linkers, and other abstractions don't have to be re-implemented again for the parakernel.
Nothing's perfect. Unix was a great design that served its purpose well for a long time, and evolved a bit along the way. Saying it was "never good" is trivializing and, in my opinion, arrogant.
"Unix went from being the worst operating system available, to being the best operating system available, without getting appreciably better."
https://news.ycombinator.com/item?id=19416485
Which isn't to say that those POV were necessarily correct, just that it isn't all hindsight.
It's not helpful to claim "it was never good design". It was a working design that drove technology to the point it is now, and in that sense it was hugely successful. What kind of perfect and pure tech do some people want, anyway? Pick anything, whatever they like -- say, Plan 9 or OS/2 or whatever -- and I can bet you in a parallel universe where that tech won, someone on para-HN will claim that it sucked and it was never a good design and if only Unix had won.
> I can bet you in a parallel universe where that tech won, someone on para-HN will claim that it sucked and it was never a good design and if only Unix had won.
I'd take the same side of that bet as you :-P
It's a good thing nobody uses it anymore.
It only took off thanks to Bell Labs being forbidden to sell it, so it got offered for a symbolic price, alongside source code to major universities, which then decided to build on top, instead of paying OS street prices.
Had UNIX been sold in the same vein as other mainframe OSes and no one would be talking about whatever quality it might have had.
You can always get mindshare being first massively underpriced thing to market.
Or in another form, C should have done the same as other systems languages, do proper bounds checking, arrays and strings without implicit decay into pointers.
Second, having a proper UI story like NeXTSTEP or NeWS, and not the X11 Frankenstein.
Also, Unix is not X11, and weren't the other ones running on top of Unices as well?
"The language runtime is the OS." - https://news.ycombinator.com/item?id=15468395
Maybe that's how Ford beat the earlier luxury cars in raw profit, but that's not how the ICE beat electric cars.
Windows is a mainstream OS that is very different from Unix in many regards (even if it still has files). And it really shows that the grass isn't greener on the other side: some parts are great, some parts are terrible, but there aren't a lot of things that offer a universally better tradeoff.
But what are you imagining as alternatives for i/o and terminals?
Terminals and process hierarchy are a complete no-go in a fresh design. Every entity in the system can be identified through unique IDs which can be handed down to other processes based on different OS policies.
I'll treat this seriously: as software development practices are just now beginning to mature, with more emphasis on test coverage being considered best practice, sure ... maybe.
But there's still a lot of software out there for which "restart the application" (or even, "restart the stinking OS") is the only practical solution when it's behaving badly.
> Every entity in the system can be identified through unique IDs which can be handed down to other processes based on different OS policies.
Okay, but how do you write a process which is capable of communicating with any other process?
Which, in principle, doesn't prevent a program from crashing or misbehaving when it encounters a corrupted file left over from its last run. In the imagined system, the application developer would be fully aware whether he is putting a data structure in the "volatile" or in the "non-volatile" memory areas, with some safety guarantees from the OS. Restarting would zero only the volatile area, enabling a clean starting state.
> Okay, but how do you write a process which is capable of communicating with any other process?
Why would you need to communicate with any process? Maybe I need to communicate with the currently running instance of "HTTP Server App" or "Database Engine App", not with an unrelated process my program knows nothing about.
Of course. But saving state to nvram also has this problem, plus the additional problem of saving broken program state. Like I say, I'm not totally pessimistic on this anymore: there are some promising trends that make me think we might get there in the not distant future. But we're not there yet.
Maybe supporting some way for a user to manually reset broken program state would be a good enough compromise -- as long as application developers didn't do something dumb, like store their license information in their program state. Autodesk immediately comes to mind there.
> Why would you need to communicate with any process?
Because some other process wants to communicate with you.
Composability and common interfaces and loose coupling is a huge advantage in Unix-like operating systems or any other software architecture that embraces those principles.
They mean that you can write software with a much longer lifespan, and usually for less effort. In your example, someone may come along wanting to write "Firewall App" long after you've abandoned your software. If your software is extensible and supports standardized IPC, "Firewall App" is possible. If it doesn't, then it gets thrown out and replaced with something else.
Today, I piped the output of some process I was running into grep. Neither of these programs knew about the existence of the other.
Later, I used an image editor to create an image, then used a web browser to upload it to a website. The web browser and image editor were unaware of each other's existence.
Only because "an unstructured stream of bytes" is the way for Unix processes to communicate. It allows for composability, but a very fragile and unsafe one, requiring programs to spit out and read free text, with all the crazy filtering and guesswork needed.
There are also other IPC mechanisms, like D-Bus.
People have been thinking about this for decades. (Probably because it's a compelling idea.) Google "orthogonal persistence."
There's room to replace traditional folder structures with paths or anything else, and most traditional file system implementations have really slow search. But I don't see a future in replacing the notion of a file, it maps too well to the intentions of the user.
You could quibble about whether this counts as "replacing the notion of a file", but I could certainly imagine it might be useful to have a system that talks directly to disk whose basic unit of data has much more useful metadata, such as a canonicalized MIME type, much more granular access controls, much more granular access and edit history, etc. Has your mother ever downloaded a file with the wrong extension and been unable to open it? I know I have.
Similarly, the abstracted "everything is a file" notion of a file without random access, which includes sockets and named pipes and stuff, is an untyped, unsized stream of octets. Message-oriented protocols like WebSockets and HTTP can and in fact are built on top of that, of course, but it could have been the reverse: instead of a stream of octets, TCP could have been a stream of arbitrarily-large but finitely sized messages. There almost certainly would have been advantages to such an approach, and applications that didn't want the message framing and just wanted a stream of octets could have easily ignored it.
So ... something Docker-esque?
I think there are a few reasons it hasn't happened:
1. Filesystems are hard. Every new filesystem architecture ends up requiring a large pool of talented developers.
2. Nobody wants to break backwards compatibility. Current filesystem design is so integral to all kinds of software that you just can't expect all software in the world to be updated just to work with a new filesystem paradigm.
3. Databases are also hard, so a database-like filesystem is doubly so.
4. Current filesystem architecture is good enough for most stuff. The pain of continuing to use it isn't as great as the pain of changing it.
None of these reasons make a database-like filesystem inherently bad. It's just not practical right now.
I think we're moving in that direction though, with things like object storage.
I forgot about BeOS! Do you remember which version of Windows was going to do it, or at least had preliminary plans for it? I distinctly remember reading about it. Was it an early version of Windows XP or what? I can't remember...
Not saying it shouldn't be done, just that it might fail.
iOS, when the iPhone first came out, turned many of the entrenched perceptions about computing devices around on their head, and people embraced it.
Even without the celebrity power of someone like Apple, an experimental system could still thrive today in the shadows with a small cult of followers nurturing and developing it, until it breaks out.
[0]https://doc.redox-os.org/book/design/urls_schemes_resources....
Once you have data and handles to access it you might start wanting some convenience features like access control, locking, namespacing, constraints, relationships.
I don't disagree that we can drop many of the current filesystem semantics but fundamentally all that really means is changing is the query language to access objects and manipulate their metadata.
I also don't disagree about process hierarchy. Being able to express relationships between processes beyond parent-child natively without farming out to an external scheduler would be awesome.
The software that marries these two things is basically a filesystem driver (in that you can implement filesystem semantics on top of it -- hell Ceph does it right now).
Nothing about a modern Linux/BSD system really stands in the way of doing this.
Something below the object level (possibly part of the object system, possibly a layer below that) needs to read from disk and bootstrap all those agents/live objects/actors into existence.
There were some experiments along these lines. Instead of files, just make everything an object and have orthogonal persistence for every object. In a way, this was what the early Smalltalk implementations were working towards.
I don't disagree that we can drop many of the current filesystem semantics but fundamentally all that really means is changing is the query language to access objects and manipulate their metadata.
One thing which Smalltalk demonstrates, is that the query language can simply be the same language used to implement the OS. Activity which looks a lot like database querying, but for objects, not database rows was just an expert level debugging trick of Smalltalk programmers.
But it seems to me that much of this can be done within the existing structure. You don't want I/O as streams of bytes? Great. Whatever new thing you think it should be, you can build that on streams of bytes. Knock yourself out. (It may not be as fast as it would be if it were directly supported by the OS, but you can prove the value of the concept by building on top of the OS.)
Same thing with the metaphor of file cabinets (I presume you mean the hierarchical file system.) Well, does your OS let you read and write raw disk sectors? No? Fine. Create one giant file that takes up the whole disk, and manage it yourself. Try out whatever different way of managing that space that floats your boat. Again, it will be slower, but again, you can experiment and prove out your concepts right now. You don't need to wait.
Use a single system image with no distinction between network, processor, or core communication unless needed.
Just saying “we need to change this” without saying what the short comings you want to address isn’t a super useful statement. Multiple platforms have attempted to have the user interface to their data be tag based, but that simply doesn’t scale for the amount of data people can manage with hierarchies.
Finally what the heck are you talking about in that last sentence: how does changing representation result in a new age of experimentation and (???) craftsmanship? Why is craftsmanship gated on the user level abstraction to bytes? Experimentation already happens today, what does this change?
While it will be better than current XFS, we've made aio improvement to the later over the years and today it's good enough for ScyllaDB.
Practically, even though Scylla has its tcp/ip stack in userspace on top of DPDK, we learned over the years that it's ok to use the less efficient kernel tcp stack. Most of the overhead and the optimizations can still happen within the DB itself as long as it controls the memory, the cache and manages the networking queues
Looking at modern app-based "file" access using Google docs and its ilk are that reimagining. The UI is a list of recent files, a small number of features, and then a search box. There's not File -> Save, nor am I forced to pick using a folder metaphor, where I want to put it.
That there's (likely) an underlying hierarchical filesystem somewhere below, in the stack seems like an implementation detail. As a programmer there's a library/middleware to be used to access resources, but once inside, object based access already exists. Looking at video game save files, that's been the case for a while, with the state of objects (in fact, visible objects that the user interacts with) being saved and restored from disk.
I agree it's not as satisfying as a total paradigm shift in computing on every single level, but the notion that file system, byte stream access is a holdover from a previous era ignores practical, user facing progress we've made since.
The author has done a thorough preliminary exploration on this matter. [1] https://www.nayuki.io/page/designing-better-file-organizatio...
Now lets add their approach: When you cause a page fault from accessing stuff not in memory you get the context switch but the actual workload could be handled by an auxiliary controller, it need not be on the CPU.
Changes: Locking parts of a file would be on a friendly basis, you would be able to get around the rules. Access to remote files with small chunks of data would still be inefficient--but the vast majority of accesses are local and remote accesses are generally documents that are read in their entirety.
> a 40 GbE NIC can receive a cache line sized packet every 5 ns, but the last level cache (LLC) access latency is already up to 15 ns, which means a single LLC access can already prevent the OS from keeping up with arriving packets
and
> NVMe SSDs perform I/O faster than the OS can accept new (asynchronous) I/O requests and notify their completion.
They also note e.g. that while nvme provides for 65k command queues OS generally have one IO queue per CPU.
I'm not saying that things are the same today - but it kind of sounds to me like they are. Back in the days, people were always claiming that we should switch to the newest and fastest I/O controller since CPUs were more general purpose and would therefore always be slower. It just didn't work out that way in practice.
Given the allocation of particular hardware devices - NIC, RAM, NVMe - to particular processors running a (static?) application process, it's not clear how the filesystem abstraction would work or whether that's simply delegated to the application. This is very definitely a server-focused system as no mention is made of GPUs or interactive devices.
We likely haven't designed OSes or CPUs to match this new reality.
There's nothing new under the sun, basically. It's an Ecclesiastes design. I haven't read through the whole article, but my guess is that the "parakernel" interface the authors are positing is going to look a lot like the IBM Channel interface.
http://bitsavers.trailing-edge.com/pdf/ibm/360/princOps/A22-...
[1] https://en.wikipedia.org/wiki/Cell_%28microprocessor%29#Syne...
You have reinvented the Mainframe Channel Processor!
Your next challenge: Try to avoid reinventing the 3745 Frontend Processor.
Because all of this had happened before and will happen again, often without learning anything about the past (example case: NoSQL)
It doesn't seem to be that the orders of magnitude are so close as to require totally rethinking mainstream kernels.
Or am I looking at this the wrong way?
No. It's true if you care about NVMe drives, or high-speed networking--which is to say, it's true if you care about a few kinds of server workloads, but it's absolutely not true for most consumer hardware.
http://sc16.supercomputing.org/sc-archive/tech_poster/poster...
It is something like separate FPU or MMU units, built for the total control of the peripherals, so that CPU had little or no work to do. Don't forget that device drivers run on CPU.
Cisco Nexus switches can run Docker workloads also.
Think of blocks of RAM with math processors. Or the same in your NVMe/NIC/etc.
Once they have increased memory system bandwidth to be able to feed the multiprocessor throughput, the rest of the architecture is designed to make the most efficient use of it.
They spawn thousands of threads and schedule them in and out really quickly... so the processor utilization is always as close to 100% as possible. When a thread is waiting for memory it is put to sleep in a few clock cycles. When it’s data is available it wakes up and does it’s business. Since the workload is split among thousands of threads therefore there is somebody ready to be scheduled. This is why GPUs only make sense to run on massive workloads.
The same trick can be used to hide SSD latency.
I do agree with you that lots of little processors is a good way forward here with a careful eye towards reducing sharing of state, but maybe its useful in this case for them to have their own instruction streams.
Eventually, you realize that you're trying to use line buffers / getc / ungetc to parse lines on that packet of data on the iio_ring to serve that cat picture for teh Internetz. :)
We need to eliminate variable length protocols to make these interfaces go away.
- Wondermark: http://wondermark.com/357/
- XKCD: https://xkcd.com/676/
I'm not a fan of Oracle, but things like that are awesome.
But you can't - except in very specialized (ie dedicated) designs
> We solicit position papers that propose new directions of systems research, advocate innovative approaches to long-standing problems, or report on deep insights gained from experience with real-world systems. We seek early-stage work, where the authors can benefit from community feedback. An ideal submission has the potential to open a line of inquiry for the community that results in multiple conference papers in related venues, rather than a single follow-on conference paper. The program committee will explicitly favor early work and papers likely to stimulate reflection and discussion over mature ideas on the verge of conference publication.
...Where?