How io_uring and eBPF Will Revolutionize Programming in Linux
scylladb.com
scylladb.com
In 1985... yes I said 1985, the Amiga did all I/O through sending and receiving messages. You queued a message to the port of the device / disk you wanted, when the I/O was complete you received a reply on your port.
The same message port system was used to receive UI messages. And filesystems, on top of drive system, were also using port/messages. So did serial devices. Everything.
Simple, asynchronous by nature.
As a matter of fact, it was even more elegant than this. Devices were just DLL with a message port.
The multitasking was co-operative, and there was no paging or memory protection. That didn't work as well (But worked surprisingly well, especially compared to Win3.1 which came 5-6 years later and needed much more memory to be usable).
I suspect if Commodore/Amiga had done a cheaper version and did not suck so badly at planning and management, we would have been much farther along on software and hardware by now. The Amiga had 4 channel 8-bit DMA stereo sound in 1985 (which with some effort could become 13-bit 2 channel DMA stereo sound), a working multitasking system, 12-bit color high resolution graphics, and more. I think the PC had these specs as "standard" only in 1993 or so, and by "standard" I mean "you could assume there was hardware to support them, but your software needed to include specific support for at least two or three different vendors, such as Creative Labs SoundBlaster and Gravis UltraSound for sound).
Even with all its management flaws, the Amiga might have survived without the mass production of PC clones.
RIP Guru meditation.
The fate of Amiga is so infuriating. It's mind-boggling to think how Microsoft was able to dominate for so long with clearly inferior technology, while vastly superior tech (NeXT, Amiga, BeOS) lost out.
There are many such unhappy stories, and I often think about the millions of hours spent on building tech that should have conquered the world, but didn't. The macOS platform is a rare incidence of something (NeXT) eventually winning out, but the Amiga was a different kind of dead end.
This is what made it unattractive to business and continues to make it unattractive to many.
The restrictiveness of Apple is likely an advantage for novice mobile users, and other vendors copied it.
A lot of people seem to have switched to Ubuntu or Arch due to Windows 10 tracking. And these are also non-technical people that have no idea what they are doing, which is kinda awesome.
I always love when using Linux gets a bit easier to use as a Desktop for the wider audience.
It merged with PayTrust, which was acquired by Metavante, who then sold the customers to Intuit. All this happened in the early 2000s. More recently it seems that Intuit sold Paytrust back to Metavante. It's still operating a service at Paytrust.com.
I first heard about this from one of the developers of the hit game SimCity, who told me that there was a critical bug in his application: it used memory right after freeing it, a major no-no that happened to work OK on DOS but would not work under Windows where memory that is freed is likely to be snatched up by another running application right away. The testers on the Windows team were going through various popular applications, testing them to make sure they worked OK, but SimCity kept crashing. They reported this to the Windows developers, who disassembled SimCity, stepped through it in a debugger, found the bug, and added special code that checked if SimCity was running, and if it did, ran the memory allocator in a special mode in which you could still use memory after freeing it.
https://www.joelonsoftware.com/2004/06/13/how-microsoft-lost...
"Interesting defense. While Chen has a good point about Microsoft taking the blame unfairly over this and perhaps similar issues its not like Microsoft is renown for their code quality. Indeed, check out an item in J.B. Surveyer's Keep an Open Eye blog from September 2004 which details how Chen's team added code to allow Windows to work around a bug in Sim City!"
https://www.networkworld.com/article/2356556/microsoft-code-...
https://web.archive.org/web/20070114082053/http://www.theope...
It actually refers back to the Joel on Software blog post in the comment above.
BeOS had no chance to be evaluated on its own merits because by that time Microsoft had already applied anticompetitive and illegal leverage on PC vendors - for which they were convicted and paid a hefty fine (which was likely a calculated and very successful investment, all things considered).
The Amiga wasn’t even expensive for what it gave: it was significantly cheaper than a PC or Mac with comparable performance, and even had decently fast PC emulation and ran Mac software faster than the Mac.
It did not have a cheap “entry level” model, though, which was one big problem. The other (not unrelated) problem was incredible incompetence among Commodore management.
You could use a composite monitor on the A500/2000. It was only in monochrome. I did that for the first couple months I had my A500.
They should have ditched m68k too. I loved it, but with 68040 and 486, the writing was on the wall for everyone to see.
By the time of Pentium, the writing was on the wall, the floors, the windows, the ceiling, the windows.
Yes, 68060 held a candle against early Pentium but it was not intended as a Personal Computer CPU, more "fast embedded".
The Amiga OS was great. No memory protection, but Win 3.1 had none, DOS had none, Win 95 had some but it somehow crashed relentlessly anyway. It took years for them to discover that it had a max uptime of 48 days because of a timer running out of bits.
AGA could and should have been incrementally upgraded with more modes, ever keeping backwards compatibility. (Like AGA did with the original chipset.) They could have sold Amigas on PCI boards, with a cheap 68000 to boot legacy Amiga OS until the transition was complete with emulation or whatever, and using the PC x86 for games code. So many possibilities, but R&D was on a shoestring budget.
The "game console like" conformity was the strength and ultimately the downfall of the platform, but not because that's bad inherently, but because the revisions stopped coming. The original PS2 was compatible with the PS1, and the original PS3 was compatible with the PS2.
The iPhone also shows the strength of vertical integration, Commodore had a great chess board opening but traded all its pieces for nothing, except in the end, pork for the CEO and board.
This very accurately describes Go
...and was a deliberate design objective made by an ex bell labs guy.
It is interesting to note that while Brian Kernighan and Ken Thompson are involved in initial Go language design, C was largely Dennis Ritchie's baby and his got a complete PhD thesis on programming language design meaning that he basically aware of the state-of-the-art of programming languages design at that time.
The main argument of several steps back is probably about the lack of the functional language aspects like closure and this feature probably at the very bottom of the programming language features list that you want to have in porting OS, given the computer systems CPU and memory limit at the time. The other is object oriented, but you can perform object oriented programming in C inside the kernel just fine but not as gung-ho as things like the multiple inheritance nonsense [1].
Jury is still out on Go. The fact that Kubernetes is very popular for the cloud now does not mean it will be as successful as 50 years of C. Someone somewhere will probably come up with better Kubernetes alternatives soon that uses different languages. To be relevant in today and in the future, Go needs to adopt generics and its designers are well aware of the deficiency of not having generics for current Go implementation.
The majority of the today's world's OS is written in C, as a testament to the success of free beer OS given alongside tapes with source code, while other mainframe platforms required a mortgage just to start.
Had Bell Labs been allowed to sell UNIX and there wouldn't exist a testament of anything.
And do the mainframe OSes were written for portability in the first place like UNIX?
https://en.wikipedia.org/wiki/BLISS
The even the popular JVM is written in a mix of Java and C++, with plans to port most of the stuff to Java, now that GraalVM has been productised, https://openjdk.java.net/projects/metropolis/
Speaking of which, there are at least two well known version of the JVM written in Java, GraalVM and JikesRVM. Better learn the Java eco-system.
UNIX was written in Assembly for the PDP-7, C only came into play when they ported it to the PDP-11 and UNIX V6 was the first release where most of the code was finally written in C.
IBM i, z/OS or Unisys ClearPath in 2020, have completly different hardware than when they appeared in 1988, 1967 and 1961 respectively, yet PL/S, PL/X and NEWP are still heavily used on them. Looks like portable code to me.
Mac OS, you know the predecessor for macOS, was written in Object Pascal, even though eventually Apple added support for C and C++, which than made C++ with PowerPlant the way to code on the Mac, not C.
BeOS, Symbian were written in C++, not C.
Outside of the kernel space, Windows and OS/2 always favoured C++ and nowadays Windows 10 is a mix of .NET, .NET Native (on UWP) and C++ (kernel supports C++ code since Windows Vista).
NeXT used Objective-C in the kernel, that's right, NeXT drivers were written in Objective-C. Only the BSD/Mach stuff used C.
macOS replaced the Objective-C driver framework with IO Kit, based on Embedded C++ again not C. Nowadays with userspace drivers the C++ framework is called DriverKit in homage to the original Objective-C NeXT framework.
Arduino and ARM mbed are written in C++, not C.
Android uses C only for the Linux kernel, everything else is a mix of Java and C++, and since Project Treble you can even write drivers in Java, just like in Android Things allowed to since version 1.0.
Safe Ada is used alongside C++ on the GenodeOS research project.
Inferno, the last iteration of the hacker beloved Plan 9, uses C on the kernel and the complete userspace makes use of Limbo.
F-Secure, you might have heard of them, has their own bare metal Go implementation for writing firmware and is used in production via the Armory products.
IBM used PL.8 to write a LLVM like compiler toolchain and OS during their RISC research and only pivoted to Aix, because that was what the market wanted RISC for.
Contrary to the cargo cult that Multics was a failure, the OS continued without wihtout Bell Labs and was even assessed to be more secure by DOJ thanks to its use of PL/I instead of C.
There is so much to the world of operating systems than the tunnel vision of UNIX and C.
The original JVM written by Sun was in C not C++ or Java.
Windows NT the kernel part is mainly written in C. The chief developer of Windows NT Dave Cutler is probably the most anti UNIX person in the world, but the fact that he has chosen C to write Windows NT kernel in C is probably the biggest testament you can get. Dave Cutler is also part of the original developers of VMS, if BLISS with its typeless nature is better for developing OS than C, he'd probably has chosen it.
For whatever reasons Multics had failed to capture wide spread adoption compared to UNIX and the fact that its name existed mainly in most of operating System books as pre-cursor OS to UNIX. For most people Multics is like B language that is just a pre-cursor to C language. I know it a shame that Multics had become a mere footnotes inside OS textbooks despite its superior design compared to UNIX.
PL/I language is interesting by the fact that it is quite advanced at the time but as I mentioned in my original comments, Dennis Ritchie had to accommodate the fact that some of languages features are over engineered based on the hardware of the day and had to compromise accordingly. Go designers, however, have chosen to compromise not based on the hardware state-of-the-art but what the language designers think are good for Google developers at the time of the original language design proposal.
Sometimes it feels like all the hate on HN toward Go is ignorance that there is a whole domain of software outside of scripting and low level systems programming, and how some enterprises value 20 year maintenance over the constant churn of change eg Rust and JavaScript. And yes, I often hear people saying “you can still do that in x or y” but the point is that Go does it better than most languages because it was purposely designed with those goals in mind - hence exactly why it suffers from expressiveness, state of the art features et al. And I say this from 30 years of experience writing and managing enterprise software development projects across more than a dozen different languages.
Go might not be cool nor pretty, but it’s extremely effective at accomplishing its goal.
Go reduces complexity in order to make it easier to build resilient systems.
A language like Perl has bucketloads more features, and more expressive syntax, but I’d still say Go is many steps ahead of Perl.
On another note, I’d actually argue that some of Go’s features, such as “dynamically typed” interfaces and first-class concurrency support are streets ahead of most other languages. Not to mention its tooling, which is better than any language I’ve used, full stop (a language is so much more than simply its syntax).
I believe that functional languages, with proper, fully-fledged type systems, are the best way to model computation. But if I had to write a resilient production system, I’m choosing Go any day.
As someone who loves software, there was a very clear feeling at that time that Microsoft was putting a huge chilling effect on the whole industry, and that the entire industry was stagnating under their control.
Thank god for Netscape, Google, Apple, Facebook, and Amazon (in that order) who were able to wrest that control from them. Now at least there are multiple software ecosystems to move between. When one of these massive companies poisons the water around them, there are other ecosystems doing interesting things.
Which other OSes from the era are you referring to here?
This made upgrading chips nigh impossible without full software rewrites, which ultimately caused stagnation.
Indeed, as an A500 kid I used to laugh and was horrified by my first PC...
I think that epoll timeout granularity is still in milliseconds, so if you want to build high res timers on top of it for your event loop you have to either use zero timeout polling or use an explicit timerfd which adds overhead. I guess you can use plain ppoll (which has ns resolution timeouts) on the epoll fd.
All the DOS games at the time used IPX for network play for a reason. TCP was too "big" to fit in memory.
All I had was the 512kB expander, he had a 386 with 387 and could only run a single tasking OS
To say a 386 is limited to single tasking is wrong though. That was my main point.
- in the late 80s, Commodore ports AmigaOS to 386
- re-engineers Original Chipset as an ISA card
- OCS combines VGA output and multimedia (no SoundBlaster needed)
- offers AmigaOS to everyone, but it requires their ISA card to run
- runs DOS apps in Virtual 8086 mode, in desktop windows or full-screen
[1] for example: https://en.wikipedia.org/wiki/Amiga_Sidecar although "card" is stretching it :)
https://bigbookofamigahardware.com/bboah/product.aspx?id=329 https://bigbookofamigahardware.com/bboah/product.aspx?id=330
Reminds me of: https://en.wikipedia.org/wiki/Unikernel
I do remember that, and it was cool. But, lightweight efficient message passing is pretty easy when all processes share the same unprotected memory space :)
Memory access is enforced, although not technically via the kernel. Rather at boot time the kernel owns all memory, then during init it slices off all the memory it doesn't need for itself and passes it to a user space memory service, and thereafter all memory requests get routed through that process. L4 uses a security model where permissions (including resource access) and their derivatives can be passed from one process to another. Using that system the memory manager process can slice off chunks of its memory and delegate access to those chunks to other processes.
All problems in computer science can be solved by another level of indirection.
The rejoinder, and I don't know who gets credit for it, is: All performance problems can be solved by removing a layer of indirection.Sure, I still write asynchronous code. Mostly to find out if I can. My experience has been that async code is hard to write, is larger, hard to read, hard to verify as correct and may not even be faster for many common use cases.
I also wrote some kernel code, for the same reason. To find out if I could. Most programmers have this drive, I think. They want to push themselves.
And sure, go for it! Just realize that you are experimenting, and you are probably in over your head.
Most of us are most of the time.
Someone will have to be able to fix bugs in your code when you are unavailable. Consider how hard it is to maintain other people's code even if it is just a well-formed, synchronous series of statements. Then consider how much worse it is if that code is asynchronous and maybe has subtle timing bugs, side channels and race conditions.
If I haven't convinced you yet, let me try one last argument.
I invite you to profile how much actual time you spend doing syscalls. Syscalls are amazingly well optimized on Linux. The overhead is practically negligible. You can do hundreds of thousands of syscalls per second, even on old hardware. You can also easily open thousands of threads. Those also scale really well on Linux.
But on the other hand, if your server is behind a buffering proxy so it's not streaming directly over the Internet, it might not be a problem.
This is one instance of a larger pattern I've been noticing. When using some languages (like Python and Ruby) in the natural, blocking way, a back-end web application typically needs multiple processes per machine, because it doesn't handle many concurrent requests per process. Combine this with the fact that each thread has to block while waiting on the client, and you have to add more complexity around the application server processes to regain efficiency. The proxy in front of those servers is one example. Another is an external database connection pool like PgBouncer. Speaking of the database, to avoid wasting memory while waiting on it, you may end up introducing caching sooner than you otherwise would. And when you do, the cache will be an external component like Redis, so all of your many processes can use it. Or you might use a background job queue just to avoid tying up one of your precious blocking threads, even for something that has to happen right away (e.g. sending email). And so on.
Contrast that with something like Go or Erlang (and by extension Elixir), where the runtime offers cheap concurrency that can fully use all of your cores in a single process, built on lightweight userland threads and asynchronous I/O, while the language lets you write straightforward, sequential code. In such an environment, a lot of the operational complexity that I described above can just go away. Simple code and simple ops -- seems like a winning combination to me.
What are some of those use cases where userland threads are no longer good enough? In what areas do they fall short?
At my first job we had a prototype that performed 2x faster (on average) by using Go-style async, but we couldn't trust our libraries enough to eliminate bugs from blocking dispatcher threads. So we stuck with traditional multithreading.
Yes, believe it or not, you can achieve correctness and speed, together, without compromise.
But I think what many people get wrong (not the person I'm replying to) is that how you write code and how you execute code does not have to be the same.
This is essentially why google made their N:M threading patches: https://lore.kernel.org/lkml/20200722234538.166697-1-posk@po...
This is why Golang uses goroutines. This is why Javascript made async/await. This is why project loom exists. This is why erlang uses erlang processes.
All of these initiatives make it possible to write synchronous code and execute it as if it was written asynchronously.
And I think all of this also makes it clear that how you write code and how code is executed is not the same, so yes, I'm in agreement with the person I'm replying to, I don't think this will change how code is written that much, because this can't make writing code asynchronously any less of a bad idea than it is now.
JavaScript async/await is different from the others. It requires two colors of functions [1], and it conflates how the code is written with how it's executed, so it has the same problem you were talking about at the start of your comment.
Also, JavaScript async/await is suboptimal in that it's ultimately built on top of unstructured callbacks. Or, as Nathaniel J. Smith put it in a post about Python's asyncio module, which has the same problem, "Your async/await functions are dumplings of local structure floating on top of callback soup, and this has far-reaching implications for the simplicity and correctness of your code." [2] That whole post is well worth a read IMO.
[1]: https://journal.stuffwithstuff.com/2015/02/01/what-color-is-...
[2]: https://vorpus.org/blog/some-thoughts-on-asynchronous-api-de...
It allows me to write synchronous code and execute it asynchronously. The mechanism is different - but the purpose is the same. I'm not endorsing the implementation. But I do use it, because it is way better than writing asynchronous code.
No ebpf component at this time, but I do wonder if ebpf could perform journal searches in the kernel side and only send the matches back to userspace.
Another thing this little project brought to my attention is the need for a compatibility layer on pre-io_uring kernels. I asked on io_uring@vger [1] last night, but nobody's responded yet, does anyone here know if there's already such a thing in existence?
[0] https://lists.freedesktop.org/archives/systemd-devel/2020-No...
[1] https://lore.kernel.org/io-uring/20201126043016.3yb5ggpkgvuz...
For fd-based monitoring of the CQE, wouldn't a simple pipe or eventfd suffice? When CQEs get added, write to the fd, it just happens to all be in-process.
I must admit I haven't gone deep into the liburing internals or the low-level io_uring API, but conceptually speaking there doesn't seem to be anything happening that can't be done in-process in userspace atop pthreads for the blocking syscalls. It just won't be fast.
Am I missing some critical show-stopping detail?
I'm curious to see how this might work its way into libuv and c++ ASIO libraries, too.
Do windows completion ports also work that way or do they involve a system call to be performed in order to consume completion events?
Edit: actually no it's not required in all cases. Thanks for the correction.
I read that io_uring has two modes, one where you signal via a system call and another that uses memory mapped polling.
https://unixism.net/loti/tutorial/sq_poll.html states:
> Reducing the number of system calls is a major aim for io_uring. To this end, io_uring lets you submit I/O requests without you having to make a single system call. This is done via a special submission queue polling feature that io_uring supports.
It's all on github, with accompanying CppCon talk. Asio, by the way, will be C++23's network layer.
I'm however wondering what the actual quality level is, whether people used it successfully in production and whether there is an overview with which kernel level which feature works without any [known] bugs.
When looking at the mailing list at https://lore.kernel.org/io-uring/ it seems like it is still a very fast moving project, with a fair amount bugfixes. Given that, is it realistic to think about using any kernel in with a kernel version between 5.5 and 5.7 in production where any bug would incur an availability impact, or should this still rather be a considered an ongoing implementation effort and revisited at some 5.xy version?
An extensive set of unit-tests would make it a bit easier to gain trust into that everything works reliably and stays working, but unfortunately those are still not a thing in most low-level projects.
io_uring has many tests in the companion user space library liburing, maintained by the same person that made the kernel patches (Jens Axboe). They test both the library as well as expected functionality in the kernel.
io_uring is not going to give you speed ups if you use it in the same way as you would epoll or kqueue. Thus, simply sticking it into e.g. libuv without changing how the applications are built probably won't give you a lot of benefit (speculating).
It comes down to how you work with the ring buffers and how much you take advantage of the highly out-of-order, memory-barrier-based shared memory approach as opposed to more "discrete" (maybe not the right word) syscalls.
As of yet, I haven't personally come across a published example of a production framework that utilizes these features adequately. We have some internal IP that does, but probably won't be open sourced.
io_uring allows for better utilization of fast storage.
One has to be in quite a techie bubble to equate Linux kernel features with actual world-changing events, as the author goes on to do.
More on-topic though, having read the rest of the article, my guess is that while these features will let companies squeeze some more efficiency out of high-end servers, they won't change how most of us develop applications.
You'll still need a few worker threads for blocking syscalls that haven't been ported to io_uring yet but that need is greatly reduced compared to the previous state of things.
So even if you're not using io_uring yourself the language standard libraries or server frameworks will.
There are WIPs for netty, libuv, nginx. Other projects are exploring it or have announced intent to use it.
I'm also really excited by how you can use io_uring to power everything (fs, networking etc.) with one easy api and a single-threaded event loop: https://github.com/coilhq/tigerbeetle/tree/master/demos/io_u...
io_uring makes thread-per-core designs so much easier.
It's not a tech bubble as much as it's a journo bubble. People are reading before they're writing, so he's seeing trendy topics like 2020 and the virus. He feels he needs a hook to get his readers engaged, so he's reaching for things readers can related to.
It's a bad hook. I think an editor would have cut that whole intro.
Most applications will not change the way they work with the kernel, because they don't work with it, they hide it as well as possible under libraries and frameworks. Even so, most applications need neither io_uring, nor eBPF. Hardly a revolution.
The intro is both highly visible, and, because of how writing and thinking work, it's also the spot where you're either collecting your thoughts or trying to hook the reader.
If you don't have an editor and you're done with your first draft, try deleting your first few paragraphs. It's often a simple way to vastly improve a piece.
Knowing that with eBPF I simply cannot crash the laptop I'm working on is a huge deal, and reduces the great psychological hurdle that kernel development always had (for me, at least).
Honestly, by hacking it.
There's a famous book about Linux internals that I don't remember the name (but has "Linux" and "internals" on it). But I have never seen anybody doing it by reading a book (despite how excellent it can be). You just go change what you want or read the submodule you are interested in understanding, and use the book, site or whatever when you have a problem.
It's not an easy thing, by any means. Just locating where you have to touch on the source tree is a problem that will lead you to plenty of books or sites. But don't try reading those before you have a problem to solve, you will lose time and drown in information.
(By the way, I am assuming you know how syscalls work. If you don't, go study that before you start anything.)
I was attempting to hack away a simple example but I found the USB-Serial (more) generic driver intended for "test only" and... it just worked.
Another reason for reading about IO calls, schedulers, etc? That "I'm still writing data into your USB flash drive even when the GUI says it finished 5 minutes ago" that I hate so much.
https://github.com/sebcat/yans/blob/master/drivers/freebsd/t... https://github.com/sebcat/yans/blob/master/drivers/freebsd/t...
this one? https://0xax.gitbooks.io/linux-insides/content/index.html
though it says insides instead of internals
Container problems? Namespaces, Cgroups, ...
Network problems? Netfilter, tc, lots of sysctl knobs, tcp algorithms (cue 1287947th thread on nagle/delayed acks/cork)
Slow disk IO? Now you need to read up on syscalls and maybe find more efficient uses. Copy_file_range doesn't work as expected? Suddenly you're reading kernel release notes or source code.
I don't mean to downplay the investment, but if you're already an experienced software engineer you can get into it if it interests you. There is a different mindset among systems software programmers though. Reliability comes first, performance and functionality come second. It's a world away from hacking python scripts that only need to run once to perform their function.
[1]: https://github.com/DataDog/glommio [2]: https://news.ycombinator.com/item?id=24976533
My work is "programming in Linux", but it's not impacted by any of this since I'm working in a different area.
I'm sure this is important work, but maybe tone down such claims a bit.
io_uring seems to be relevant mostly to people using or wanting to use AIO. Outside of a "few niche applications" this is unimportant for the majority of Linux developers. Libraries like ASIO would likely wrap it anyway since these low-level APIs are not pleasant to use.
It would be nice to see it at a high-level at the syscall interface i.e. currently if I want to attach a probe I have to find the function myself or use a library but it would he nice to have it understand elf files.
Only with in kernel polling mode is it close to removed. But kernel polling mode has it's own cost. If the system call overhead is no where close to being a bottle neck, i.e. you don't do system calls "that" much, e.g. because your endpoints take longer to complete then using kernel polling mode can degrade the overall system performance. And potential increase power consumption and as such heat generation.
Besides that user mode tcp stacks can be more tailored for your use case which can increase performance.
So all in all I would say that it depends on your use case. For some it will make user mode tcp useless or at least not worth it but for others it doesn't.
Additionally, compared to using epoll/select/.. for network IO, one can just submit a send/recv, instead of patterns like recv -> EAGAIN, epoll, recv
Or you can run the kernel IO thread on another CPU, but that itself has overhead compared to performing IO and handling the data all in the same thread.
If the queue completely empties, then a normal application will use a system call to go to sleep.
But as long as it's not empty, the application can keep receiving events with neither system calls nor active polling.
You'd only actively poll in very specialized/niche cases.
eBPF is code, and follows similar rules to kernel modules. That is, non-GPL-compatible eBPF code is allowed, but a subset of APIs (helpers, like module symbols) are only available to GPL-compatible eBPF programs.
Also, you're still pulling in parts of Linux into your code, so GPLv2 still applies.
That said, libbpf is LGPLv2+, even though a lot of the stuff you'd pull in to use eBPF at the kernel level forces it to be GPLv2.
Glibc being the entry point for the syscall, and glibc being LGPL is specifically why it's "okay". If you were to directly link an application to the kernel code, it would be viral.
That doesn't sound right, how is a io_submit syscall different from a read syscall? Obviously if you write a kernel module it links with the kernel, but just issuing a syscall shouldn't be considered linking, otherwise every single proprietary software that issues a raw syscall would be GPL-infringing.
While the exact nature of when a software project is a "derivative work" of the libraries it depends on is still somewhat of an open legal question, I would be very surprised if anyone were to find that a computer application were a derivative of the OS it runs on. The typical understanding of the industry is essentially a process boundary, and the boundary a system call represents is closer to a process boundary than it is to a library call.
I agree that this is the typical thinking but I've always found it a little silly and arbitrary. It implies that if I write a GPL-licensed library and release it along with a thin wrapper program that gives it a command-line interface, say it does something like a complicated calculation which reads some data and outputs a single number; then someone could come along and write a program that would not work without it, say something that transforms another input format and then passes it to my calculation. As long as that program calls my "library" as a "program" (using "system()" for example) then they are not bound by the GPL, but if they link to my library and call the calculation directly, then all of a sudden they are?
This linking vs. process boundary thing always seemed like the wrong way to determine if a program is a derivative work of another. If someone writes a program that does not work without the GPL code, they should be bound by the GPL, regardless of whether it's linked, loaded into the same process, called through the command line, or over the wire.
This last one would obviously be controversial, but frankly a lot of companies do hide their use of open source code behind a REST API, and avoid adhering to any particular licenses that way, since they are not "distributing" the software.
I suspect trying to make the case that GPL's viral copyleft isn't limited to strictly linking but potentially any interaction with it would probably have a chilling effect on the use of GPL code, and this reinterpretation would only reinforce some people's prejudice against the GPL, a la Ballmer's "Linux is cancer" line.
Maybe it's the pragmatism in me, but I think it would have a net negative effect long term, unless it managed to flip all of the tables and convince everyone to use all GPL code, instead of making people reject copyleft wholesale.
But that's not what I said. I said programs that do not work without some other program, is, in my opinion, a derivative work. I just don't see how the calling mechanism even plays into that judgement.
I do agree that there are other licenses such as the AGPL that try to cover these cases.
And arguably the online thing is a whole different ball of wax, because you can talk about software using a service, etc. It really is tricky in that case.
But I don't see the reason to distinguish between calling a function via the C stdcall mechanism, vs. "popen" and capturing stdout. It's exactly the same, logically, the only difference are details that imho should not matter for the legal case.
Right now, if I release a GPL library, what stops someone from coming along and writing a CLI program that just wraps every function with some textual interface, and including that with their closed-source program? The GPL becomes pretty toothless if it's bypassed so easily.
Let's say I'm writing a refinery simulator to sell to people, and I use a GPL command line utility to do some particular calculation about flow rates.
Now I'm GPL just for outsourcing a single equation. But only because that's the only program around for doing that calculation. As soon as someone else reads a paper on the subject and makes an alternate program for that math, my program is no longer GPL?
Those consequences sound like a mess I don't want to deal with.
I don't really see the problem. You are saying that if you change your dependency to a non-GPL program, then you are no longer GPL. The answer to your question is simply "yes".
We are not talking about patents here, but copyright. If someone comes up with an alternative implementation with a different license, you are perfectly free to start using it instead, what's the issue?
Depends on what you mean by changing the dependency. Let me lay out the scenario in more detail.
The program is still exactly the same. It asks to be pointed at a fluid sim program, and then uses that for some of the math it needs.
When it was coded, the only dependency it could use was GPL.
Now there's a new non-GPL dependency it could be pointed at, with the same API.
Now it's possible to run the program without using any GPL code. Does that make the program no longer GPL, even though it didn't change?
I'll answer with my own hypothetical. If I write a program that dynamically links a library performing the same GPL'd fluid sim calculations, it is presumably forced to be GPL, because it links to it. What if someone comes along and runs the program but at runtime uses LD_PRELOAD to override the dynamic linker, linking it to an alternative library that presents the same interface. Is the program still required to be GPL?
I don't really have an answer to your specific proposed loophole, it's pretty clever and is a very good question; but I don't think the calling mechanism is part of the issue. You could make the same argument whether you are talking about a "program" or a library. The calling convention is a meaningless detail imho.
I think you are specifically responding to my "does not work without" interpretation overly literally. Clearly if the program is written for and tested against a specific interface of a GPL'd program, it is intended to work with that program.
On the other hand if it's written to call into some kind of standard interface, it no longer requires that GPL program specifically, but could work with any program implementing that interface. And I will admit that whether a program is written only to work with a GPL program/library/whatever, or is more general, may be up to interpretation, what is considered "standard", etc., but that is exactly my point -- law is nuanced. If it were possible to codify laws perfectly with overly simple rules like "the copyright applies because it's a DLL and not a program", then we wouldn't need lawyers.
In law, intent is important. If I write a non-GPL program that depends on the functionality of a GPL library, I can go find all sorts of ways to not "link" to it but still use it, e.g., as a program, a service, etc. -- and it happens -- but the intent, which was to find a way to use GPL software without adhering to its license, is still quite clear.
I've never believed that linking made your code necessarily GPL in the first place. I don't care what the FSF says, they're not exactly unbiased.
> I think you are specifically responding to my "does not work without" interpretation overly literally. Clearly if the program is written for and tested against a specific interface of a GPL'd program, it is intended to work with that program.
> On the other hand if it's written to call into some kind of standard interface, it no longer requires that GPL program specifically, but could work with any program implementing that interface.
Well that's basically how the standard already works. If your code is using a specialized enough interface, sharing data structures you got from the GPL code, then it's derivative of the GPL code and needs to follow the GPL.
So while "process boundary" is an inexact tool, your suggestion of "does not work without" doesn't seem significantly better to me.
I just know that, to me, "dynamic linking" seems like an arbitrary and imprecise way to define "derivative work". And, I'm not sure whether it's really something that _can_ be defined and possible to determine without consider it on a case by case basis. It's a good "right hand rule", perhaps, but doesn't strike me as either necessary or sufficient to really define it. We'll never really know, I guess, until someone makes that actual argument in court.
The part of where this submission and completion information involves a ring buffer mapped to kernel space is unique to Linux, I believe.
RIO is very similar and predates Linux version by many years: https://docs.microsoft.com/en-us/previous-versions/windows/i...
Main downside, that thing is only for sockets.
(or polling, since the possible high "interrupt rate" bottleneck of completion notifications is one of the things that motivated RIO)
With RIO and now io_uring, kernels map buffers to both kernel and user addresses spaces just once on initial setup, and reuse the same buffer for many I/O operations.
https://lwn.net/Articles/316806/
My approach has both the virtue and curse of being too simple to worry about all that, but it _does_ remove basic syscall overhead.
I think it has some bearing to those using eBPF to just batch calls, too. Unless I am missing something, I do not think there needs to be any super-user/root/capability restriction on syscall batching since all the syscalls check permission "on the inside". That gives it maybe more scope for applications.
That sys_batch is kind of a tiny "jump-forward-only" assembly language where you can use the output of prior calls in later ones. The jump forward only (no loops) I do should also guarantee termination { at least conditioned upon all syscalls terminating...but that's a whole other domain ;-) }. (EDIT: IIRC, the article that this conversation is about was excited about this aspect. In my examples/ I have an "mmap a whole file in one syscall" example.)
No ability to do such common things as concatenate strings from syscall results, or branch according to stat() output for example, makes for a severely limited interface.
io_uring is already in mainline, and it delivers syscall batching as a side-effect while bringing async to the table. I just don't see the point in adding another, severely limited syscall batching thingy, certainly not now, even moreso with talk about ebpf logic joining the party.
BTW as mentioned in a sibling comment, you might want to check out mingo's syslets proposition, which has similar naive batching, but was also async.
Not needing superuser/any special capability is also nice, though. Not sure of the current status/plans, but I am pretty sure eBPF needed root that for a very long time.
Anyway, I was not trying to "compete" with you or try to "get into mainline". A module works fine for me. Was just exhibiting an easy possibility about some points discussed.
EDIT: and thanks for the pointer. I will check it out.
EDIT2: and much like the word copy is a fake syscall, other fake syscalls like a "value test" could be added to forge a sort of if condition jump forward thing. My little repo there is more a proof of concept than anything else.
It's not like I have a dog in this race, I'm just another consumer of these kernel interfaces...
But I am happy something finally has landed upstream we can start writing generic userspace programs targeting and actually expect them to work on distro kernels in the future. But we probably still need a compatibility layer for emulating it in userspace, looks feasible.
Compatibility-layer-wise, I did actually do that for my batch system. In the tiny user-space entry point I check if sys_batch is working and if not I fall back to just a loop of userspace making syscalls. That also checks a BATCH_EMUL environment variable to force that emulation mode for benchmarking purposes { so I don't have to unload/reload the module. :-) }
So, user code would always just work, but work faster on kernels with the module loaded.
[0] https://github.com/facebook/folly/blob/16d6394130b0961f6d688...
One problem with io_uring is that it's completion based I/O where you move ownership of an buffer to the kernel which then writes to it until the operation completes.
This means you might not be able to (sync) cancel an operation occurring in the background.
This makes it harder to integrate into some I/O libraries, as the previous fact conflicts with RAII patterns.
Another think making adaption harder is that the interfaces for reading/writing with io_uring are conceptually slightly different.
Because of this e.g. Tokio currently hasn't switched yet to use io_uring but still uses readiness based async I/O as far as I know. (Which doesn't mean it won't support it in the future.)
This issue might be relevant (mio is internally used by tokio for async I/O): https://github.com/tokio-rs/mio/issues/923
[1]: https://crates.io/crates/rio [2]: https://docs.rs/rio/0.9.4/rio/struct.Completion.html#impl-Fu...
#define io_uring_for_each_cqe(ring, head, cqe) \
/* \
* io_uring_smp_load_acquire() enforces the order of tail \
* and CQE reads. \
*/ \
for (head = *(ring)->cq.khead; \
(cqe = (head != io_uring_smp_load_acquire((ring)->cq.ktail) ? \
&(ring)->cq.cqes[head & (*(ring)->cq.kring_mask)] : NULL)); \
head++)>"It’s beyond our scope to explain why, but this readiness mechanism really works only for network sockets and pipes — to the point that epoll() doesn’t even accept storage files."
Could someone say here explain why this readiness mechanism really works only for network sockets and pipes and not for disk?
What?!
To the point that as part of the Obama-Trump transition, a literal playbook was created for pandemics, with Coronaviruses (MERS-COV, SARS) explicitly mentioned:
* https://assets.documentcloud.org/documents/6819268/Pandemic-...
They had tabletop exercise on pandemics:
* https://www.politico.com/news/2020/03/16/trump-inauguration-...
> Friends of the government win state contracts at high prices and borrow on easy terms from the central bank. Those on the inside grow rich by favoritism; those on the outside suffer from the general deterioration of the economy. As one shrewd observer told me on a recent visit [to Hungary], “The benefit of controlling a modern state is less the power to persecute the innocent, more the power to protect the guilty.”
* https://www.theatlantic.com/magazine/archive/2017/03/how-to-...
TLDR: Any recommendations on the best way to clone one harddrive to another that doesn't take forever?
> Storage I/O gained an asynchronous interface tailored-fit to work with the kind of applications that really needed it at the moment and nothing else.
Say you have 2x 2TB SSD harddrives and one needs to be cloned to the other.
Being the clever hacker I am who grew up using linux I simply tried unmounting the drivers and trying the usually `dd` approach (using macOS). The problem: It took >20hrs for a direct duplication of the disk. The other problem: this was legal evidence from my spouses work on a harddisk provided by police, so I assumed this was the best approach. Ultimately she had to give it in late because of my genius idea which I told her wouldn't take long.
Given a time constraint the next time this happened, we gave up `dd`, and did the old mounted disk copy/paste via Finder approach... which only took only 3hrs to get 1.2TB of files across into the other HD - via usb-c interfaces.
I've been speculating why one was 5x+ faster than the other (besides the fact `dd` doing a bit-by-bit copy of the filesystem). My initial suspicion was options provided `dd`:
> sudo dd if=/dev/rdisk2 of=/dev/rdisk3 bs=1m conv=noerror,sync
I'm not 100% familiar with the options for `dd` but I do remember a time where I changed `bs=1M` to `bs=8M` helped speed up a transfer in the past.
But I didn't do it for the sake of following the instructions on StackOverflow.
It might be faster to have multiple rsync operations (via xargs or the like), but if the disk is relatively empty I can see this being faster. Finding the right level of parallelism isn’t something I can help you with, probably needs some experimentation.
dd if=/dev/nvd0 of=/dev/nvd1 bs=65536 skip=0 seek=0 count=1302349 status=progress &
dd if=/dev/nvd0 of=/dev/nvd1 bs=65536 skip=1302349 seek=1302349 count=1302349 status=progress &
dd if=/dev/nvd0 of=/dev/nvd1 bs=65536 skip=2604698 seek=2604698 count=1302349 status=progress &
dd if=/dev/nvd0 of=/dev/nvd1 bs=65536 skip=3907047 seek=3907047 count=1302349 status=progress &
dd if=/dev/nvd0 of=/dev/nvd1 bs=65536 skip=5209396 seek=5209396 count=1302349 status=progress &
dd if=/dev/nvd0 of=/dev/nvd1 bs=65536 skip=6511745 seek=6511745 count=1302349 status=progress &
dd if=/dev/nvd0 of=/dev/nvd1 bs=65536 skip=7814094 seek=7814094 count=1302349 status=progress &
dd if=/dev/nvd0 of=/dev/nvd1 bs=65536 skip=9116443 seek=9116443 count=1302349 status=progress &
dd if=/dev/nvd0 of=/dev/nvd1 bs=65536 skip=10418792 seek=10418792 count=1302349 status=progress &
dd if=/dev/nvd0 of=/dev/nvd1 bs=65536 skip=11721141 seek=11721141 count=1302349 status=progress &
dd if=/dev/nvd0 of=/dev/nvd1 bs=65536 skip=13023490 seek=13023490 count=1302349 status=progress &
dd if=/dev/nvd0 of=/dev/nvd1 bs=65536 skip=14325839 seek=14325839 count=1302358 status=progress &
dd if=/dev/nvd0 of=/dev/nvd1 bs=4096 skip=15628197 seek=15628197 count=6 status=progress &Io_uring can keep track of events through using event_fds, but yes, that is perhaps not optimal.
- kqueue only tells you there's data to read (/space in the write buffer). You still need to call read() or write(), including paying the cost of the syscall. io_uring lets you batch a lot of read/write calls together and either issue a single syscall to the kernel for all calls, or have the kernel poll and never syscall at all.
- kqueue doesn't let you issue fsync, or any of the other syscalls now in io_uring. fsync is essential on the write path for correctness in lots of cases, and for that you still need to dispatch to a local thread pool or something.
So yeah, I prefer kqueue over linux's epoll. But io_uring seems like the new king.
So, WASM isn't entirely unconstrained, and those constraints do allow some interesting analyses that true native code cannot support, e.g. symbolic execution: https://blog.trailofbits.com/2020/01/31/symbolically-executi...
By it's very design it's aimed at running untrusted, even malicious code without exposing the host to security risks, while allowing the imposition of resource constraints. It doesn't sound completely absurd to me (not an expert in this) that some project might seek to use those safety guarantees as a way to avoid paying for the more heavyweight guarantees provided by the kernelspace/userspace split.
A quick google find lots of people trying to get this to work; who knows - it might pay off.
No one even says "polygot bytecode", let alone says that WASM 'invented it'. Most programmers have already used something that uses a bytecode format, why would anyone say that?
> Joyful things like the introduction of the automobile
Cars cause so much pollution, noise, traffic, and take up so much space... How can you say its introduction is joyful?
About the new api: while I’m not very knowledgeable about the kernel, it seems like very good news for performance, the improvements are drastic!
Sure, if you want to sidestep the innumerable ways the automobile, or more accuracy the internal combustion engine, have completely revolutionised society, you could make an argument...
Except you can’t. Think about the increase in distance and speed at which goods and services can be rendered compared to prior modes of transportation. One simple example that comes to mind is the ambulance, in which that increase may well be the difference between life or death for an unknown but surely enormous population. A similar argument can be made for logistic supply chains delivering medicine, food, sanity products, waste removal, all of which without, mortality would surge. Plague, famine, disease are some of largest killers in human history.
Please tell me again how the absolute gains measured in human lifetimes is anything other than joyful, outside of subjectivity which is an unresolvable debate.
Horses generated literal tons of pollution in cities, were complicated to look after, and slow. The arrival of the car cleaned up the streets, made it easier to travel (and travel far) and very rapidly became the preferred mode of transport all over the world.
Not speaking was Kinvolk’s CTO Alban Crequy https://news.ycombinator.com/item?id=23042722
It is not fun to program against io_uring just as BPF can be a nightmare, but that's a problem for userspace to solve with libraries and better abstractions, much in the same way userspace usually doesn't have to deal with parsing the shared ELF segment exported by the kernel (glibc does that) or writing complex routines for fetching the system time (the ELF segment does that), both of which are implementation details necessary for extracting the full performance from modern hardware.
We'll catch up eventually, but first the lower level interfaces must exist. In meantime, I cringe Every. Single. Time. I run UNIX find or du over my home directory, realizing it could have completed in a fraction of the time for almost 10 years now if only our traditional software environment was awakened to the reality of the hardware it has long since run on.
Any inevitable standard will likely come from userspace, much in the same way libpcap successfully papered over raw packet capture for a huge variety of operating systems. The underlying kernel interface is basically just details. We're decades past the point where application portability required standards like POSIX to make progress, I imagine comparatively few modern programmers even know much about those kinds of standards any more.
We're _very_ good at buffering things up and handling large chunks to gain high throughput. We have layers and layers of magic that make us _very_ good at that.
For cases where you have lots of devices throwing small chunks at high frequency and you have to respond (in this one, do a small computation, and out that one...) we've sucked bad.
That's why bare metal RTOS's still exist.
The io_uring / eBPF combo looks _very_ promising for opening up that domain in a tidy fashion.
I also hope there will be no reason to throw away "heaps and mounds of existing and perfectly working code and solutions".
I hope to replace my finely tuned epoll / read/write reactor pattern inner loop with io_uring / eBPF and leave 99% of the code untouched... just better throughput / lower latencies.
I hope to do an update of Asio at some point and it will start to use io_uring automatically as a backend
Though they are great for some things.