423 karma · joined October 3, 2016
This is the kind of hostility (which is frankly toxic) that’s become associated with parts of the Rust community, and has fairly or not, driven away many talented people over time.
This is plain wrong, and it undermines the credibility of the author and the rest of the piece. Rust did not invent ownership in the abstract; it relies on plain RAII, a model that predates Rust by decades and was popularized by C++. What Rust adds is a compile-time borrow checker that enforces ownership and lifetime rules statically, not a fundamentally new memory-management paradigm.
Is the most interesting quote IMO. I often feel like productivity has gone down significantly in recent years, despite tooling and computers being more numerous/sophisticated/fast.
The language isn’t even stable, which is pretty much the opposite of something you can rely on.
We’ll know in many years if it was something worth relying on.
We are on our way there.
In some situations, the “logical” block size can differ. For example, buffered writes use the page cache, which operates in PAGE_SIZE blocks (usually 4K). Or your RAID stripe size might be misconfigured, stuff like that. Otherwise they should be equal for best outcomes.
In general, we want it to be as small as possible!
Newcomers often push back on this aspect of the language (among other things), but in my experience, that usually fades as they get more familiar with Go’s philosophy and design choices.
As for the Go team’s decision process, I think it’s a good thing that the lack of consensus over a long period and many attempts can prompt them to formally define a position.
All of this for a clock! I don’t get it, but I’m in awe.
It’s not surprising they didn’t see a linear speedup from splitting into so many crates. The compiler now produces a large number of intermediate object files that must be read back and linked into the final binary. On top of that, rustc caches a significant amount of semantic information — lifetimes, trait resolutions, type inference — much of which now has to be recomputed for each crate, including dependencies. That introduces a lot of redundant work.
I also would expect this to hurt runtime performance as it likely reduces inlining opportunities (unless LTO is really good now?)
First, significant work has been done in the kernel in that area simply because any gains there massively impact application performance and energy efficiency, two things the big kernel sponsors deeply care about.
Second, asynchronous IO in the kernel has actually been underinvested for years. Async disk IO did not exist at all for years until AIO came to be. And even that was a half-backed, awful API no one wanted to use except for some database people who needed it badly enough to be willing to put up with it. It's a somewhat recent development that really fast, genuinely async IO has taken center stage through io_uring and the likes of AF_XDP.
That only works when what you're trying to do has no side effect. Consider what happens when you need to cancel a write to a file or a stream. Did you write everything? Something? Nothing? What's the state of the file/stream at this point?
Unfortunately, this is intractable: you'll need the underlying system to let you know, which means you will have to wait for it to return. Therefore, if these operations should have a deadline, you'll need to be able to communicate that to the kernel.
UB is meant to add value. It’s possible to write a language without it, so why do we have any UB at all? We do because of portability and because it gives flexibility to compilers writers.
The post is all about whether this flexibility is worth it when compared with the difficulty of writing programs without UB.
The author makes the case that (1) there seem to be more money lost on bugs than money saved on faster bytecode and (2) there’s an unwillingness to do something about it because compiler writers have a lot of weight when it comes to what goes into language standards.
For (1) GC or not it doesn’t make a difference, I’ll opt-out. For (2) GC is really convenient and correct.
IOPS indeed matters a lot, but so does latency! For our use case, it was much easier to saturate those disks than the old i3s, and we attribute it to the better latencies, making IO scheduling a lot more accurate.
When the invariants breaks, the abstraction collapses and the code can become much more convoluted than it was originally. I’ve seen it many times.
In my career I’ve seen many many bad code abstractions and very few good ones. As measured by how long before they break.
I’ve asked the engineers that came up with those if there was a trick to it, and the answer has always been “dude it’s the 10th time in my career I’m writing that stuff”.
Good abstractions come from domain experience. If you’re writing something new, don’t abstract it. If you feel smart about it, that’s a bad sign. You don’t feel smart when you’re writing something for the 10th time.
Go made something super simple. I write and review a lot of Go code daily, and I don't quite get how these error branches are such a big issue. The code is always very simple to follow through.
Historically, buffered IO was sufficient to circumvent slow threads. Simply because buffered IO operations usually don't block (you merely memcpy, and the kernel flushes asynchronously in the background). That can only take you so far, however, and it's apparent today that hardware can go much further and that the gap is widening.
It's a valid point, though, to question whether that additional performance is even needed. John Ousterhout (https://www.youtube.com/watch?v=o2HBHckrdQc) is currently working on a new network protocol to alleviate some of the problems TCP creates in software wrt performance, and he also questions whether there is even a need for very fast networking for real-world applications.
IMO, the mere existence of stuff like DPDK is proof enough. Many folks use it and would rather use the kernel if it could provide comparable performance.
POSIX is outdated and problematic. It's outdated because it was conceived when hardware was vastly different than today. Storage used to be orders of magnitude slower than compute. Today, it's the opposite. POSIX APIs make squeezing all the juice out of the underlying hardware impossible. A standard duplex 100Gbps network link will carry ~300M small IP frames per second. You'll be lucky if you can do more than a few 10k/s using POSIX APIs. That's a massive bottleneck.
And it's not something that can be fixed with clever implementations. None of the POSIX APIs are async. So, to drive concurrency, programmers have to resort to threads that don't scale. That's a fundamental issue that no hardware improvement and/or software trickery will ever fix.
Today's reality is that software is seldom written against POSIX but against Linux, which offers many more APIs. Linux is, likewise, not formally standardized. It dodges the problem of S3 because it's open-source and ubiquitous. But that's not a good solution: it stifles innovation. For that reason, there hasn't been any new (production-ready, serious) kernel in decades.
We are at a deadlock, and POSIX is part of the problem.
> There are no downsides, except for confusing newbies.
False. Populating the page cache involves lots of memory copies. It pays off if what's written is read back many times; otherwise, it's a net loss. It also costs cycles and memory to keep track of all these pages and maintain usage statistics so we know what page should be kept and which can be discarded. Unfortunately, Linux makes quantifying that cost hard, so it is not well understood.
> You can't disable disk caching. The only reason anyone ever wants to disable disk caching is because they think it takes memory away from their applications, which it doesn't!
People do want that, and they do turn it off. It's probably the number one thing database people do because they want domain-specific caching in userland and use O_DIRECT to bypass the kernel caches altogether. If you don't, you end up caching things twice, which is efficient/redundant.
Can you tell us which? Go, Haskell and the other usual suspect all have runtime with automatic, transparent preemption.
Folks would rather have every future time sliced so that other tasks get some CPU time in a ~fair way (after all, there is no concept of task priority in most runtime).
But you're right: it isn't required, and you could sprinkle every loop of your code with yielding statements. But knowing when to yield is impossible for a future. If nothing else is running, it shouldn't yield. If many things are running but the problem space of the future is small, it probably shouldn't yield either, etc.
You simply do not have the necessary information in your future to make an informed decision. You need some global entity to keep track of everything and either yield for you or tell you when you should yield. Tokio does the former, Glommio does the latter.
It gets even more complex when you add IO into the mix because you need to submit IO requests in a way that saturates the network/nvme drives/whatever. So if a future submits an IO request, it's probably advantageous to yield immediately afterward so that other futures may do so as well. That's how you maximize throughput. But as I said, that's a very hard problem to solve.
This is used to keep track of task runtime quotas so they can yield as soon as possible afterward.
This is the same technique used in Go and many others for preemption. If you don't add this, futures that don't yield can run forever, stalling the system.
You are right that it is not strictly necessary, but in practice, it is so helpful as a guard against the yielding problem that it's ubiquitous.
> I certainly hope that we didn't end up with colored functions in Rust because of such a misconception.
Misconceptions are everywhere unfortunately!
But it also praises Go for its implementation, which is also based on a coroutine of a different kind. Stackful coroutines, which do not have any of these problems.
Rust considered using those (and, at first, that was the project's direction). Ultimately, they went to the stackless operation model because stackfull coroutine requires a runtime that preempts coroutines (to do essentially what the kernel does with threads). This was deemed too expensive.
Most people forget, however, that almost no one is using runtime-free async Rust. Most people use Tokio, which is a runtime that does essentially everything the runtime they were trying to avoid building would have done.
So we are left in a situation where most people using async Rust have the worst of both worlds.
That being said, you can use async Rust without an async runtime (or rather, an extremely rudimentary one with extremely low overhead). People in the embedded world do. But they are few, and even they often are unconvinced by async Rust for their own reasons.
Typically, if you want to build something with Rust, it'll have to use async, at least because gRPC and the like are implemented that way. So the vanilla (and excellent, IMO) Rust language doesn't exist there. Everything is async from the get-go.