Unbuffered I/O Can Make Your Rust Programs Much Slower
era.co
era.co
Making syscalls like write() every time you want to send even a character of output down a stream is much more latent than just sticking bytes in a buffer.
In general, the only time anyone should use unbuffered output is when there’s a real person consuming it - for example, when you’re prompting a user for feedback or input.
I think this might be a little -too- general, as it's pretty workload dependent. I mean, if you just want to write a blob of bytes once, just write the blob IMO.
I don't think using a buffer in those cases ever really hurts, but sometimes makes what should be simple a bit more complicated than necessary.
The worst I’ve seen was actually streaming the LOB on getLength. Since there’s no JDBC API for lob.getContentAsArray(), you need to call lob.getContent(1, lob.length()) (notice that you’d start at index 1…), which by default would stream the LOB twice under the scene!
As a side note, the real issue here is that unbuffered is the default. Other languages and libraries I’m aware of make it difficult to get in such a mode.
With that said, the OP could definitely have done a better job making it clear that this problem isn't specific to Rust
I'm used to a raw stream allowing a few specific operations:
- read/write/peek one byte - write a byte array, but you must explicitly specify length and offset
So seeing a call that simply takes an entire byte string and writes it gives me "buffered I/O vibes". It's hard to quantify, but with how ubiquitous that pattern is I'd be surprised if there are a lot of people who expect similar.
some_script.py | tee log.txt
But by default, Python will turn off line buffering when it detects it's not talking to a tty. So unless you know to call flush, or to add the right env var or interpreter flag, you're not going to see any early output from tee. This can be super confusing for newer programmers.Okay, I’ve been running into a bug in one of my Rust programs and I think this comment has the answer. My program uses a Command [0] to listen to stdin and stdout. Theoretically that Command should terminate when the parent process dies, but if the Command is a Python script that doesn’t include a `print()` with `flush=True`, it doesn’t die. I think you might have found my problem for me.
[0]: https://docs.rs/tokio/1.15.0/tokio/process/struct.Command.ht...
I'm in the camp that logging should be always unbuffered. If you buffer logs, and your process dies for any reason, you lose your log exactly when you really need it to know what happened.
(Also, not every environment is running with unlimited core size.)
A beginner learning C++ today will be getting familiar with buffered I/O primitives by the time they write a "Hello World"
Java mostly keeps the "useful" access patterns (anything more than reading/writing a single byte or manually passing in byte buffers) in buffered wrappers so even a clueless beginner is inclined to use them.
People are missing the forest for the trees here, it's about ergonomics and encouraging people to do the right thing by default.
Ironic to see this attitude in a Rust comment section of all places, isn't this the same attitude that leads people to dismiss Rust? "Just write correct C code without UB, the major pitfalls haven't changed in the last 4 decades duh?"
"Hello World" in C from 1986 uses printf() which is buffered.
So yeah, it's been a long, long, long time.
I think it's fine like this, because most of the time the language makes it easy to wrap/convert a writer into a buffered writer and use that instead.
Also, it preserves the principle of least surprise. If I call write on a Writer, I expect it to write immediately. When I want it to be buffered, I would either look for it in the stdlib (and in Rust's case expect to find it) or look for a lib (sans implementing it)
But you can't know what "to write" actually means.
When you call write(2) on a POSIX system, what is it that you think happens immediately? Patterns of NVRAM storage are changed? Fluctuating voltage on a cat-5 cable wobbles a bit in response?
OK, maybe you can get away with saying "If I call write on a Writer I expect it to hand over the data to the kernel immediately", and that's somewhat reasonable. The problem is that such a design encourages beginner programmers to use unbuffered IO, which is generally the wrong thing to do.
I'm used to a raw stream allowing a few specific operations, if it only allows:
- read/write/peek one single byte
- write a byte array, but you must explicitly specify length and offset
I assume it's unbuffered. Usually those are the building blocks of a buffered stream API after all. It also matches write(2) closely (offset replaces pointer math)
What surprises me is that it takes a full byte string and writes it with no hints about offset or length, usually that's only provided with buffered interfaces.
-
Again, there's no written rules about it, I'm just sharing my experience from writing code in a lot of different languages.
Someone else might have an experience with a different set and have no such expectation, so I'm not saying Rust is wrong, I'm instead pointing out why this isn't a "well duh" moment
Unbuffered APIs are meant to be tricky: They ask for length (even in languages where arrays have a fixed length like Java) because the idea is:
- you're either building out a buffered wrapper and can't rely on the encoded length value (because your array length will be the entire buffer while you might only want to write part of it)
- or you want very precise control over the writes and are willing to risk those mistakes. So they need to open up that risk of a mistake.
It's an API that mirrors to the user how the underlying syscall operates, so the underlying complexity is being very plainly exposed.
-
Usually it's the buffered wrapper that will have a convenience method that takes a String or a full array without additional information, or there'll be an explicit call to force a buffered wrapper to use a 0-sized buffer.
C++: Without explicit call to setbuf, system defined default buffer used (matches POSIX behavior)
Java: FileWriter vs FileOutputStream, FileOutputStream is unbuffered will not take a string without manually encoding it to bytes first, FileWriter will take a string and is buffered
Python2: need explicit 0 sized buffer
Python3: need to flag as unbuffered and automatically converts file to binary mode, so calling write("SomeString") will not be allowed with unbuffered output.
C#: FileStream buffered by default, must explicitly ask for 0-size buffer
Obviously this won't hold across every language. But you can see that languages consistently make it easier to do the "right" thing in this case, while Rust isn't.
I'm not necessarily holding it against Rust either to be clear, but I'm holding it against the droves of comments here that are acting like this is such natural behavior you'd be a rookie to not catch it... it's been a loooong time since most people have used a language where that was "easy" behavior to uncover accidentally. Like I said even C++ went out of it's way to improve things in this regard.
Rust's solution to the problem, even if a departure from traditional solutions, is more ergonomic, and I like that.
There's no reason buffered or buffered IO needs to be trickier than the other! Why sacrifice one? Sure you could have Write buffered by default and have UnbufferedWrite instead, but I personally find this "reversed API" idea really weird. It seems like a subtractive API rather than an additive one, and it also doesn't compose as nicely.
Your argument about size proves my point too, so we're going in circles I guess. I'll summarise why I prefer Rust's approach (sans ergonomics): it's harder to introduce logic errors if all the information can be inferred from the value itself. If I want to send "hello", and somewhere I fucked up, so I'm thinking 4 bytes instead of 5 and now I sent "hell", and threw 'o' to the RAMs. Why? I can just tell it I wanna send "hello", and it does it. It's more robust. If I wanted to send just "hell" (assuming it's bytes) in I can just create a slice to the necessary data and pass that. The slice explicitly itself signals to other devs I wanna send a subset of the buffer, too.
Doing it this way, buffered or unbuffered doesn't really matter: The API stays the same. Adding buffering is usually a one line change of converting your writer to a buffered writer. Someone called Rust's approach an "arcane incantation", if anything the only magical things here is the special/magic sized buffers or numbers in other languages you highlighted. Furthermore, I think half of the existing implementations probably didn't even think about this design decision at all, they just inherited it.
At the end of the day, this discussion about buffering isn't super important in the grand scheme of things. The stdlib gives you just building blocks (Rust has great ones, IMO) to build a networking library you want, so your (un)buffering is going to be tucked away in a couples of places and you'll be using your own `report` or `send` or `set` methods instead of a sprinklings low-level write/read syscals everywhere anyway.
It also doesn't take a string; it takes a slice of bytes. In Rust, prefixing a string literal makes it a `&[u8]` instead of a `&str`, which is why you can just pass it straight into `write`.
Adding a length field would only make it more error-prone to use without any added benefit. Slices already know their length, and subslicing is trivial. If you have a larger slice, and you only want to write a subsection of it, you just index it. Forcing the user to pass in two lengths is just silly, and all the API implementations would do is subslice the input to that second length anyway.
You could argue that `File::open` and `File::create` should return a `BufReader<File>` and `BufWriter<File>` respectively, and I would agree that this would be reasonable as the most common use-case probably benefits from buffered IO. However, that ship's already sailed, and those functions can't be changed.
The description for the `File` type should also more explicitly state that IO is unbuffered. I'm surprised it doesn't, the docs are typically better at that sort of thing.
My point isn't that every language should force that syntax ('offset' for example, doesn't make sense in languages with pointer math for example)
Rather the point is that when the language doesn't follow that convention, it should buffer by default. I guess it is a blind spot in the standard library then
Java. There
In that regard it got turned weirded with multiple buffers which slows down as well, just on as much as totally unbuffered. Java 1.4 added nio where the buffers were part of the framework.
If you try to write a string to a FileOutputStream the way that Rust code did this is what it looks like:
byte[] byteString = "Hello World".getBytes(StandardCharsets.US_ASCII);
myStream.write(byteString, 0, byteString.length);
But if you use a FileWriter, which has a built-in buffer and wraps the OutputStream: myWriter.write("Hello World");
For me, if the Rust example had required a length and offset, I'd immediately assume it was unbuffered. I keep clarifying: I realize it's not a rule, but I guess that's just a pattern I've seen across a lot of languages. Maybe because it somewhat matches write(2)?However, in my experience the writers are not very commonly used which is a stark contrast to BufferedReader that's everywhere (it comes with readLine()). Instead, writing is commonly implemented via encoding/conversion/etc. to a byte[] and then written directly to the output stream.
Also there is a lot more then mentioning cpu cache as this is very generic term. which part of the cpu cache are you talking about? instruction cache? data cache? And why do you have think they matter in this context?
It is important to note that every time buffered I/O is used, it copies data through buffers, and touching those buffers means the memory used for those buffers being touched clearly has to be pulled into the CPU's L1 data cache. In some cases the extra cache traffic for those buffers degrades performance. Instruction caches are relevant to some workloads, but are clearly not what I was referring to since data doesn't go through the instruction cache.
glibc further complicates the matter by using mmap() in some cases to directly map the kernel's buffers for data (the page cache in the context of Linux) into userspace. This avoids some of the downsides of buffered I/O by avoiding one of the copies at the expense of page fault overhead. Sometimes it's better for the application itself to use mmap() directly.
Direct use of read() and write() makes sense for some applications, especially when dealing with larger chunks of data. Saying unbuffered I/O is bad is a gross oversimplification of the reality when writing high performance applications. Like everything in programming, it's a good rule of thumb, but there is definitely an asterisk next to it.
Not only does using plain Read make this mistake possible, but it also means they have to be doing byte-by-byte processing. [1] So not really possible to do SIMD-oriented parsing acceleration like simdjson uses. (I gather simdjson's approach is quite sophisticated, but a simpler version is to scan for the next delimiter via a memchr2 on " and \, then have a routine that validates the characters until then are valid via SIMD methods.)
[1] or do larger reads into their own buffer. But they're clearly not doing that given this strace output. And it'd be redundant when folks do use BufRead.
The converse is potentially disastrous - if you read too much from a TCP stream you could hang your program waiting for more input that never comes.
Recommended by whom?
When you want to process stuff in bulk, it's better to use BufRead::fill_buf (reusing the BufRead's buffer) rather than copy into your own.
For example, I wrote some code which basically skips some escape bytes. [1] It wraps a BufRead and is itself a BufRead. Its caller actually can process bytes straight from the buffer of the BufRead you supply, skipping two layers of copying.
> The converse is potentially disastrous - if you read too much from a TCP stream you could hang your program waiting for more input that never comes.
You can implement any idea badly, but that doesn't mean the idea itself is bad. There's a difference between taking advantage of buffering and expecting to operate only on full, fixed-sized buffers.
[1] https://github.com/dholroyd/h264-reader/blob/60ed66dc4dbfe74...
* if the caller supplies a BufRead, they can use fill_buf.
* if the caller supplies a Read that isn't BufRead, they can wrap it in a BufReader or the like, and then use fill_buf on that.
Stable Rust doesn't have specialization today, so at most one of those can be handled optimally. (It looks like neither of them is handled well at all right now.)
Maybe they're designing the simplest possible API for post-specialization. They might have been a lot more optimistic when they chose it than I am now about when Rust will get specialization. Or they might be heavily prioritizing API stability for years.
Right now this interface is what I would call an attractive nuisance, leading to the problem in this blog post.
What I would do for today is to have a method that explicitly takes a BufRead and uses it appropriately. And then perhaps a convenience method that explicitly wraps a Read in a BufRead. And the caller chooses. I'd deprecate the attractive nuisance interface for now (maybe reversing that later).
It would have been IMHO probably less clean but safer to make the standard IO constructs like File buffered by default (which is what users expect in the vast majority of cases) and provide on the side unbuffered variants. Thankfully modern OSes with caches and stuff somewhat mitigate this thing, but it can make IO on less advanced systems way more problematic than it should be.
1. Didn't know about a nuance. 2. Didn't care about said nuance.
Creating magic incantations to achieve typical expected behavior will result in many libraries doing the unexpected behavior. Which means that the typical code that any end user works with will probably do the unexpected thing.
The fact that Rust goes in the opposite direction makes sense from a logical standpoint, but as the article itself shows will almost certainly end up causing performance issues. You can rest assured that people *will* forget to buffer their files because it is something they almost never had to care about in the past.
What's magical here? Rust makes it there's no need for magic. Most things like these are designed to compose. Here it's no different, and certainly not any more magical.
That programmers need to consider this stuff on each new project is a failure of the standard library design, it's a problem not even C++ suffers from
You've also gotta be careful with unexpected and implicit assignment/copy/move constructors.
More than half of all the C++ wisdom one needs in their career is when to avoid stdlib and which part of boost to use. Due to C++s religious devition to 80s compatibility, it becomes a minefield of a Don'ts, with a narrow path of Dos.
The fact boost is a billion libraries in one speaks in of itself about the abysmal dependency situation in C++. There's also a buncha dep managers that are at each others throats. Also, C++ people like to poke fun about js devs and npm/leftpad, but forget that they use boost for even more rudimentary string operations. Or they implement the same thing over and over again. Or my favourite: copy+paste the same code in multiple projects, and then have divergent implementations of the same thing with different bugs and/or API surface; bonus points if the thing is part of an RPC implementation.
As much as the edgy-dev crowd loves to belittle the very concept of dependency management, at least JS/Ruby/Python ecosystems give a shit about DX and made sure the package managers can be used with the same repositories/sources.
But hey, at least C++’s IO is buffered by default. Amazing.
I got distracted by my rant above: I actually PREFER the default is unbuffered, because it doesn't violate the principle of least surprise (if we remember that Rust is a systems programming language). The fact you expect the default to be buffered is probably due to exposure to C/++.
There are no dep managers for C++ that have any kind of universality to them.
To reiterate for the umpteenth time: many C and C++ libraries come with compile-time flags that totally change the nature of the resulting library. It is not possible to refer to a "canonical" version of the library the way you could with something for Java, Python, JS etc.
When I say "this code relies on libfoo", it is necessary to know if there are build/compile flags for libfoo that could change the nature of libfoo.{a,so}. If there are not (common, but not universal), great. If there is just one, probably manageable. If there is more than one, game over.
Also, only C++ iostreams are buffered by default. Just like in C, where you are free do IO with syscalls if that floats your boat. There's nothing more or less obvious about this than in C, where stdio vs. syscalls is an essentially identical choice.
That just iterates GP's point: C/C++ dependency management is abysmal, but it's "hard" because the languages themselves do not have the capability to express the required metadata in any standard way. As an example, the first time I tried to compile Ardour took me several hours finding and installing the right dependencies for my distro.
PL designers have learned from these mistakes, it's not worth defending them or accepting that it is the way the world should be...
Of course the languages don't. It's not a part of the language. If I build libfoo with -Dexclude_feature_bar which just drops 2 source modules from the library build, how can the language express that?
On Debian, you get the Ardour dependencies like this: apt-get build-dep ardour6 If your distro makes it hard than that, that's a distro problem.
Of course, the libraries Debian or any other distro will install when you do that are not the ones we use to build Ardour, because we have patches for several libraries that will never be accepted upstream but are necessary/useful/important for Ardour. You can't express that with the language either.
You cannot use Cargo or Rust to describe how a C or C++ library was compiled and linked.
> If I build libfoo with -Dexclude_feature_bar which just drops 2 source modules from the library build, how can the language express that?
By expressing within the language that the library can be configured with an optional `exclude_feature_bar` flag. Then downstream dependencies can express they require importing libfoo configured with exclude_feature_bar.
At least in Rust where the compilation unit is a crate, crates must express what their compile options are, and dependents express that they require specific compile options from their dependencies. You could argue this isn't a property of the language, but I would argue it is since it's so tightly coupled to the language implementation (despite living in a config file, that file is required to build anything)
Since C and C++ eschew this fundamental requirement of packaging software and pass the buck to the library authors to figure out a way to describe available configuration options, there's no standard way to do this, and dependency management becomes quite difficult. That doesn't mean that other language authors can't solve the problem. And it plays into what the original comment said, which is that managing dependencies in C/C++ is abysmal.
1. use double or float for computation
2. thread-safe or not
The result is 4 possible versions of libfftw: double + thread-unsafe
float + thread-unsafe
double + thread-safe
float + thread-safe
What aspects of rust's crate design will allow me, developer of an application that uses libfftw, to define which of these I want to use? How many crates will I have installed in order to switch between any of them?However, if fftw was wrapped in a Rust crate, the manifest can express those using features in the Cargo.toml file [1] which can be read in a build.rs[2] script which would provide the proper configurations to libfftw at build time for all dependents that specify which options it needs. In this model you won't need more than one crate installed for libfftw, since the crate expresses everything needed. Crates are not static collections of object files like C/C++ libraries.
[1] https://doc.rust-lang.org/cargo/reference/features.html
[2] https://doc.rust-lang.org/cargo/reference/build-scripts.html
The crate design you've described works as long as all your dependencies are wrapped in a "crate" of the correct type (in the current case of Rust, a Cargo crate).
But is it even close to realistic to imagine that there will be correct "crate-like" encapsulations of every C/C++ library, thus enabling them to be used seamlessly from each language with a "crate-like" system? I don't know, maybe this is genuinely likely, at least for Rust. To me, having to assume that the hundreds of millions of lines of code that already exist must be and will be wrapped in a crate-like system seems like a bad bet, and I'd rather have a build system that doesn't require this to be true for me to use any pre-existing code.
The model you are describing for crates seems to also break down for 3rd party non-open-source libraries, where (re)compilation as part of the build process for my own application is not an option. In such cases, it's hard to see how proprietary developers would be content with anything other than "a static collection of object files" for their licensable libraries. But maybe I'm wrong about that too.
This is not true for most languages. I can use a build system that has its own options, and in the build specification I can say:
if (defined (build_option1)):
sources += source_file1
else
sources += source_file2
The compiler and linker have no idea that there is any choice between the two source files, or that the build time option exists.The two files can each provide the same functionality (e.g. hand-coded SIMD asm, but for different target systems); nothing in the language knows anything about this.
() or de facto standard build tool like Cargo
But that said the example is a little contrived since you wouldn't use a compile configuration to exclude a source file in most languages that have more flushed out module structures than C
The example comes specifically from the build system for Ardour, where we have hand-written assembler and will incorporate different files into build depending on the target (SIMD) architecture and expected runtime environment. I would not call that contrived.
I have no doubt at all that Cargo can handle many, even the majority, of projects of a certain scale. I have very strong doubts that a build system so centered on a particular language and with so many assumptions about how things are written and built can work for all projects. It's a problem when the build system falls down on the most demanding stuff, because it leads to people building "meta-build systems" around it, and we've seen how that turned out in the past.
I think you're misunderstanding a lot of what I'm describing, it has nothing to do with Cargo or Rust but the fundamental limitations of C/C++ that explain why dependency management that involves either of them is extremely painful compared to any modern language.
One major philosophy aspect of Rust is that (almost)everything is explicit and you build abstraction from the bottom.
There are also other reasons for designing it this way such as composability. In general it’s bad idea to try to undo an abstraction, and instead one should build on top of primitives instead.
"Veteran developers", the sort of who would use Rust after having worked in C/C++/etc would know very well that buffering on very granular IO is a must except some specific situations. Otherwise they're anything but. And it has nothing to do with Rust. This is of course when IO libraries used do not already take care of that.
It might be interesting to have a version of a Writer trait that handed ownership over in addition to supporting slices and handled things under the hood for you. The API may be too complex to bother with. Buffered I/O is typically a fine middle ground. When you start needing writev you’re changing the I/O model more drastically.
https://doc.rust-lang.org/src/std/sys/unix/fd.rs.html#77
https://doc.rust-lang.org/src/std/sys/unix/fd.rs.html#144
Buf(Reader|Writer) are vectored I/O aware, bypassing the internal buffer whenever the total length of the provided buffers is larger:
https://doc.rust-lang.org/src/std/io/buffered/bufreader.rs.h...
https://doc.rust-lang.org/src/std/io/buffered/bufwriter.rs.h...
That said, I can guarantee rate I have missed such issues in my code both as the writer and the reviewer. Saying "it's in the documentation" doesn't mean it's inexcusable to make 5he mistake.
Btw, pointing to documentation and saying "see, it's documented!" Every time something has lame or broken or unexpected default behavior doesn't make it any less lame or broken or unexpected. Once you feel like you know how to use something, you aren't always checking the docs to get a refresher on the caveats. So tone it down with the accusations and the "helpful" links, everyone else has access to the same Google you do.
Also, senior people are more likely to make the kind of mistake explained in the article than less experienced developers. This is because:
- They are more focused on the "big picture" (algorithm correctness, general approach, system architecture, etc.) than on minute details.
- They tend to read less documentation because they don't need to reach for it so often.
- They have learned that a mistake such as this one can be easily corrected, and hence dedicate their attention to more important matters (avoiding data corruption for instance).