Rust Performance Pitfalls
llogiq.github.io
llogiq.github.io
I believe this needs a stronger warning. Functions that operate on strings are allowed to assume that their input is valid UTF-8 and may perform out-of-bounds memory accesses if the input is not valid UTF-8. Therefore, creating a string containing invalid UTF-8 can lead to a memory safety vulnerability. The suggestion to use this function seems out-of-place in this article, which otherwise avoids suggesting unsafe code (e.g. indexing arrays without bounds checking).
steveklabnik's profile: "Rust core team"
I'm guessing you know that a lot better than I do lol.
Really? I think you mean the opposite here.
The correct response when warning people of potential problems is never "oh, they should be able to figure out whether this applies to their case".
All the solutions and problems in computing stem from the fact that we're programming completely deterministic Turing machines that do exactly what they're told.
Just saying "or doesn't" is not a useful description. The vast majority of horrible broken behaviors fit inside of that bucket, where it doesn't segfault but then goes on to do the wrong thing at an unexpected place.
And because we're using an optimizing compiler, you can't give a simple description like "overflows a buffer". When the compiler assumes your code is correct, it might output instructions that behave in 'impossible' ways when fed invalid data. For example an if/else that takes neither branch, or both branches, or code that verifies a number has a certain value yet outputs a completely different value. You can only make assertions about what a particular compile will do, and that's obsolete information immediately.
This is exactly why warnings should strive to be as clear as possible. Just because you think it's not ambiguous, doesn't mean you aren't just missing some context that someone else might have that makes the statement somewhat ambiguous.
Good argument. If you advise someone to use unsafe the responsible thing to do is explain precisely what the consequences are besides "it's faster."
Literally means "something violates the safety model".
---
Here is the truth about safety guarantees. You can do a lot valid things type systems prevent you from doing. You can do a lot of valid things if/else/for/while prevent you from doing.
But 99% of the time is isn't worth it. Following the rules, even with strange BS they create is easier than managing the mess of GOTO's you'll find yourself in a year or two.
Yes. But the safety model is not capable of identifying whether something is or is not safe in every case, just that it can't prove that it is safe. There are plenty of things you must do in unsafe that are safe, just not provably by the compiler. Thus, needing to use unsafe does not always imply actual unsafe operations.
If someone was under the impression that they were recommended to use unsafe but in a safe way but there were other caveats to the usage they weren't aware of, bad things could happen. I don't believe it's sufficient in a guide to rely on the fact that unsafe is recommended to convey the level of danger that recommendation entails.
This is mistaken, and an impression that we continually strive very hard to counter. The `unsafe` keyword is to be used in the process of writing (and therefore consuming) APIs if and only if those APIs have external unenforced invariants which, if broken, could cause memory unsafety. The reason that we strive to reinforce this so fervently is precisely because we want people to see the `unsafe` keyword and instantly become wary of memory safety violations (and undefined behavior in general); conversely, we want people to be able to view the absence of `unsafe` in code that they have written and have confidence that that code is memory-safe.
In particular, this means that `unsafe` is not to be used for operations that may be dangerous but that have nothing to do with memory safety. If a misused API could delete the production database, that's not unsafe. If a misused API could leak all your users' passwords, that's not unsafe. We had this argument before 1.0 regarding a few things in the stdlib that are more-or-less "dangerous" (e.g. `std::mem::forget`) that cannot be used alone to cause memory errors and arrived at the current strict interpretation deliberately.
Though, of course this requires social pressure to enforce (which is exactly what I'm doing here), as by definition unsafe code is something that the compiler itself cannot reason about.
> This is mistaken, and an impression that we continually strive very hard to counter.
> In particular, this means that `unsafe` is not to be used for operations that may be dangerous but that have nothing to do with memory safety.
That's not what I'm talking about. I think you've misinterpreted my point. I'm talking about unsafe for memory access, but in ways that are provably (or it not provable, are accepted as) safe. For example, the standard libs that are unsafe under the covers, but expose a safe API.
Saying "you must use unsafe to accomplish this" is not equivalent to saying "this can cause memory unsafety". That latter is a subset of the former, and the only time that isn't true is if Rust is capable of completely identifying every case of safe memory access and only requiring unsafe for actual unsafe operations.
Since unsafe could be required for what is an entirely safe, and possibly provably so, set of actions, it's dependent on what the context the recommendation is whether you believe someone stating an action requires unsafe implies that actual problems could occur.
If you're trying to say "there exists Rust code that contains `unsafe` blocks that is memory-safe", then this is obviously (hopefully!) correct, because 100% of the time we hope that our `unsafe` blocks are correctly implemented. In the same sense that "valid" C code does not contain undefined behavior, "valid" Rust code doesn't either; therefore the correctness of `unsafe` blocks in Rust must be tautologically (if uselessly) guaranteed. But that's not a very useful statement. Bugs happen.
Conversely, if you're trying to say "there exists code that can only be written in `unsafe` blocks but that can never cause memory unsafety despite any modification to the surrounding code", then I'd say this is trivially refutable; I can modify any unsafe block to exhibit memory unsafety.
Finally, if you're just saying "static analysis must, by its very definition, reject some correct programs in order to guarantee that the programs it does accept are correct, and Rust's static analyses are no different" then, of course, this is true (again, by the nature of static analysis), but again this is not an especially useful statement, especially since this isn't what anyone here is disputing. Your original comment was made in reply to this statement by rabidferret: "`unsafe` literally means "this might violate memory safety". This is a true statement, both in a social context (see my original comment) and in an implementation context (see the second example from this comment).
The bottom line is, if you see an `unsafe` block, assume memory safety is at risk unless you're absolutely sure the author knows what they're doing.
Yes, and additionally there are some algorithms that are safe, but to implement require unsafe blocks. I think it's obvious that there are patterns of memory access that are safe, possibly provably so, but not by the compiler at this time.
Since a recommendation for someone to use unsafe could be either a general recommendation with a very broad caveat that quite a bit of additional care and thought, or a fairly benign recommendation to "do X, Y, and Z in this order, and while it requires unsafe, it's no less safe than when we do the same in C/C++", I think it's important to distinguish those, lest someone mistake the former for the latter.
What that boils down to in practice is that stating something requires unsafe is not sufficient by itself as a warning to denote the level of care someone should take. Different people will interpret that statement differently at different times on different topics. There is no need to leave that ambiguity standing when it is easy to clarify. That's all I was trying to express, in response to '"requires unsafe" already implies everything in your comment.'
I believe kibwen's point is that every algorithm is safe (when implemented correctly), since any memory safety violations inside `unsafe` blocks are incorrect. The keyword indicates the compiler can't guarantee that there isn't any, not that memory unsafety is okay. Maybe an example of the sort of thing you're thinking of would clarify the distinction you're drawing.
> I think it's obvious that there are patterns of memory access that are safe, possibly provably so, but not by the compiler at this time.
This is true of "everything": given any task X, a sufficiently smart language/compiler could allow expressing it without the risk of memory unsafety (i.e. no use of `unsafe`).
One, has, for the most part a self contained implementation and is part of a very well known algorithm. Checking that the implementation doesn't have any obvious flaws may be sufficient.
The other has implications that far outlive the small suggested bit of code. Until you've correctly made sure that anything resulting from this call has been confirmed to conform to UTF-8, there is a risk in it's use in any number of string processing routines.
The article, to it's credit, does mention that problem with UTF-8 when working with it unchecked, albeit vaguely. What I don't think would have been sufficient for an article that aims to help people would be to say "it required unsafe" and leave it at that. That would be suggesting a routine for performance reasons without sufficiently addressing the real downsides.
Thus, my assertion, that simply noting that something requires unsafe is sufficient to denote in all cases the consequences of the suggestion, as I interpreted rabidferret comment to imply.
Another way of stating my argument is "when suggesting unsafe as a possible solution, it behooves you to mention any non-obvious consequences this specific suggestion might entail." That's a fairly uncontroversial view, in my eyes, so I'm not sure exactly why I've had to explain it four separate times now.
1: Like in the rust standard sort algorithm.
Violating memory safety in unsafe code is UB.
So, unsafe Rust is a superset of safe Rust. Adding `unsafe` around some code lets you do four things:
* Dereferencing a raw pointer
* Calling an unsafe function or method
* Accessing or modifying a mutable static variable
* Implementing an unsafe trait
That's it. Nothing else changes, you get these additional abilities. This is very important, conceptually. Tons of other checks are still on, etc.
With that in mind,
> (What would happen if you put only code that could be verified by the compiler in an unsafe block?)
It would function identically.
UTF-8 is a fine transport format, but for raw runtime performance it's obviously going to be an issue if you ever need to iterate over characters, do substring matches, things like that because you can't do constant time "next char" or indexing.
UTF-16 doesn't let you do that either in the presence of combining characters, but they're pretty rare and for many operations it doesn't really matter.
A better suggestion is to rethink why you need those operations in the first place.
UTF-8 is certainly not a problem for runtime performance. Substring search, for example, is as straight-forward as you might imagine. You have a needle in UTF-8 and a haystack in UTF-8, and a straight-forward application of `memmem` will work just fine (for example). In fact, UTF-8 works out great for performance , because it's very simple to apply fast routines like `memchr`. e.g., If you `memchr` for `a`, then because of UTF-8's self-synchronizing property, any and all matches for `a` actually correspond to the codepoint U+0061.
Indexing works fine so long as your indices are byte offsets at valid UTF-8 boundaries. Byte offset indexing tends to be useful for mechanical transformations on a string. For example, if you know your `substring` starts as position `i` in `mystr`, then `&mystr[i + substring.len()..]` gives you the slice of `mystr` immediately following your substring in constant time. When all your APIs deal in byte offsets, this turns out to be a perfectly natural thing to do.
Generally speaking, indexing by Unicode codepoint isn't an operation you want to do, because it tends to betray the problem you're trying to solve. For example, if you wanted to display a trimmed string to an end user by "selecting the first 9 characters," then selecting the first 9 codepoints would result in bad things in some circumstances, and it's not just limited to the presence of combining characters. For example, UTF-16 encodes codepoints outside the basic multilingual plane using surrogate pairs, where a surrogate pair consists of two surrogate codepoints that combine to form a single Unicode scalar value (i.e., a non-surrogate codepoint). So if you do the "obvious" thing with UTF-16, you'll wind up with bad results in not-exactly-corner cases.
It's worth noting that Rust isn't alone in this. Go represents strings similarly and it also works remarkably well. (The only difference between Go and Rust is that Rust's string type is guaranteed to contain valid UTF-8 where as Go's string type is conventionally UTF-8.) Notably, you won't find "character indexing" anywhere in Go's standard library or various Unicode support libraries. :-)
I would very strongly urge you to read my link in my previous comment to you. I think it would help clarify a lot of misconceptions.
I do have my unrelated niche complaints about Rust's string story (and I have vague plans to resolve them), but Rust's string implementation is my favorite among any other language I've used.
Writing a iterator to provide the individual bits supplied by an iterator of bytes means you can count them with
fn count_bits<I : Iterator<Item=bool>>(it : I) -> i32{
let mut a=0;
for i in it {
if i {a+=1};
}
return a;
}
Counting bits in an array of bytes would need something like this let p:[u8;6] = [1,2,54,2,3,6];
let result = count_bits(bits(p.iter().cloned()));
Checking what that generates in asm https://godbolt.org/g/iTyfapThe core of the code is
.LBB0_4:
mov esi, edx ;edx has the current mask of the bit we are looking at
and esi, ecx ;ecx is the byte we are examining
cmp esi, 1 ;check the bit to see if it is set (note using carry not zero flag)
sbb eax, -1 ;fun way to conditionally add 1
.LBB0_1:
shr edx ;shift mask to the next bit
jne .LBB0_4 ;if mask still has a bit in it, go do the next bit otherwise continue to get the next byte
cmp rbx, r12 ;r12 has the memory location of where we should stop. Are we there yet?
je .LBB0_5 ; if we are there, jump out. we're all done
movzx ecx, byte ptr [rbx] ;get the next byte
inc rbx ; advance the pointer
mov edx, 128 ; set a new mask starting at the top bit
jmp .LBB0_4 ; go get the next bit
.LBB0_5:
Apart from magical bit counting instructions this is close to what I would have written in asm mysef. That really impressed me. I'm still a little wary of hitting a performance cliff. I worry that I can easily add something that will mean the optimiser bails on the whole chain, but so far I'm trusting Rust more than I have trusted any other Optimiser.If this produces simiarly nice code (I haven't checked yet) I'll be very happy
for (dest,source) in self.buffer.iter_mut().zip(data) { *dest=source }However, the compiler does have perfect aliasing info, and it does provide a subset of this info to LLVM (I don't think LLVM IR currently supports more fine grained aliasing info being provided from the compiler). This does help certain optimizations; though I'm unsure if this one is one of them.
You didn't specify what the type of `buffer` was, so I picked a `u8`. This code[1]:
pub struct Thing {
buffer: Vec<u8>,
}
impl Thing {
pub fn copy(&mut self, data: &[u8]) {
for (dest, &source) in self.buffer.iter_mut().zip(data) {
*dest = source;
}
}
}
Produces this assembly: _ZN10playground5Thing4copy17hf523bcb10e2298f3E:
.cfi_startproc
pushq %rax
.Ltmp0:
.cfi_def_cfa_offset 16
movq 16(%rdi), %rax
cmpq %rdx, %rax
cmovbeq %rax, %rdx
testq %rdx, %rdx
je .LBB0_2
movq (%rdi), %rdi
callq memcpy@PLT
.LBB0_2:
popq %rax
retq
The call to `memcpy` is what makes me happy.----
Your `count_bits` already exists as a combination of iterator adapters (`filter`[2] and `count`[3]):
a_bit_iterator.filter(|bit| bit).count()
Although if you have a numeric value, I'd suggest using `count_ones`[4], which can use the `popcnt` intrinsic.If you wanted to count all the bits in an array, I'd suggest `map`[5] and `sum`[6]
let p = [1u8, 2, 54, 2, 3, 6];
let result: u32 = p.iter().map(|b| b.count_ones()).sum();
If you wanted to keep the bit iterator, you could also use `flat_map`[7].[1]: https://play.integer32.com/?gist=03f8ffbe3ade6ced4d315c8e020...
[2]: https://doc.rust-lang.org/std/iter/trait.Iterator.html#metho...
[3]: https://doc.rust-lang.org/std/iter/trait.Iterator.html#metho...
[4]: https://doc.rust-lang.org/std/primitive.u8.html#method.count...
[5]: https://doc.rust-lang.org/std/iter/trait.Iterator.html#metho...
[6]: https://doc.rust-lang.org/std/iter/trait.Iterator.html#metho...
[7]: https://doc.rust-lang.org/std/iter/trait.Iterator.html#metho...
That's the sort of thing I was hoping to see.
>Your `count_bits` already exists as a combination of iterator adapters (`filter`[2] and `count`[3]):
That's the problem with simple examples. I don't actually want to count bits. It was just the minimum workload I could think of to generate a result from the conversion.
Seeing how the for (a,b) in ai.zip(bi) works well I'll probably be writing a bitmap glyph renderer that is basically if b {*a=color}
I actually wrote about it last week: http://tmccrmck.github.io//post/rust-optimization-partii/
But if you find popcount too "magical", the commonly-known fast way to count bits is via masking, shifts and adds, so that you do it in log(n) steps. Which also would perform much better than this solution.
So what you're really saying is "the compiler managed to make a pretty efficient representation of the naive solution" which is fine but it does not mean your code is fast.
What do you consider a "modern CPU"? Atom chips sold less than a decade ago didn't support popcnt. AMD shipped some C-series chips without support for it as recently as 2012. The low-end chips were likely to be sold later without refreshes to newer features in some cases, and those are also likely the ones to be repurposed for small and cheap x86 devices later. I wouldn't want the default compilation settings without specifying CPU extensions to use an extension that might not exist on my target platform.
let nopes : Vec<_> = bleeps.iter().map(boop).collect();
let frungies : Vec<_> = nopes.iter().filter(|x| x > MIN_THRESHOLD).collect();
where he recommends avoiding the first collect(). Can't the optimizer do that for you if you don't do anything else with nopes?A way to mark functions as pure for this purpose would be great! Especially if it's not as fraught as const in C++.
It was also a very, very long time ago, and so today's Rust might be different enough that those reasons don't apply any more.
Haskell can do cool optimizations that make it feel like magic sometimes.
Meanwhile mapping in Rust generally takes an iterator and produces another iterator that will apply the given closure to the current element as it's yielded. So the naive codegen for iter().map().map().map().collect() is exactly what list fusion is trying to produce -- no temporary lists.
TL;DR: making "map" having the monadic `T[U] -> T[V]` signature is really expensive. ¯\_(ツ)_/¯
Ah, yeah I had a vague memory that that was the case, which is why I asked.. I seem to recall it being removed, but I only follow Rust news, don't use it, so I wasn't sure of the reasons or implications.
You just demonstrated what I mean when I comment that const is fraught.
Purity is a different concept.
Purity more easily maps more closely to an abstract concept and is easier to for people reason about.
<Disclaimer: I don't implement compilers so, i can be totally wrong>
The real problem is that you're shoving the variable into a Vec<>, which means that you are doing heap allocation. This means that optimizing it out requires matching the heap allocation to the free, noting that the side effects of these two function calls are only about allocation. There's also issues with respect to inlining, dead-code elimination, and then doing data dependence analysis to prove that the two loops can be fused.
List<CompletableFuture> futures = requestList.map(Client::makeRemoteCall).collect(toList);
List<Response> responses = futures.map(CompletableFuture::get).collect(toList);
Basically, we need to start all of the futures, then wait for all the futures. Eliminating the first call to `.collect` would force these remote calls to happen one at a time.Actually that's part of why ApplicativeDo was created: https://research.fb.com/publications/desugaring-haskells-do-... It's pretty cool.
While that Java code is neat, you don't need to get anywhere near concurrency to exhibit problems with side effects. Here's a program with two iterator chains and two calls to collect:
let foo = [1, 2, 3];
let bar: Vec<_> = foo.iter()
.map(|x| {
print!("{} ", x);
x + 10
})
.collect();
let qux: Vec<_> = bar.iter()
.map(|x| {
print!("{} ", x);
x * 10
})
.collect();
This program prints "1 2 3 11 12 13 ".Here's that same program, but with the first collect removed and the iterator chains collapsed into one:
let foo = [1, 2, 3];
let bar: Vec<_> = foo.iter()
.map(|x| {
print!("{} ", x);
x + 10
})
.map(|x| {
print!("{} ", x);
x * 10
})
.collect();
This program prints "1 11 2 12 3 13 ".When you have two iterator chains with two calls to collect, everything in the first iterator runs to completion, then the second iterator chain runs. When you collapse those two chains into one, you change the order in which things happen.
It's still true that a smarter type system would be able to track side effects, but like all effects systems, these things are viral, and it's not immediately clear whether the annotation burden makes up for itself in optimization potential.
> for x in &xs { // do something with x }
I am curious why the compiler can't rewrite the former to the latter?
The second just increments a pointer which starts at the first elem of the array; the first increments a counter and then uses it as an offset into the array, which is bounds checked (the second never generates bounds checks). Proving that that counter never exceeds the end of the array, so the bounds check can be dropped, is not trivial.
Though this might be information that rustc knows about but not LLVM.
(I may have gotten something wrong, this is my first Rust program! Yay!)
use std::mem;
enum MyEnum {
A { name: String, x: u8 },
B { name: String }
}
fn a_to_b(e: &mut MyEnum) {
// we mutably borrow `e` here. This precludes us from changing it directly
// as in `*e = ...`, because the borrow checker won't allow it. Therefore
// the assignment to `e` must be outside the `if let` clause.
*e = if let MyEnum::A { ref mut name, x: 0 } = *e {
// this takes out our `name` and put in an empty String instead
// (note that empty strings don't allocate).
// Then, construct the new enum variant (which will
// be assigned to `*e`, because it is the result of the `if let` expression).
MyEnum::B { name: mem::replace(name, String::new()) }
// In all other cases, we return immediately, thus skipping the assignment
} else { return }
}
Don't get me wrong, I think Rust is an incredibly impressive language, but this is nuts. For the first time ever, I had a moment of appreciation for the "simple clarity" of C++11 move constructors. If it weren't for the comments and the documentation[1] I wouldn't have had the slightest clue what this code is doing (a hack to fool the borrow checker while allowing for Sufficiently Smart Compilation to a no-op[2] ... I think).This is a good example of the main conceptual aspect of Rust where I feel it could use improvement.[3] A lot of its features marry very high-level concepts (like algebraic data types) to exposed, low-level implementations (like tagged unions). Now, there's nothing wrong with the obvious choice to implement ADTs as tagged unions, but the nature of Rust as a language that exposes low-level control over allocation and addressing, in combination with the strictures of the borrow checker, means enums and other high-level features live in a sort of uncanny valley, falling short of either the high-level expressiveness of ADTs or the low-level flexibility of tagged unions (without expert-level knowledge or using `unsafe`).
Similarly, the functional abstractions almost feel like a step backwards from the venerable C++ <algorithm>, abstraction and composition-wise. Very nitty-gritty, low-level implementation details leak out like a sieve — you can't have your maps or folds without a generous sprinking of `as_slice()`, `unwrap()`, `iter()` / `iter_mut()` / `into_iter()` and `collect()` everywhere, although at least for this case you would be able to figure out the idioms by reading the official docs. But nevertheless it seems like reasonable defaults could be inferred from context (with the possible exception of `collect()` since it allocates), while still allowing explicitness as an option when you want low-level control over the code the compiler generates.
In this case, enums are a language-level construct, not a library, so the borrow-checker really shouldn't be rejecting reasonable patterns. It (legal to alias between the structural and nominal common subsequence of several members of an enum) should be the default behavior and not require an nonobvious, unreadable hack such as the above. At the very least that behavior (invaluable for many low-level tasks) should be easy to opt in to with e.g. an attribute.[4]
[1] https://doc.rust-lang.org/std/mem/fn.replace.html
[2] Or rather since it is a tagged union, in MASMish pseudocode something like
jnz x_nonzero
mov MyEnum.B, e.tag
x_nonzero:
ret
Although it could still be a no-op depending on what else Sufficiently Smart Compiler/LLVM inlines.[3] And I'm not saying I have the solution or even that all-around-better solutions exist.
[4] Something like
#![safe_alias_common_subsequence(structural)]
#![safe_alias_common_subsequence(nominal_and_structural)]Another problem with your proposal is that it's not safe to access the structural and nominal common subsequence of several members of an enum in the same way, because the Rust compiler can and will reorder the fields differently in different variants in order to fill padding.
That said, it sounds like what you're looking for is one of the various inheritance proposals, which have safe access to common enum fields as one of the main guarantees. This would make this pattern much more ergonomic.
What I was proposing was that common subsequences of enum fields be treated by the compiler as synonyms, but accesses to uninitialized data would still be illegal. One way this might be implemented is by, when necessary, creating additional tags behind the scenes, e.g. MyEnum::__AB__ meaning A was initialized but B is the active member, so only accesses to their common subsequence are legal, MyEnum::__BA__ meaning the reverse, and so forth. Or it could be restricted to fields that contain no members outside their common subsequence with nontrivial destructors. Or the compiler could null out any uninitialized references to such.
>That said, it sounds like what you're looking for is one of the various inheritance proposals, which have safe access to common enum fields as one of the main guarantees. This would make this pattern much more ergonomic.
Interesting. Link?
I feel like that's too much hidden behind-the-scenes magic for a systems language, and I suspect the majority of the Rust community would feel the same way, but feel free to file an RFC.
> Interesting. Link?
That's not been my experience at all. In fact I've found <algorithm> to not only be extremely limiting but also being painfully verbose due to the need of begin/end pairs everywhere.
What class of algorithms is faster in Rust than it is in C?
Theoretically, Rust's memory ownership/aliasing rules are stricter and more granular than C's restrict, so some pointer-heavy code could optimize better. Rust is very good at inlining by default. Rust makes it easy to use stack-allocated structures. But C can do the same, it's just a matter of effort.
Both languages are low level enough that you can use them as a portable assembly and endlessly tweak them to one-up the other. If some C code doesn't optimize well you can write a more contorted code that will.
Maybe something that leverages signed overflow, overlong shifts, or type punning.
But if you have enough control over your codebase to mandate the compiler/flags its built with, then you can generally tell the major C compilers to act like Rust in these cases.
That said, the expected win for Rust over C(++) in practice is that you can be more "reckless", because you have a stronger type system protecting you from messing things up. A production-quality C(++) codebase might rightly do more copies, use more reference counting, and use less concurrency just because the risk of doing otherwise isn't worth the potential performance wins.
Organizations have limited resources to commit to optimizing/verifying code. Rust is intended to get you more bang for your buck.
I'm glad to see somebody articulate this observation. SaferCPlusPlus[1] is meant to, in part, bring this benefit to existing C++ code bases. The question is, would a borrow checker for C++ make sense?
[1] shameless plug: https://github.com/duneroadrunner/SaferCPlusPlus
Back when Rust used "internal" iterators, this could be the default. Now, you can encapsulate it by passing a function to handle each line.