C++ still isn't cutting it
da-data.blogspot.com
da-data.blogspot.com
1) Rust's standard idioms (in this case, Results and Iterators) encourage less bug-prone or ambiguous APIs (standard and otherwise)
2) Rust's ecosystem makes it easy to find and integrate solutions not offered by the standard API
3) And then finally, Rust's borrow checker prevented a threading issue
Others are rightly pointing out that only #3 is fundamentally unique to Rust. All of these other things could be done for C++. But, importantly, they haven't.
This raises an interesting dimension to the comparison which is that many of Rust's advantages don't really come down to its unique traits (heh) but to the simple fact that it's "C++ without baggage". It had a fresh start when it came to establishing standard idioms for everyone to use, providing cross-platform tooling for everyone to use, providing centralized package management for everyone to use, etc. If you introduced to C++ a Result type, or a package manager (I assume people have done this already), most C++ developers wouldn't end up using them. Most libraries wouldn't be using the same vocabulary. It would be arduous to even move the standard library over, because it would be a breaking change. Much of the tooling out there would probably never get specialized support.
I don't think these network effects and cultural forces get talked about enough. Sure, this stuff is "just a library", and C++ has libraries. But the culture and the baggage and the stakeholders around a language have a huge impact on what ends up being practical to do with it, independent from what's technically possible.
This is a double-edged sword. I'm a huge fan of Rust's tooling, but when it's so easy to add dependencies, this inevitably leads to deep dependency trees, slower compile times, and less knowledge about what's actually happening inside most rust code bases. Often times when asking about how to solve a problem in Rust, the answer comes in the form of "use crate X" rather than an explanation of how to solve the problem using the language itself.
Do you think this is a bad thing? How frequently would you say it's a better use of a software engineer's time to carefully reimplement a commoditized library that already exists for a speed improvement or to understand it?
I don't really understand arguments that you shouldn't introduce convenience and quality of life features because programmers will lean on them too much. Leaning on them is the point; the implicit thesis is that programmers no longer have to understand what's happening under the hood. It's specialization and division of labor applied to programming.
And if you really care, you can probably just look at the source code or pick up a book to learn how it works?
People make the mistake of treating this issue as black and white: either you're for or against code reuse. But the reality is far more nuanced. Often, a dependency will solve a much more general problem than what you actually need, and thus, avoiding the dependency might result in solving a considerably simpler problem than what the dependency does. In exchange, you use less code, which means less to audit/review and less to compile.
Given my position as the author of a few core crates, I actually often find myself advocating against the use of those very crates when the problem could be solved nearly as simply without beinging in the dependency. (I did not author globwalk, but I did author its 'ignore' and 'walkdir' dependencies.)
Let's say someone makes an XML parser (just trying to pick an example). IMO a bad XML library would read files. Instead it should just, at best, take some abstract interface to read text and outside the library the docs should provide examples of how to provide a file reader, a network reader, an in memory reader, etc...
But, I rarely see that. Instead (this is more the npm experience) such library would include as many possible integrations as possible. Parse XML from files, from the network, from fetch, from web sockets, there would be a command line utility to parse XML to JSON, and some watch option that continually parses any new or updated XML file and spawns an task to do something with it, all integrated into this "XML Library". Parts of it will respond to environment variables and it will have prettifiers integrated with ANSI color codes and you'll be able to choose themes, and it will have a progress bar integrated into the command line tool for large files
And the worst part is the noobs all love it. They down load this 746 dependency XML library and then ask for even more things to be integrated.
Maybe someday there will be a language with a package manager/community/guidelines that mostly somehow rejects bad bloated packages. It seems like a nearly impossible task though.
Note: I don't know rust well but I tried up fix a bug in the docs. The docs are built with rust. The doc builder downloaded a ton of packages which to me at least was not a good sign
Yet another possibility that I think has a place in this nuance conversation is adapting existing code with changes. It's interesting that we don't culturally today have an accepted way of doing copying + adaptation, but over most of the history of computing it has been very common. It has the obvious downsides of not easily getting improvements and bug fixes of later versions of the upstream code, and rightly is passed over in most cases, but in some cases it's still a win.
But this is the kind of reasoning that leads to the situation in npm of thousands of tiny libraries such as "leftpad" that many systems programmers are so derisive of.
yeah for sure, it's a hard problem. but at least a high level listing of similar libraries in the same category with # of recursive deps and total bloat size clearly visible so it was easy to shop around and sort by whatever attribute the developer wants to prioritize.
it would be similar to category-specific filters & comparison tables for features of e.g. SLR cameras, usb3 power supplies on amazon.
Having lists of alternative libraries to solve a particular problem is great. It is valuable on its own. But it doesn't really fix the nuance that I'm focusing on here in this example. And my general claim is that this sort of nuance is a fairly common problem that leads to unnecessary dependencies.
If you made good choices (e.g. maybe you depend on Python's Requests) you had to go out of your way to make problems for yourself, by default the software understands that it should prefer shiny modern cryptography, allow anything that's probably fine, and reject nasty broken stuff. What counts as "shiny modern" or "nasty broken" will vary over time, but it's not likely you, as the application developer, were best placed to decide when you wrote the code, so unless you specifically did so the runtime decides.
In contrast Microsoft shipped not one but several distinct C++ APIs and .NET APIs where you'd need to go back and periodically modify old code (e.g. to specify newer TLS versions) or now it doesn't work because there was no provision for programmers to just let the runtime do the Right Thing.
I do think that Rust making it easy doesn't necessarily mean that you have to use it. It's fine to write applications using only the stdlib, and only link to some known libs like openssl or libpq.
In Rust's defense, it is easy to list all dependencies, and they don't automagically upgrade from one checkout to another.
Last large-scale C++ codebase I worked on took 60+ minutes to compile when you touched the precompiled header.
In fact, I've found it to be common to have binary-only library dependencies in C++ which makes it harder to tell what those downstream dependencies do.
This toy program written using Boost.Beast or something similar would compile much faster. Of course I would have spent all day setting up the library instead of waiting for it to compile...
I've never thought that "having more friction when adding dependencies" is a valid strategy for preventing overuse of third-party code. Unlike other "trust the developer" scenarios, dependency bloat is a minor issue at best and is very easy to detect and diagnose. IMO you should cut down your dependencies when they actually become a problem that's bigger than the one they're solving.
> and less knowledge about what's actually happening inside most rust code bases
I actually think the package ecosystem (along with macros) is crucial for the breadth and accessibility that Rust provides. Rust is a low-level language, but you have high-level libraries at your fingertips if that's the level where you need to be working. This means that instead of writing hot-paths in a low level language and application code in a high-level language (and dealing with the added complexity in terms of FFI, builds, tooling, developer knowledge, etc), it's very possible to just write the entire thing in Rust. Again in this case, I think "forcing devs to learn how everything works" is a weak argument for increasing the friction for adding dependencies. I'm pretty sure that everything on crates.io is required to include source, not just a binary, so it's trivial to dive in and learn how it works if you're so inclined.
I ended up paying for Clion just to get better debugger integration and a quick action for cargo expanding a file in the UI so I can copy paste the macro results into my codebase. I'm hoping that will improve incremental compilation times until I buy a Threadripper workstation.
In the latter case, maybe cargo needs an intermediate macro-expansion cache (instead of just a crate-level build cache)
Whenever you run anything there's millions of lines involved across the OS API's directly or indirectly, so in practice it doesn't matter what's happening behind the scenes. In languages with fast compile times like Go and Java people add dependencies without even caring. And I think that's how it should be.
All our Java backends at work are like 200MB and it doesn't matter and nobody cares. It still compiles in less than ten seconds, and all that code barely affects startup time.
This does result in large dependency trees, but it does also mean that if you need that one specific functionality, you may not need to pull in the full library, but only that smaller section.
I think programmers often under-emphasize the importance of culture, so I want to boost this line.
Lots of "debates" around programming languages are, I think, better understood as differences in beliefs about what a language should prioritize. C++ could add an allocation checker that provides stronger memory guarantees by outlawing certain operations (or requiring they be marked a'la `unsafe`). They don't add it because, I think, they view other additions to the language as having a higher "return" on allowing people to create high quality programs.
It seems to me that a lot of this comes back to first-principles about what makes a good programming language. These ideas are, to some degree, set before a line of code is written. The rust folks thought memory and concurrency safety were very important when they were designing rust - even though they couldn't have formed those opinions by using the kind of language they were making (because nothing quite like it existed yet).
I also think about "infamous" qualities of languages (generics in go, the GIL in python). These ideas exist within wider cultural beliefs about what makes a good programming language. In a sense, it's inevitable that most of the people working on cpython aren't too concerned about the GIL, because they chose to work on a project that has developed around the GIL and believe its limitations are worth the gains from being able to rely on its coordinating function.
The implications of this are, to me, that when advocate for languages to adopt features we'd like to see we should remember the wider cultural context. It seems silly to expect the current golang team to implement generics "the way the community wants" because they've spent years building a fully featured language that functions without generics. Culture runs deep and it's not an intellectual or professional flaw to value improvements differently than our peers.
Cyclone has been around some 2002 according to Wikipedia. I don’t think a similar language (RAII + annotations + a mostly-mandatory validation step in the compiler) has had a well funded and sustained marketing campaign until rust, though. RAII also predates Rust by a couple decades.
Rust is also not very similar to itself from its early days.
I mean that, when they decided it was time to make a new language, the team would have had some understanding of what the language would be like. That understanding couldn't come from using the language itself (it doesn't exist yet). Instead, teams do cultural work to communicate & come to shared understandings about what the language they are making "should be like." That shared understanding survives past when the language is working and influences decisions to improve the language.
It's just worth considering and talking about project culture! It's common for there to be many possible improvements to a language and have the culture of the project (or wider community) be what tips the scales.
This made me think back to something Rich Hickey said in his "Are we there yet?" talk. In short rust has memory management, and so can better function as a library language, similar to how Java did back in the 90's. Rich's quote below.
>> And it is a big problem. I think the lack of garbage collection really impeded C++ in one of its design objectives, which is: it was supposed to be a library language. All the original design stuff and any time you heard Stroustrup talk about it, it is like, "C++ is going to be a library language". But it only ever ended up being a parochial library language. Every shop had a library, but there were not a library, and still are not a lot of libraries, that go between places, because of this problem.
https://github.com/matthiasn/talk-transcripts/blob/ccc4a0172...
I've converted many mono-thread C++ algorithm this way.
Except they have, but some people aren't keeping up with C++ news, only bashing it.
1) can be easily done with bounds checked enabled STL, like modern C++ compilers do in debug builds, or force enable them in release builds.
2) Conan, vcpkg, and better than cargo, because they support binary dependencies.
Cpp API design also seems to suck. Doing a simple curl request is bloody complicated.
Cargo keeps failing my litmus test for usability of compiled languages.
I just look forward that something like Cranelift will fix it.
Without an ability to tweak build flags. Which makes them useless. And since they are not really crossplatform either, you better stick with portage, which is still far behind cargo.
Sword fighting gets tiresome after a while, and not all businesses enjoy shipping source code anyway.
I don't know too much about the details of rust, but: std::optional? Whether people will use it is of course up to them, but my feeling is that they will.
I believe std::expected is the correct answer but it’s unclear to me if it actually made it into the standard yet or not. I thought it had but can only find proposals...
As always, it depends. But traditionally, most languages that use this strategy have both, and they aren’t interchangeable. I can see this happening more when you only have optional, though.
Hacking failure reporting into an existence API will make your code confusing, and will leave you no way to report multiple types of errors.
CMake kinda makes me think there's definite interest in this kind of thing in the C++ community. It's janky and horrible, but it has tackled the module/package/library issue a little bit (not in an elegant sense but its a little more standardized)
Conan, the C/C++ Package Manager [1]
That's how I've always viewed it. Or rather a better replacement for C++.
I think there is no language level feature in Rust to prevent TTCTTOU. TTCTTOU can happen in both c++ and rust.
for example:
let f = File::open("username.txt");
/*what if the file is deleted here?*/
let mut f = match f {
Ok(file) => /*what if the file is deleted here?*/ file,
Err(e) => return Err(e),
};
/*what if the file is deleted here*/
I mean even if you write it like: let f = File::open("username.txt")?;it looks atomic, it is not.
the c++ example, on the other hand, can do a filtering too, using:
http://www.cplusplus.com/reference/fstream/ifstream/is_open/
I mean Rust forces you to do error handling, which is nice. But that's not enough for preventing TTCTTOU. I don't think the first example is a strong argument.
If you want atomic open-and-read-the-whole-file, you can reach for something like std::fs::read_to_string (https://doc.rust-lang.org/stable/std/fs/fn.read_to_string.ht...).
Nothing. In all major operating systems I'm aware of, opening a file creates a handle to the file contents. This handle is distinct from the handle implied by the existence of the path in the filesystem. Some filesystems allow you to create another handle to the same file by a different name (called a hardlink). Removing a file only deletes the handle implied by the path; if another handle to it exists (via a hardlink or the file currently being open in a process) the data itself is unaltered.
In general, TOCTTOU attacks are caused by introspection being done on paths instead of file handles. Once you have handles, and you are querying exclusively using handles, your TOCTTOU issues largely go away. There is still a related issue that filesystem atomicity guarantees (or rather, a general lack thereof) is still a major headache, but it's not the same issue as TOCTTOU attacks.
The point is, if I go by names, there is no guarantee of identity. File-descrpitors ensure that.
Same could be done with c++, but the libraries are not designed that way.
Still, I don't find it terribly convincing.
More broadly, this is kind of a bad example. Handling files is surprisingly difficult and that's really where the dragons lie here. IME the dragons exist for both Rust and C++, because they're platform problems rather than stdlib/language problems.
NB: the correct thing to do here for POSIX platforms is to call fstat, which is true regardless of systems language.
I absolutely understand some of the niceness about Rust preventing programmers from shooting themselves in the foot(it is still possible nonetheless), but sometimes it feels people just want to complain about C++ for the sake of complaining.
All the common FS (in linux .. ext* xfs) do not have that these.
As you intuited, a problem that arises in environments that lack transactions/locks/etc
You know that files are nameless on Unix, right?
I was referring to the transaction semantics
> /what if the file is deleted here?/
you don't need transaction semantics here.
On Unix file is basically inode with counter. If it doesn't have record (aka name) in directory but non-zero counter, it would exist.
So on open(2) counter is bumped up and is equal to 2 (or higher if there are hardlinks), one for the directory record and one for program running it.
If you remove record from directory (aka `rm filename`), counter would decrease by 1, but it is still non-zero, program still has it open, could read its content, etc. Only after close(2) counter would decrease to 0 and file would be gone.
Just tested it on Ubuntu 20.4 with simple C program
[0] https://github.com/rust-lang/rust/blob/master/library/std/sr...
[1] https://doc.rust-lang.org/std/os/windows/fs/trait.OpenOption...
You know that files are nameless on Unix, right?
For example, undefined behavior by default, value categories (rvalues, lvalues, etc), exceptions in destructors, SFINAE, and so much more.
Even if C++ were to end up making its syntax closer and closer to Python, which it seems to be doing, beneath the sugar is still a haggard mound of a ridiculously complex and unsafe language. The only way to change that is to start fresh, as Rust did. Rust is complex, indeed, but it removes a whole lot of the safety concerns and legacy crud.
Presumably, 20 years from now something else will be created which improves upon Rust in such profound ways and we can all talk about how complex Rust is and how its idioms aren't as <something> as they could be. Maybe we'll be using machine learning to write and compile code for us and Rust will be shat on for being so complex and manual. But, for now, even with its complexity, it should be a breath of fresh air for anyone working on non-trivial, multi-threaded C++ applications.
There will be those who don't dig it, just as there were those who have stayed with C all this time. But how else would we keep the vulnerability finders busy? Everyone needs to eat.
The same strictness in C++ can arguably be availed by using static linters. This has 3 problems:
1. Static linter errors are inherently optional to fix. One has to have some CI rules to block check-ins if linter flags something
2. Your project might follow best practices but not your dependencies.
3. Static linters are not foolproof (last I checked back 4-5 years)
> The strictness is built into the compiler and the entire ecosystem is subjected to same high standards.
> 2. Your project might follow best practices but not your dependencies.
The crates ecosystem is the weakest part of Rust and it is not always subjected to the 'same' high standards, despite the 'strictness' of the compiler.
A simple command line app or library (even worse) can depend on hundreds of other third-party libraries and cargo reminds others of the same ills of npm to some extent as developers throw in unnecessary libraries in your project. Point (2) can be still said about Rust given that the developer can still abuse unsafe{} or still have dangerous .unwraps() written by the developer. Just look at the immense scrutiny that actix-web was under because of this.
I would take openssl over most (though not all, for sure) rust cryptographic code. OpenSSL's bugs have not been RCEs for the most part, and many of them could happily exist in rust. The cult like perspective that language choice can make software automatically safe in the absence of things like domain expertise or long term production usage increases risk even where the language itself theoretically reduces them. ( See also: https://en.wikipedia.org/wiki/Risk_compensation )
Especially with the deeply nested dependencies the rust ecosystem seems to me like a ticking time bomb for security disasters. The same bad practices exist in NPM and have already produced a number of high profile incidents, as it repeats in rust there really won't be any excuse but to admit that it resulted from an excess of hubris.
There is stuff in ISO C++ standard library that requires third party libraries in Rust, yet another thing to certify in security sensitive domains.
Being politically incorrect to web developers?
I’d agree with that? I think the power and danger of C++ is that it assumes the programmer knows best. Rust assumes the programmer is wrong if it looks dangerous. So.. yes?
First, trivial: what's exactly mediocre in the Rust example? I don't think it's mediocre at all.
Now, the second, which is significant: what constitutes a "mediocre" C++ program? The author is the engineer who built the Meson build system - his code is representing the output of an experienced engineer.
This brings to the usual C/++ criticism: it makes it easy to make certain class of mistakes even for experienced programmers.
A better way to phrase it would be "code produced by a middle-level engineer". It just so happens that in Rust a regular, adequate, middle-level engineer can easily write the best possible code because the language guides him to it, but in C++, a regular engineer still wouldn't know about some gotchas that would make his code inferior.
It's unknown unknowns: Rust shows most of them to you, while C++ stays silent.
Unfortunately not every programmer is or knows their limitations.
And never make a mistake.
Honestly, it's probably just because I'm not good enough, but I prefer to use my focus on improving the functionality of my code instead of ensuring any change I make is correct.
I've successfully maintained some 100k lines C++ code in the past. I would rather avoid that in the future.
Additionally, if the project is owned or influenced by a corporation, the mythical programmers have little influence:
Some authoritarian ex-mediocre-programmer-turned-middle-manager calls the shots and is more anxious to "grow the team" than to ship robust software.
Are you seriously saying that this isn't also true for C, (and indeed for all programming languages)?
A challenge - try writing a program in C to read a text file containing lines of unspecified length, sort them, and output them in sorted order. Compare the C++ and C code and tell me which you think is the cleaner and safer.
To meet your challenge, think about the problem. How about this?
file size + 1 => s
allocate char buffer b[s]
read file to b (all data must be in memory, or merge sort. if on linux and allocation worked, read may fail! oom will get us, or we bail on read failure)
add final newline if missing
count newlines => n
allocate line pointers p[n] (if line pointers will not fit, we bail -- we could use offsets and fancy virtualization, if this is not a Z80)
fill pointer array p
sort => sorted[n] (if sorting fails, bail)
output sorted lines
Same for C and C++. UTF-8 is ok.Do you want more data than will fit into memory? Even then, C++ doesn't bring much to the table. Separating the sort algorithm? Yes C++ BEGINs to pull ahead by a bit.
safe? yes. easy to write? yes. no fancy bits? yes. Transcribe to C C++ as you will. Java? won't be as pretty. Rust, Go? don't know. Javascript. Not even close.
Since this is a "beginning programmer" problem, give a design that works for another language that makes sense to a beginning programmer.
I would argue that this is setting up for catastrophic failure. If not when writing, sometime later when maintenance happens.
The implementation there also tells you that the `Item` type of the iterator is a `Result<DirEntry, WalkError>`, presumably because any of the file system reads can fail. The `WalkError` type is from the globwalk library.
Going back to `filter_map`[2] that function takes a function or closure that maps from `Item` to the `Option` enum, and then just discards anything that is the `None` variant of the enum, filtering them out. `Result::ok` is a function on `Result` that converts the `Result` to an `Option`, converting `Err` to a `None`. So this is discarding the errors without reporting them.
As someone who is familiar with Rust, this wasn't particularly hard to follow, but that comes down to familiarity with these adaptor functions and how they can be plugged into each other.
[0] https://docs.rs/globwalk/0.8.0/globwalk/struct.GlobWalker.ht... [1] https://doc.rust-lang.org/stable/std/iter/trait.Iterator.htm... [2] https://doc.rust-lang.org/stable/std/iter/trait.Iterator.htm...
It also includes a test case which the OP's implementation will get wrong because the order of the combining diacritics is different in both words. If you don't normalize them you'll get two different words.
[0] https://gist.github.com/Measter/e2e287ee21311d34ea8eb8cd9d57...
This is IMO a good example of poor choices that an "npm-like ecosystem" can encourage. Obviously there are deep and fundamental trade offs here, but other than watching crates compile, you're almost completely removed from just how much code you're bringing in relative to what is really necessary.
To the GP, the place it shows up in the docs is mostly in its example: if let Ok(img) = img <-- this is a sign they're unpacking a Result type.
First of all, the proliferation of header-only libraries that people vendor into their Cpp projects has the same attack surface. Additionally, there is less tooling to help you track vulnerabilities in Cpp dependencies and upgrade when fixes come out. If you do rely on external cpp deps, then you must have a code review process before those deps are whitelisted for use.
Now, you can extend this code review process to rust crates.
Run a security review of a crate after which that crate is pushed to an internal registry. Every external crate upgrade can go through the same security review process until which point developers will build with the old version present in the internal registry.
In the context of a large enough organisation, an increasing number of crates will become internal, thus solving the trust/responsibility issues.
by no means, do i suggest that crates is unhackable. That's why I want to raise awareness of already existing infrastructure and procedures to vet code before including in your systems.
Safety
Before we get to ownership and some thread-safety, how do you explain the fact that integer addition and number conversion are safe operations in Rust, which throw/panic when they fail instead of the default Cpp behaviour, which _silently_ corrupts data?
I just want to use tools that help me get stuff done. C++ has had a 30-year head start, why can they not see the overall value-add of a default build and test tool and a format for package declaration and management.
All integer types have "checked_" arithmetic operations, which return None in case of overflow. https://doc.rust-lang.org/std/?search=checked_
C++ does not have a standard checked arithmetic operation, so I'll concede with you there. It really should, although most people just use the fairly widespread compiler builtins that behave similarly. (That being said, having it not be the default means that people who really need it won't use it, which is the same situation that Rust is in.)
I think the author's problem is that C++ even allows you to do it the wrong way in the first place, but I'm sure that you could design an unsafe FS API with Rust as well (though it might be a little harder).
In general, i think a lot of valid criticism of cpp is dismissed with ”But you could do it correctly with X”
Perhapse, but if everything is a special case, what does that say about cpp?
Thing is, all of the solutions to the issue here would be a real PITA if done as one function as both examples are. That is why example code often skips checks to allow the point of the article to be highlighted and not hide it. This is more true with filesystems related code as filesystems are not reliable. One example of how many corners filesytems have is the unit tests for sqlite, they are amazing.
C++ still isn't cutting it for a Rust fanboy
However, if you really want high performance, you can go ahead and write it in C with a custom matching code and a custom data structure for keeping your word count. This will most likely result in hundreds of lines of code, and enough performance to max out a PCIe SSD. However that would be an different exercise.
Second, it's not clear to me at all why this problem asks for regexps. Both examples used a regexp to check a file extension, and then to find word boundaries. Both are trivial† to code directly.
Am I misunderstanding the problem here? It seems like there's hardly any string manipulation in it at all.
† admittedly, I didn't bother being careful about word boundaries
I know some C++ folks were annoyed it’s using regex due to criticisms of the standard regex libraries that can’t be fixed, or something, so it’s also not like folks agree that the original solution was optimal. I don’t think it was trying to be, so seems fine, but it is what it is.
Using a different technique, for example by not using regexp is a bit like cheating in that context. You are not comparing two languages, you are comparing two different solutions to the problem.
That's what I meant with my previous post. Either you do it as specified in the article, with regex and maps, which require pulling libraries and working in a way that is much less convenient than with the C++ and Rust example for no good reason. Or you reimplement it using different techniques and this is not the point of the article.
Or to put it simply, the example in the article is not good for C.
So, no, I don't think you're right about this.
More to the point, though: I'm talking about regex libraries because the parent comment is. My point: the Rust example uses 3p dependencies, so what C does "out of the box" is already out the window.
I'm fed up with C dependencies, which like Makefiles, always seem to be very easy in principle, and then kill by thousand cuts (like a regex library that defines a symbol that happens to conflict with a POSIX regex function, which I didn't even use, but it corrupted memory of a completely different dependency elsewhere).
This particular program needs no third party dependencies at all; in fact, I bet it'd get longer if I added them (like a glib hash table, or pcre).
https://gist.github.com/tqbf/4de61a3e34d2e4664044666c107abe7...
(I don't vouch for this code; I wrote it off the top of my head, and, in keeping with the exercise, I wrote it in pico).
It's not 10x longer. I'm honestly not sure why either Rust or C++ bothered with regexps for this problem.
Anyways, you get the gist of what this looks like in C now. Obviously, don't write things like this in C.
With that fixed, this dumb program does my whole (very large) homedir in about 10 seconds, for some definition of "does" that may or may not include counting every word in every txt file. For the curious:
0. get / 566006
1. the / 158168
2. pkg / 119419
3. syscall / 105828
4. const / 98259
5. and / 86670
6. ideal-int / 43078
7. that / 31609
8. for / 31129
9. this / 24980
10. text / 22907
:P while (fgets(buf, 1024, fp)) {
char *c, word[1024], *w = word;
for (c = buf; *c; c++) {
*c = tolower(*c);
if (*c >= 'a' && *c <= 'z')
*w++ = *c;
else {
*w = '\0';
count(word, (size_t)(w - word));
w = word;
}
}
*w = '\0';
count(word, (size_t)(w - word));
}
and you have that `count` function skip words with len < 2 then I can match results with the other 3 (Nim, Rust, C++) versions. Also, your C version runs in just 7 ms on Tale Of Two Cities, about half the time of the PGO optimized Nim.I thought about also keeping the top 10 as I go instead of copying the whole table. But I'm guessing that virtually all the time this program spends is in I/O.
I just did a profile and saw about 15% in strcmp in the hot cache run, but sure if it's not in RAM then IO is likely the bottleneck.
This C code is not only not protecting you against that, but also it seems to have little to none error checking. What happens if fopen fails? It just skips to the next file and happily ignores all the entries in the file. What if fget fails? Again, it stops processing the file and goes to the next.
I won't even get into what happens if calloc fails and returns null. Or if ++ wraps around and you get a negative value (let's hope we are using 64 bits)
This program will likely work, but if it doesn't you will just get an invalid number and never find out.
Don't get me wrong, I understand this is "not serious" code, written in a Sunday and for fun. And I would be ok with it if the article wasn't about language safety.
(I think you estimated badly, for what it's worth.)
:P
You can golf out about 10 lines of this by getting rid of the table free. :)
I could have done a number of things to keep the line count down–in the spirit of the challenge I tried to keep it clean, pedantic, idiomatic POSIX C because if I didn't I'm sure someone would have jumped on it for it not being that :P So it's written in nano in the style of how I might write a homework assignment rather than something I specifically golfed. In retrospect the fixed-depth buckets probably added more complexity than they saved, and using a linked list or open addressing would have probably been much easier since it'd clean up some of the resizing and traversal code. Plus, the load factor was in general fairly poor, the final table (when I ran it on my Downlods folder) had just over 8 million buckets of depth 5 allocated and only about 600k got used at all, and of that 570k were filled just once.
Oh, and here's the results on my Downloads folder:
514464 Developer
425236 Xcode
386616 saagarjha
386075 Users
361324 Library
337374 build
333056 Wno
306758 WebKit
302892 DerivedData
296179 eugbibmfmfphgbhczsxiimkhynol
I have a couple WebKit build logs that dominated the results… import os, strutils, tables, heapqueue
iterator topByVal*[K, V](c: Table[K, V], n = 10, min = V.low): (K, V) =
var q = initHeapQueue[(V, K)]()
for key, val in c:
if val >= min:
let elem = (val, key)
if q.len < n: q.push(elem)
elif elem > q[0]: discard q.replace(elem)
var y: (K, V) # yielded tuple
while q.len > 0: # q now has top n entries
let r = q.pop
y[0] = r[1]
y[1] = r[0]
yield y # yield in ASCENDING order
template maybeCount(count, word: untyped) =
if word.len > 1: count.mgetOrPut(word, 0).inc
proc main() =
var count = initTable[string, uint]()
for path in ".".walkDirRec:
if not path.endsWith(".txt"):
continue
for line in path.lines:
var word = ""
for c in line.toLower:
if c in {'a'..'z'}:
word.add c
else:
count.maybeCount word
if word.len > 0: word.setLen 0
count.maybeCount word # line ends in a word
for word, cnt in count.topByVal:
echo cnt, " ", word
main()
Doing a gcc-10 profile-guided-optimization build on this, I get the min time of 5 trials running on Charles Dickens' Tale Of Two Cities { opening paragraph is very optimization claims relevant! :-) } available here: http://www.textfiles.com/etext/AUTHORS/DICKENS/dickens-tale-... 7.3 ms EDIT- tptacek's C version (w/bug-fix & new tokenization)
15.8 ms Above Nim 1.4 PGO --gc:arc
24.8 ms Above Nim 1.4 --gc:arc
27.3 ms Original C++ PGO
30.2 ms Original C++ -O3 only
42.6 ms dga's Rust with rustc-1.47.0 build --release
This is on a i7-6700k. So, Skylake core and probably more importantly 8 MiB L3. So, also FWIW, I could not reproduce the Rust/C++ timing ratios with the above mentioned input file. I would also point out that not only is globwalk unnecessary as @burntsushi points out, neither are regexes nor even sorting. A heap is a better way to do "top N", but you may need to reverse the answer if you care. I mean, maybe the point of the original C++ is to also bench/exhibit use of regex engines and part of the Rust point is utf8 or some such, but to me that seems more like a distraction.Anyway, maybe my rust build environment is screwy. Hence, my inclusion of an exact input .txt file for others to try out if they care. I had to add some extern crates and [dependencies] of anyhow, lazysort, regex, globwalk. I've never done a PGO build with rust, either. So, that may help.
- I7-7920HQ, MacOS
Rust version: 90ms (rustc 1.47.0 (18bf6b4f0 2020-10-07))
C++ version: 340ms (Apple clang version 12.0.0 (clang-1200.0.32.2))
- Xeon(R) Gold 6130 (skylake, Ubuntu 20.04)
Rust: 76ms (rustc-1.47.0)
C++: 60ms (gcc version 9.3.0 (Ubuntu 9.3.0-10ubuntu2)), -O3, no PGO
Seems quite CPU and compiler-dependent. Odd that the results on MacOS were so horrible for C++.And thanks - I've updated the post to include this more-reproducible benchmark, and included both macOS and linux results.
One note: I modified the single threaded one to use walkdir, but that shouldn't affect time in a major way. The macos timings were about the same.
And yes, I agree about the top N part. I deliberately tried to remain "algorithm-compatible" with the C++ example; there are lots of tricks to use to speed this up more. Getting rid of the line-at-a-time processing would be a good start, for example - it results in an unnecessary double-scan of the input.
I mostly thought Nim deserved to be seen and then it happened to also be faster..perhaps giving folks a slight Bayesian update on presumptions of performance. :-)
[0] https://gist.github.com/Measter/e2e287ee21311d34ea8eb8cd9d57...
Really, though, all the code for all 6 versions (two Nim, one C, two Rust, one C++) as well as the input file is available to all. So, you should/could double check yourself. As dga mentions in his updated blog there is a lot of compiler/CPU sensitivity.
For example: the author writes "What happens if the file has been replaced by, e.g., a pipe or other not regular_file?" But it seems to me that the same problem -- replacing the file which is regular during the directory walk with the pipe exactly before the said pipe is opened -- can exist in Rust version too.
It doesn't matter how the syntax looks like, I suspect the equivalent system calls are the same, Rust will also walk the directory at T1 obtaining metainformation, at T2 open the file, even if it is at that moment actually a pipe, and then at T3 check if the metainformation obtained at T1 is a regular file ("file" in the source?), which could still be that? That says to me that the both version can suffer from the same bug.
Luckily it's something that's not going to happen too often in practice, knowing how typical programs don't use the same file name for a pipe and a file in rapid succession.
You could do the same in C++, but in Rust it's the default. So...
Iterating through a directory generally yields a Path or a DirEntry, both of which have a metadata method readily available, and neither will nudge you to open the file first and then calling metadata on the open file.
Metadata request must happen at the different time point than the open of the file and the API as far as I see works on the path, not on the handle:
https://doc.rust-lang.org/std/fs/fn.metadata.html
which says to me that the Rust version can manifest the bug of the same kind -- there are different time points in which different independent information is checked, and that all the calls correspond to the same opened file handle just can't be assured?
As far as I understand only checking the properties of the opened file via the same file handle is safe for the described kind of bugs, nothing else.
std::fs::File::metadata operates on the open file. If the file is replaced with something else, your open file handle will still refer to the (orphaned) file, and the metadata still reflects its properties.
Call me old fashioned, but only C is more sure to be explicit there as it wouldn't allow two different functions having the same name. If in C++ the corresponding names are also overloaded, both C and C++ "aren't cutting" if the obviousness is not there.
Filesystem handling is not trivial, and some knowledge is required beyond the language itself. I do think that File::metadata and Path::metadata make a nice API together (better than lstat, stat, and fstat).
As for C not allowing two functions with the same name, sure, but then it doesn't have namespaces. In a nicely designed C library for file handling, these functions might well be called std_path_metadata and std_file_metadata, which boils down to the same thing.
No, Rust won't help you with that. On the contrary, Rust has excellent type inference, and in this case the name belongs to a type, which itself is inferred. As a casual reader, you'd have to track the documentation on the preceding calls.
(This doesn’t invalidate your point, of course, just saying you can do this.)
it would have been just as possible to write the code incorrectly by using the path version to get metadata.
Why are these C++ bashing threads practically always about Rust? I have nothing against the language but I do feel the Rust devs have some sort of inferiority complex and a compulsive need to convince everyone about the superiority of their tool.
Guys, these endless pissing contests are not really productive.
That means Rust is a technologically natural counterpoint to C++ because it was literally designed to be a C++ replacement, and coming from a high-profile company means its entire target audience found out about it very quickly, so any discussion relevant to Rust can find plenty of participants. The rest is just due to all the usual reasons programmers bicker about things, and we've seen this language bickering again and again before, it's just a different language's turn in the spotlight.
I'd love to see these discussions include a more coherent analysis of when a language is appropriate, instead of assuming everyone should always make the same choice. If I'm starting greenfield, when should I pick C++ over Rust? If I'm looking at C++ alternatives, how would I choose between Rust and Golang? Is D ever a better choice? When? Why do we act like you can only pick one, when IPC and FFI and microservices mean you can split the difference if a language is only best for a portion of your project?
> Is D ever a better choice? When?
Probably not, very small community, few resources, etc. …
> Why do we act like you can only pick one, when IPC and FFI and microservices mean you can split the difference if a language is only best for a portion of your project?
FFI is a huge pain and pretty slow in golang if not already done for you. IPC & microservices do have a lot of overhead, so you probably don't want use it for e.g. your text parser. It also makes it harder to use established knowledge in another language.
> If I'm starting greenfield, when should I pick C++ over Rust?
- Do you have experienced C++ developers/are you one?
- Is your environment geared for C/C++ code, e.g. kernel modules or RTOSes, maybe game development?
- Do you have dependencies that strongly expect you to use a certain language, e.g. QT or (less so) GTK? Might be a pain to use with Rust
> If I'm looking at C++ alternatives, how would I choose between Rust and Golang?
- Rust is more speed/correctness focused while Go feels a bit like a faster scripting language. So something that might have started in python is probably a better match for go (a webserver, something script-like)
- Go really isn't meant for constrained environments (although there do exist solutions) like small linux IOT devices or even bare-metal stuff,
I feel like Go & Rust are pretty good complements and while you can do a lot of things with both, they are clearly very different. For some examples, with what I'd recommend, not necessarily what I'd use, since I do have a huge fondness for Rust, not so much for go
- Normal web backend: Go
- High performance web server: Rust
- DB: Rust
- App server: Go
- Crypto (the non-money kind) stuff: Rust
- WASM (it's pretty embedded-like): Rust
Have you used both? Not seriously, I'm going to say, or else you wouldn't say something like that.
As one example - one reason you'd be looking for a C++ alternative is to interact with existing scientific computing codebase written in C. Good chance you won't care about all the stuff that gives Rust the steep learning curve, and you don't want to mix weird symbols all over your code. I used Rust first, then moved to D. There's no way someone used to programming in C is going to feel more comfortable with Rust than with D. (They might still choose Rust, but there's just no way they'd do it because of comfort.)
A bunch of people have been doing it in a bunch of languages. It’s kind of the natural way to explore the topic.
Did author even know that one can check stream state or have it throw exception. It is a very poor attempt at pissing contest.