Porting C to Rust
wiki.alopex.li
wiki.alopex.li
Disagree. In my experience all was wonderful once I was past the point where I decided to just not worry. Much the opposite, so much pointless and confusing OOP boilerplate (as well as blood, sweat, and all that) goes into "securing" code against misuse.
Invest some time into data structures with obvious meaning, or a procedural API that is easily understood. "Nobody" will "ever" misuse, and misuse will be easy to detect. Much better tradeoff IMO.
And lastly, of course, abstraction has nothing to do with "security systems". It's against popular opinion, but you can in fact have good abstraction with nothing but plain functions and void pointers. I like to view abstraction as mostly a conceptual thing that doesn't even happen in code.
I've spent a lot of time digging around in other people's C code over the years, and with very few exceptions, the kind of low-level byte munging in the original post tends to be absolutely riddled with bugs. There are off-by-one errors, buffer overflows, integer overflows, and memory leaks. The error-handling code paths tend to be broken, too.
And if you run a fuzzer like libfuzz or American Fuzzy Lop, it will almost always find even more vulnerabilities. (A handful of C programs do better, including djb's tools, SQLite, dovecot, and significant portions of Apache. Apache seems to mostly succeed because it has good abstractions for strings and buffers.)
Back in 2002, I published a ~7,000 line XML-RPC library in C, with extensive unit tests. I ran it through multiple code quality tools, including Electric Fence. I spent a long time carefully examining each function for correctness. I chose Expat as my XML parser, because it was one of the best-written at the time.
Overall, I thought I was an unusually paranoid and careful C programmer. But here's the list of CVEs against my library and—most importantly—its dependencies: https://people.canonical.com/~ubuntu-security/cve/pkg/xmlrpc...
If anybody here writes C code, and if your code runs on potentially hostile data, I strongly recommend experimenting with a fuzzer. It can be a brutally humbling experience, even if you think you're an exceptionally careful programmer. And if you rely on anybody else's code, like I did, you inherit all their bugs.
TeX is an exceptionally stable program, the design has been frozen since version 3.0 in 1990 reflecting Knuth's desire to have a format suitable for reproducing the typesetting precisely even for archived texts that were developed on hardware that is no longer running. Since 1990 only minor changes have been made to TeX and this is reflected in the current version number, 3.14159265, which is converging to pi.
A comment by Knuth in the error log in March of 1990 was "We’re now up to Version 3.0; I sincerely hope all bugs have been found." At that point there had been 908 errors logged. The last entry is currently an entry made in 2014 for error 947.
I've always been fascinated by this list (see [1]).
[1] http://texdoc.net/texmf-dist/doc/generic/knuth/errata/errorl...
As to C being more bug prone, for my own code I believe the biggest improvement to correctness would come from fat pointers with array size information, enabling bounds checking. That would help to spot errors in little-used code. On the other hand, fat pointers and other sophisticated methods can get in the way of a clean normalized data organization. A lot more flexibility with respect to data and code organization is possible in plain C, which is a great enabler of software in the first place. In the end, if the code is simpler and shorter and has more deterministic performance, but each line of code is more bug prone, that can still be a worthwhile tradeoff.
By the way, I don't view myself as a very careful C programmer and I don't have the patience to check my code against all kinds of UB and such. I've never done fuzzing (I might try it). But for example valgrind has revealed only few problems once my code was working, in the past. The main problem in my eyes is when coding C in an object-oriented style with endless pointers and lifetime issues. What I do is I try to avoid tree structures (including XML) since I feel it's a lot of peeking and poking and traversing with little reward. Instead, I typically use global data, flat arrays, and integer indices, and as a result the code usually becomes very straightforward.
By the way, for a project the size of an XML parser I wouldn't feel shame for such a tiny list of CVEs :-).
Yes. In fact, I suspect that's the single biggest thing most C programmers could do to avoid CVEs. Rust does something similar with bounds-checked "slice types" like &[T], and it makes a huge difference.
For example, I wrote a VobSub subtitle decoder in Rust and I fuzzed it with 'cargo fuzz' and AFL. VobSub is a hairy binary format, and it's easy to make mistakes. AFL found 5 bugs:
- 3 were incorrect uses of Rust's &[T] type, all of which were detected by the bounds checks.
- 2 were integer overflows, both caught by the fact that Rust panics on overflow in debug builds. Neither of these looked exploitable—one was harmless, and the other one would have been caught the first time it tried to index a &[T].
So Rust's borrow checker is useful. But the single biggest win turned out to be having bounds-checked array slices, as you suggested.
> By the way, for a project the size of an XML parser I wouldn't feel shame for such a tiny list of CVEs :-).
I didn't write the XML parser. That was written by James Clark, who was something of a legend in the SGML and early XML world. So even if you pick your C dependencies carefully, it's still risky. (And my code has been fuzzed much less than Expat.)
But to paraphrase various recent medical claims, there is "no safe level" of CVEs.
> I've never done fuzzing (I might try it).
I recommend it to anybody who works with pointers or who parses untrusted data. It's fun to watch AFL generate a billion(!) test cases and slowly ferret out "impossible" conditions. But it may also make you completely paranoid about bugs.
As you should be. It's not paranoia if they really are out to get you. The the bugs are, they really are...
I think zig is quite a good candidate to fill that niche. It is very compatible with C, but it offers meaningful improvements while still being very picky about what goes into the language.
[1]: https://old.reddit.com/r/rust/comments/9mioiv/porting_c_to_r...
Early ANSI, like Fortran77, only required 6 characters of a symbol to be significant, with compilers not going much farther beyond that. At that point, it sort of becomes a "when in Rome" thing.
As to this particular function name, you're in MP3 land, so MDCT means Modified Discrete Cosine Transform. The "i" is probably "inverse." "L3" is probably "Layer 3" since we're talking about MPEG Layer 3. I had to look up "gr" since it seems specific to the encoding, but it seems to refer to "granule" which is a basic unit of the MP3 data stream (that's also clear from the context, since many functions operate on a gr_info structure). A comment on the gr_info structure might have helped there, but there's nothing wrong with consistently using a shorthand that is used in the underlying spec. In context, it's not actually a bad function name.
So we should weigh the negative impact to the familiar-reader of scanning a 64-character identifier distinguished only by a suffix like "_gr" (or "_greater" if that's what is meant here) against the negative impact of the difficult-to-decipher-abbreviations to the unfamiliar-reader. IMO the net win is to optimize for the unfamiliar-reader in this case. The frequency of unfamiliar-readers is much lower than familiar-readers, the positive impact is much more significant.
> As to this particular function name, you're in MP3 land, so MDCT means Modified Discrete Cosine Transform. The "i" is probably "inverse." "L3" is probably "Layer 3" since we're talking about MPEG Layer 3. In context, it's not actually a bad function name.
Agreed, this codec implementation will likely use some abbreviations for sane reasons. But even if you know what mdct/l3 mean in this context (I did), can you say what this function does (or should do) by looking at the name? What about it is distinct from L3_imdct12/L3_imdct36/etc?
With the exception of the "i" (possibly inverse?), the rest of the parts seem like things you would very quickly become familiar with after a short period of looking at the code and reading up on the subject matter (both of which would be required if you wanted to contribute or port).
I guess having names that are immediately understandable to any fresh reader with zero domain knowledge is a noble goal, but it's a high bar.
It's not a perfect function name (you've pointed out some other, valid, problems) but it's far from being objectively bad.
I'd argue that it is often detrimental to the readability for those with domain knowledge, and if the code (as often is) requires the domain knowledge it can be a disservice.
Consistency is key, if you are consistent these shortened function names can be disturbingly pleasant.
Long names are sometimes a consequence of the programmer not taking the necessary time to think about it (also, consistency is still paramount).
YMMV, but I personally find names like this helpful. It ties the code to the underlying problem domain. E.g. I'm not exactly sure what imdct12/imdct36 are, other than the fact that they operate on a block of data smaller than a granule, but I bet if I opened up the MP3 spec the "12" and the "36" would be more helpful in finding the relevant section than other names that could've been used.
EDIT: In fact, the MP3 format defines two MDCT block sizes: 12 points and 36 points.
Having a 64 character identifier is the bigger problem here.
Modified discrete cosine transform means a DCT type 4 in which adjacent blocks are overlapped. See https://en.wikipedia.org/wiki/Modified_discrete_cosine_trans...
type GenParser tok st = Parsec [tok] stSourceThat name is better than 99.99% of all function names I deal with, actually.
Protip: When viewing something like this in GitHub, press "y" and the URL bar will change to be a proper permalink to the current version of the code. This permalink can then be shared. Alternatively, with the line highlighted, press the … button in the gutter and it will offer a "Copy Permalink" option (which gives you the same permalink you get by pressing "y").
Why not minimp3-rs -- thus people familiar with the original C library will be sure where it is coming form...
That said, yes, it's considered bad form to name your project with a -rs suffix. GitHub is okay.
This is all I could find. It isn't terribly compelling.[1]
[1] https://stackoverflow.com/questions/6390331/why-use-array-si...
That kind of control over memory is what C is very good at, and it is, I think, the main reason why C tops almost every benchmark.
Note that modern C++ offers everything C has to offer plus even better, more advanced tools. As a result, it should be even faster. However, "proper" C++ is a bit more hands off when it comes to memory management, with things like generic containers, smart pointers, ... It tends to result in a slightly worse performance.
As for the "foo[1]" as it is described in the article, it is just syntax. Some people may find it more readable than &foo, it doesn't matter. Kind of like &table[1] vs table+1. Use the form you prefer, or even 1[table] if you are making an IOCCC entry.
"but a shitty way to engineer software as a whole. Achieving abstraction is super hard when you can just reach down into some bytes and noodle around with them instead.": Data oriented design is usually better both in ease of use and performance than any random class abstractions that exist; also, see Linux.
"Bloody hell you can’t tell whether a pointer points to a single object or an array just by looking at it": You can, usually. There are plural words in languages usually used do denote this, items, also the size thing he mentioned. If it doesn't have it then its usually just bad practice or poor code quality.
"heckin’ ternary operators, just make your if statements not suck.": Ternary operators are great, and usually quite concise. Not sure what they are specifically referring to here :/
"The pre and post increment operators are just the worst damn thing in the world.": Again, this knowledge comes to experience, and actually makes things more concise.
C is not designed to be a ""beginner"" friendly language, but its essence is simple -- and I would recommend it for any beginner as it really drives home the majority of actual programming principles, and makes you think about what you are doing on a deeper level rather than coating things in a magical dust layer of classes with vtables and garbage collection.
"it’s called Progress.": However, with modern programming languages it's one step forward with two steps back most of the time.
I agree with stuff about automatic conversions between types, however some compilers will warn you (unless you told it to shut up) about any narrowing conversions that you do, and that's the main trap that people fall into.
The majority of debugging/compiling tools are designed primarily with C in mind and are fairly simple to use also.
The majority of the rant about C was mainly not based off issues with C itself, but with the code quality of minimp3, which is quite depressing as C itself does have some bad traits imo such as function ptr definitions are bulky, : for bitfields, no predictability for most 'undefined behaviours', dodgy bitshifting too and probably more things I can't think of at the top of my head.