Modern C
modernc.gforge.inria.fr
modernc.gforge.inria.fr
In case, someone is interested here are the various C standards drafts:
* C89/C90: http://port70.net/~nsz/c/c89/c89-draft.html
* C99 (N1256): http://www.open-std.org/jtc1/sc22/wg14/www/docs/n1256.pdf
* C11 (N1570): http://www.open-std.org/jtc1/sc22/wg14/www/docs/n1570.pdf
* C17/C18 (N2176): https://web.archive.org/web/20181230041359if_/http://www.ope...
Also, here is a nice Stack Overflow answer with links to C and C++ standards documents:
* https://stackoverflow.com/a/83763/303363
This answer usually gets updated as new revisions of the standards appear.
As for why the finalized versions cost money when the final draft is virtually identical all the time? I honestly don’t know.
https://github.com/koekeishiya/
plz recommend if you have more! (mostly interested in the graphics/audio realm)
He has also written an article on how he writes the code comments [0]; I've been using Design, Why, and Teacher comments a lot in my last projects and they have been very helpful, especially if you come back to the project after a long time or introduce a coworker to the code base.
I try to comment my code in a similar way, often ending up with up to 30% comment to code ratio. I know some (many?) people are very dismissive of this style and reject the idea you have to write any comments where the code is ‘self explanatory’, but I don’t care. It doesn’t take me a lot of extra time to document the general idea behind a piece of code and it makes my future life and that of other devs reading the code so much easier it’s easily worth the effort to write these comments and keep them up to date.
Like I said it’s rare to see OSS code written that way though, I’ve even thought on multiple occasions that some projects actively strip all code comments automatically for whatever reason, considering the total absence of any form of code documentation in them.
I'm one of those, but in your defense I think we need to maintain a very high bar for "self explanatory".
At the same time, I think it's important to be aware that comments not immediately adjacent to what they describe are usually not visible in code review, are often not visible to the person editing the code in question, will fall out of date without an active process to keep them up to date, and at that point will do more harm than good.
I'm curious if you have any particular mechanisms for addressing the above?
I'm trying in my current setting by introducing a mechanism for cross-references (anywhere in the repo, ^^{label} refers to @@{label}) that my CI setup can recognize and surface during code review. It's provided a bit of benefit, but it's underused, and it's too soon to have a good sense of long term impact.
I put 'self explanatory' in square quotes exactly because of that, at the time of writing the code many things may seem self explanatory for the person writing it, but even seemingly simple things can be very confusing for people who read the code for the first time. But obviously I'm not advocating comments that simply describe in words what the next line of code does mechanically, those are useless. But I do like comments that break up the algorithm in sections that are clearly marked by comments, even if some parts of those comments may seem trivial ('loop over all objects of type x and process those with property y', etc). For me the value is in being able to read the comments first and get an idea of the structure of the code and the though process behind it, before actually having to interpret the code itself.
>> At the same time, I think it's important to be aware that comments not immediately adjacent to what they describe are usually not visible in code review [..] I'm curious if you have any particular mechanisms for addressing the above?
Not really. In our team we don't really use code review tools for feedback about algorithms or higher level design of the code, my experience they are pretty awful for that anyway. We use bitbucket primarily to point out local improvements and ask questions why sections where added or changed, but for higher-level discussion we prefer offline tools like whiteboarding, and rotating people so everyone works on different parts of the codebase.
The main challenge I have is not all engineers are equally willing/comfortable/able to communicate their ideas in this way. Sometimes because of ingrained preference (comments are never necessary), other times because of language barriers.
WeeChat is written in a very clear and easy-to-follow C.
I love Michael Fogleman for graphics: https://github.com/fogleman/Craft
There’s a very vocal group that tells you not to write in C because it’s unsafe, immoral and you’ll burn in hell if you do but the fact is...I would never be able to understand the problem Rust is trying to solve until I’ve recreated the problems in C. By now I’ve hit a good amount of segmentation faults.
Would really love to know if there’s a list of all the possible ways to get one oF these in C.
Segfaults are not something C is aware of, it's usually the consequence of triggering an undefined behaviour. But UB being... undefined, all sort of things can and do happen, for instance a buffer overflow can segfault, but it can also overwrite some other variable, or the return address of the function, or do nothing at all.
That's really the problem with C. If UB meant segfault, then it wouldn't be so bad because at least you'd immediately know that something went wrong. Instead it's very easy to have small issues go unnoticed for a long time, until your application finally crashes in a seemingly unrelated location or, worse, some clever guy or gal manages to exploit the flaw for fun and profits.
Might you or someone else have a resource on learning how to do this?
[1] - https://old.reddit.com/r/embedded/comments/ilgyun/would_you_...
A resource you'll find immensely useful is ctyme's interrupt jump table[2][3], which is easier than reading Ralf Brown's ASCII formatted notes[4]. Of course there are easier ways to do it, but printf() debugging (or the assembly equivalent) is still by far the easiest to set up.
If you want to go write code for modern systems -- UEFI code, you can find a tutorial here[5] (I didn't write this, however, so I don't know how good it is, I can't find the resources I used :c), and the EFI spec here[6]
[1]: https://wiki.osdev.org/Main_Page
[1.5]: https://web.archive.org/web/20200901014356/https://wiki.osde...
[2]: http://www.ctyme.com/intr/int.htm
[3]: http://www.ctyme.com/rbrown.htm
[4]: http://www.cs.cmu.edu/~ralf/files.html
[5]: https://github.com/safayetahmedatge/efitutorial
[6]: https://uefi.org/sites/default/files/resources/UEFI%20Spec%2...
Here's a random bootloader init code I wrote a while ago for an ARM SoC (probably heavily inspired by uboot):
https://gist.github.com/simias/c31ae83127c77519621edfb76d599...
The vectors are a table that must be placed in a specific location in memory and are used by the CPU to know what code to execute when certain events occur. Here I only fill the "reset" entry to call my init code when the CPU boots up.
Everything in my "reset" function up to line 93 is low level ARM things: making sure the CPU is in a well defined mode, disable IRQs, initialize the cache (optional, but good for perfs) etc... All that stuff is usually described in the docs for the CPU/SoC/Controller. Note that this SoC is fairly complex, if you used a smaller microcontroller it's usually much more straightforward.
Then after that you have a loop to clear the BSS (the _bss_start and _bss_end are generated automatically by the linker script, normally when you load an executable the operating system clears those regions for you, but obviously here we're on our own).
Then I set the stack pointer register SP to _stack_base which is also defined in the linker script to point at some memory location (typically the top of the RAM). And then we're good to go, we can do the rest in C with only small bits of ASM here and there for special instructions.
I think the best way to learn that stuff is by doing, get some ARM chip with an u-boot loader and dig into it, break it, see what code does etc... Obviously you need a good knowledge of C to begin with, and at least a basic understanding of assembly, so start with that if you're not there yet.
Also: run global initializers. Getting kinda fuzzy on the boundaries, but I'd also add in the process to copy the initial contents of the .data segment from read-only to read-write storage.
So I push for Rust (and sometimes Zig), because they're better than C in many ways, but they do lose the availability and some of the flexibility.
So for instance, if what I need to do is just manipulate memory with the least overhead possible, C is a really good tool for the job where Rust adds unnecessary ceremony. For instance, if I'm just walking arena-allocated memory, the borrow checker does nothing for me.
Also I think C plays an important pedagogical role. When writing C you learn a lot about what a computer actually does, whereas with Rust you're working with an abstract system.
I think a lot of the reason C gets a bad rap is because it has been used historically in a lot of use-cases where a language with better safety guarantees would have been a better fit, but that alternative didn't exist until maybe quite recently.
That has not been true in decades, if ever. C is very much defined in terms of an abstract machine, and the behaviour of C is not that of the underlying assembly, which itself is an abstraction over the µops which actually tell the hardware how to work.
Sure if you dive down into memory-model and CPU-internal details like micro-ops and pipelining then this model quickly falls apart, but on the surface, the "C abstract machine" still maps pretty well to what's happening under the hood.
Higher level languages are much further removed from C's model of the CPU and memory. And despite that "distance" there is virtually no difference between C's and C++'s (and often Rust's) performance. And at the end of the day, it doesn't matter how low level your language is if it isn't getting you any performance uplift.
Sure if you dive down into memory-model and CPU-internal details like micro-ops and pipelining then this model quickly falls apart, but on the surface, the "C abstract machine" still maps pretty well to what's happening under the hood.
A core part of the "C abstract machine" is sequential execution. Instruction-level parallelism is par for the course in modern processors, many of which have hundreds of instructions executing at any given time. So C hasn't been reflective of typical desktop/server CPUs for the last 30 years at least.
Hardware vendors spend a lot of money to make C code execute quickly, but we're now seeing the absolute limits of that.
I'm not seeing that in real-world projects though, and it's not only about performance. These are just two random examples from my own experience of where high-level language features get in the way (I guess the TL;DR is: yes, it's possible to write high-performance software in high-level languages, but it's often more effort, and may require to write unreadble and "non-idiomatic" code, because you need to "appease" the compiler much more than in a lower level language):
https://floooh.github.io/2018/05/01/cpp-to-c-size-reduction....
https://floooh.github.io/2016/01/14/metal-arc.html
Also let's not forget about build times.
The borrow checker is most useful when you are using it in a team or just using external Rust code. By forcing code to be explicit the borrow checker ensures that most, if not all, implicit restrictions are propagated through the codebase. Naturally this makes Rust far easier to compose than C.
> Also I think C plays an important pedagogical role. When writing C you learn a lot about what a computer actually does, whereas with Rust you're working with an abstract system.
C is to modern hardware like the Intel 8086 is to an Intel Xeon Platinum. The mere fact that ISO C runs on some many different types of architectures and platforms proves that C is far removed from any particular computer's mode of operation (which isn't a bad thing). Even programming in machine code wouldn't reflect a CPU's true mode of operation because of things like Macro-op Fusion, Register Renaming, TLBs, etc. Processors nowadays are probably the most dynamic and complex "programs" seen in the industry.
If C and Rust were graded on "similarity" to hardware, then C would be neighbours with Rust and hardware would be located ten timezones ahead.
> I think a lot of the reason C gets a bad rap is because it has been used historically in a lot of use-cases where a language with better safety guarantees would have been a better fit, but that alternative didn't exist until maybe quite recently.
Ada is only eight years younger than C, Pascal pre-dates C, etc. There were alternatives but people used what they knew and what their systems were written in. But to be fair, it is far easier to find new languages and tools today than it was pre-internet. However, the problem with C being used in inappropriate places hasn't gone away. People still push to write new applications and software in C despite the safety problems and developer ergonomics.
[1] https://andrewkelley.me/post/unsafe-zig-safer-than-unsafe-ru...
This is the wrong conclusion to derive from one example where Zig checks alignment and Rust chooses not to. (Nor is the conclusion particularly profound or useful: comparing languages by an attribute in cases where they have said they explicitly do not guarantee semantics of that attribute-safety in this case-is not useful.)
C could have an optional bounds checking mode. It is possible to design such a feature in a way that allows progressive adoption without breaking compatibility with existing libraries.
C could also choose to move some forms of Undefined Behavior to implementation-defined or even defined.
Many of the current standards committee members are active impediments to progress. We'd all be better off if we hit the reset button on that. Just look at what the C++ group has been doing by comparison...
I get the reasoning to learn it. Just after you do learn it, stop and use something else. Unless you absolutely have to. I don't think anyone has attempted to compute the amount of money spent on buffer overflow hacks. But I'm guessing it's in the 10s to 100s of billions of dollars if you count all the viruses that took advantage of it in the past in Apache web server or Windows NT's NTFS (or sharing etc...) codebases. And that applies to C++ too.
And hey, I know that sounds extreme - but there really are situations of "unless you absolutely have to" but I wouldn't expect anyone to write a brand new queuing system or database in C or C++ anymore given the alternative languages available.
Then you expect wrong, at least about Database Management Systems. New performant ones are often (/ mostly?) written in C++. At least that's what it's like for analytical DBMSes.
Some examples:
* HyperDB: https://www.hyper-db.de/
* DuckDB: https://duckdb.org/
* Vectorwise/Actian Vector: https://www.actian.com/analytic-database/vector-analytic-dat...
And you might also be surprised that many problems originally carried over from C to C++, like dereferencing nulls and memory leaks can be done away with - not by being super-careful, but by sticking to newer language facilities and using some static analysis. Other, like buffer overruns, are easier to avoid with things like spans and ranged-for loops.
Why even bother stating this? “I don’t have any data to back this up but look at this huge number!!!!”
I know the $ amount is greater than all other security risks in the history of software - even if I don't know the number the only reason to NOT post something about that history would be to defend C/C++, when in fact I am attacking C/C++'s record here.
Valgrind is your friend, I couldn't imagine writing C without it anymore.
Many compilers also ship with verifiers for undefined behavior which has saved me a couple of times.
With all those combined it's quite hard to hit segmentation faults and memory corruption and the source code will generally become "cleaner" and more robust.
The issue with C is the silent issues like buffer overruns, i.e. even if you program defensively and with a proper security model you're still opening up to the possibility of leaking your missile codes by accident because you committed bad code
Personally I think bugs should be addressed, with better debug modes and tools. Take this example:
int i, array[10];
for(i = 0; i < 11; i++)
array[i] = 0;
A clear bug, right? You could write this type of bug in any language. Many languages would stop you from writing outside the array. In C its undefined behavior.Undefined behavior means that the compiler doesn't have to do anything, therefor it doesn't have to at run time test if the code is correct. It can assume that the programmer knows what they are doing. Thats why C has great performance, but also why bugs can be silent and hard to find.
But, It doesn't meant that the compiler has to ignore the issue. Its undefined behavior, so the compiler is free to do all the checks that a language that doesn't let you write outside an array do. Sure, then C will become just as slow as other languages, but you can now find and remove the silent bug.
This is why I think that debug mode, should be very different from release mode. The two compilers have entirely different design goals. Some compilers are starting to do this. Visual studio will for instance break if you access a lot of uninitialized values in debug mode, something that it will ignore in release mode.
There are loads of things that could be done to improve debugging that doesn't require the language to change or to slow down the release runtime.
I know what you're trying to say, because at least 95% of segmentation faults are easy to fix, but I take issue with this. The hardest bug that I've ever fixed was a segmentation fault!
You've obviously never experienced random crashes where the best idea anyone has is to bisect code changes in production and see when the crashes stop.
And you've obviously never experienced random crashes on customer machines where you've tested the crashing function for weeks, poured over crash dumps and ran the exact same function with the exact same data and could never reproduce it.
Buffer over runs, can be trickier, but I have plenty of tooling to find the issues.
Thread safety issues are what nightmares are made of.
I'm not sure learning by repeating the failures of the past is a very scalable practice... :-)
When I was done with my professional C++ stint, I fell in love with the simplicity of C. It was like a breath of fresh air. I stopped using C largely by accident: everyone who paid for my work wanted Common Lisp or Java, and that lasted until the start of the deep learning craze seven years ago.
I thank the author for making a CC licensed PDF of his book available.
A few C books I like:
* C: A reference manual.
* C: Interfaces and Implementations.
* Deep C secrets.
* C: Traps and Pitfalls.The main reason is that, it was a very dry read with various rules and syntax etc and I got bored looking at it quickly. It almost felt like I am reading an abridged version of the standard itself. The exercises didn't appeal to me much.
I didn't mean to dis the efforts of the author at all. There is no other book that covers the modern parts of the standard well. I am not aware of any other C book that covers the wide chars than this book.
"Modern C" takes a staged approach, going through
+ Level 0 Encounter
+ Level 1 Acquaintance
+ Level 2 Cognition
+ Level 3 Experience
This sounds like an interesting organizing principle, but in practice it is implemented quite rigidly, and produces, in my opinion, two serious weaknesses.First of all, organizing by levels of detail, rather than treating individual topics in depth all at once, means that a fair amount of related information gets spread around the book.
This might be OK if the book had a comprehensive index, or consistently used forward-references to point the reader to a detailed discussion, but the book is generally weak in these sorts of cross-references.
For example, suppose you want to look up formatting strings for use with `printf()`. The index (I am looking at the Manning version) lists exactly one entry for `printf()`, which is on page 8 -- and the actual text there just mentions basically that `printf()` is a function that exists.
You could also look up "format", which gives one reference, to page 5, where the term is introduced but hardly even defined. The index does have more entries for `snprintf()` and `sprintf()`, if you happen to think to look for those. But after you play this game for a while, it becomes clear that the index takes a quite mechanical view of these functions, content to list locations where the function gets called out in the text, but without any great concern for where they are discussed as a cohesive collection of related functionality.
To me, this levels-of-experience approach makes sense for some things, like deferring discussion of threads, and setjmp-longjmp, till later chapters. That's why I say that the book errs by sticking to the game plan too rigidly.
The second major shortcoming of this rigid levels-of-experience approach is that the book delays discussing pointers for a long time. They aren't really taken up until chapter 11. Literally the book doesn't show a basic "swap" function until page 170. This means that a lot of code examples that come before chapter 11 do quite a dance to avoid using pointers.
When, finally, pointers are discussed, I felt that the text was somewhat lackluster. There is nothing wrong with it, but if this chapter was a blog post by some random C programmer, and got posted to HN, I doubt it would garner a lot of upvotes by enthusiasts saying it was the best description of pointers they had ever read.
Thus it had an anti-climactic effect for me -- I expected that deferring the discussion of pointers for so long would have some kind of pedagogical payoff, but there wasn't anything to it that, to me, justified putting it off for so long.
I had a few other problems with the book, but nothing that, by themselves, would keep me from recommending it.
I might recommend the book to somebody who already knows C, and wanted a deep dive on certain topics. However I would steer a beginner away from it.
Somebody else on this thread mentioned another new entry to this field, "Effective C" from No Starch Press. I am in the middle of this one, and so far like it a bit more (though I am not yet a full-throated enthusiast for it). I would note that "Effective C" introduces a swap function on page 16. That would be wildly premature to the author of "Modern C", but in the end I think it will serve the beginning reader somewhat better.
Along Modern C I found these modern(ish) C books:
- C Programming: A Modern Approach, 2nd Edition by King
- 21st Century C by Ben Kelemens
- Learn C the Hard Way by Zed Shaw
Can anyone recommend some of these for my background?
Besides Rust I've mostly used Webdev languages JS (mostly TypeScript these days), Python, Ruby, PHP, some Java back in the day.
"Secure Coding in C and C++, Second Edition"
https://www.oreilly.com/library/view/secure-coding-in/978013...
"Extreme C"
If you're coming from COBOL and wanting to learn C (like the target audience at the time the book was written) ok perhaps it is a good book for you. But if you're coming from something like Python or Java or JavaScript (or even no language at all) there are better options such as K N King's A Modern Approach.
K&R is a fantastic book in its own right and I certainly think once you feel more comfortable with C it is a superb book to read and more importantly complete as many of the exercises as you can.
It is that I have seen too many people come from higher level languages or with no programming knowledge and find K&R frustrating due to its assumption the reader is already a programmer in some other (1970s) language with a fundamental understanding of some programming concepts.
However it has gotten so deep into IT infrastructure, thanks to the hegemony of UNIX/POSIX clones, that even if starting today no more greenfield software would be written in C, and its copy-paste compatible languages, Objective-C and C++, it would take generations to clean it up and it would never be 100% replaced, as proven by mainframe environments and their languages.
So for the use cases where C isn't going away no matter what, we should strive for newer generations to improve their code quality and not to repeat bad practices from the past.
OTOH a microcontroller is an acceptable approximation of a PDP-11, so much of the old approaches are very directly applicable.
Not going to buy another copy just to make someone happy on Internet.
But still, most examples don't proper error correction, don't teach about use of bound checked strings and vectors, and if I remember correctly there are examples with gets().
My copy has no examples that use gets, although it is mentioned and I would agree that any such mention without a disclaimer that the function is impossible to use safely is a defect. Error handling, however, is generally present (or left out for brevity and noted). The functions in the standard for dealing with bounds checks are a new addition to the standard and a pox on the language regardless so it's not the best example of something new that the book should cover.
This is something that the book fails to teach, as it also has no mentions of modern static analysers practices, naturally given the book's age.
So at the end we get yet another C newbie writing future CVEs.
If you drop down to C on x64, you likely deeply care about cache efficiency, pipelines not stalling, etc.
Architecture which it targets doesn't matter.
I want to see what are best practices, such as what is done in the generic driver's on GitHub by Bosch for their sensors. (Such as the BMP280)
They tend to warp the structure of your software and prevent usage of many useful language features.
If you're not doing safety work, the Linux kernel is a great example of 'nice' modern C.
https://github.com/BoschSensortec/BMP280_driver/blob/master/...
...?
Because there is nothing inherently modern about this code, this is how C has been written for decades if not from the very start (I think you will struggle to find patterns in this code that are not covered by KnR).
That being said, there are books that cover relatively new developments in the C language and its ecosystem, such as "21st Centry C" by Ben Klemens (but even this book is something like 5-10 years old).
The complexity about embedded starts want to incorporate the testing with things like IO, which is more end to end testing rather than TDD.
I'm not exactly sure why it's "modern" C, but it is an introduction to C via establishing a somewhat rigorous and complete fundamental understanding of it.
http://www.iso-9899.info/wiki/Books
I like that it includes books to avoid, and ideas for further topics.
>"Free eBook With Every pBook! If you are an owner of a Manning pBook you can get a free eBook at any time easily from your account. If you prefer NOT to have a Manning account no problem, we won't be offended: we will send you a one-time download link after purchasing the pBook at manning.com. If you did not buy the pBook from manning.com, you can still get the free eBook in all available formats by setting up a Manning account, and registering your copy."[1]
It's also worth noting that you are supporting a a small independent publisher that's DRM free, as well as supporting independent authors. They also regularly have sales, often with 50% discounts, often on holidays. I wouldn't be surprised if there was a sale on Monday.
OK, I just bought the bundle. If anybody is interested, I can give one (1) of you, one (1) copy of the ebook for free.
EDIT: Besides, one of the electronic versions in the bundle is sold separately for $24. Assuming that the other electronic version has the same price, the print version alone is worth $12. I don't want to pay $48 for something that will be "thrown away".
* "Non-modern" C books, like K&R ?
* Other Modern-C books, like "21st Century C": https://www.barnesandnoble.com/w/21st-century-c-ben-klemens/...
I agree with the review author. The quoted paragraph about hash tables is shocking!
FWIW if you're looking for practical projects; part III of https://craftinginterpreters.com/contents.html is in C, and if you're interested in graphics stuff SDL is really fun. It all runs in the web / browser too with Webassembly.
But I'm terribly sorry if you didn't get it. Just a simple play on words. It's certainly not insular or community destroying - an accusation like that is simply hyperbolic. Perhaps like most living things, growth only comes when there's opportunity to being tested and stressed.