Some Were Meant For C (2017) [pdf]
cs.kent.ac.uk
cs.kent.ac.uk
https://blog.regehr.org/archives/213 gives great insights into how undefined behavior in C and C++ can be difficult to reason about and cause problems.
The lack of bulletproof memory safety and easy-to-stray-into undefined behavior of C and C++ make it easy to create code that is difficult to fully grasp how it will behave, especially when optimizing compilers are used. The C / C++ code runs really fast, but there are hidden dangers lurking.
I don't doubt that C and C++ will be with us for a long time to come, but the growing use of Rust, Zig, Ada, and others show that better alternatives exist and that they will replace the use of C and C++ for many domains and use cases.
Edit: Downvotes? Did you read my whole comment? I am saying that C / C++ are not secure for real-world use cases.
The security implications of undefined behavior are poorly understood by software developers: https://arc.aiaa.org/doi/pdf/10.2514/1.I010699
No, undefined behaviour is a serious problem, especially regarding security.
There's a long history of serious security problems in C and C++ codebases due to unintended invocation of undefined behaviour. These issues continue to arise even in well-resourced C/C++ projects with highly skilled developers, such as the Linux kernel and Chromium.
I admit I don't have hard numbers to hand, and sadly it's rather rare to see decent rigorous comparisons, but I don't think C and C++ have that much of a performance advantage over, say, Ada or Safe Rust.
My inner pedant feels it's necessary to note that Ada is still an unsafe language, but still, it's much less unsafe than C. It also has excellent support for optionally disabling all sorts of runtime safety checks, whereas the design of C and C++ make it extremely difficult to implement such checks as 'opt-in' features in a compiler.
One example is that undefined signed overflow is important when the compiler might need to rearrange a loop index - PPC prefers to count down, x86 doesn't care.
Anyway, I think undefined behavior + ubsan is better than defining all behavior. If something's undefined you know it's a bug every time you see it. If it's defined, how do you know it's wrong?
> One example is that undefined signed overflow is important when the compiler might need to rearrange a loop index - PPC prefers to count down, x86 doesn't care.
Signed overflow is undefined behaviour regardless of the target hardware architecture. The compiler is permitted to assume the absence of signed overflow and to optimise accordingly. Unless the compiler documentation specifically says its ok, you still have an undefined behaviour problem. It might happen to work fine, sure, but if you're serious about writing correct programs you should be aiming to deliver a program which is correct-by-definition rather than correct-by-coincidence.
> I think undefined behavior + ubsan is better than defining all behavior
It isn't. That's why Rust makes such a big deal of its Safe Rust subset. It's also a major selling point of SPARK Ada for safety-critical software. It's tremendously valuable to be able to categorically close the door on a whole family of potentially serious and difficult to detect bugs.
> If something's undefined you know it's a bug every time you see it.
No, you absolutely don't.
I already mentioned that high-profile projects like Chromium and the Linux kernel continue to face security vulnerabilities arising from unintended invocation of undefined behaviour. Section 7 of the paper discusses undefined behaviour but doesn't really explore its full consequences, so instead I suggest reading [0] and [1].
Undefined behaviour means exactly that: if undefined behaviour has been invoked at runtime, the behaviour of the program is not constrained by the C/C++ standard. The program is not required to explode loudly, it can do anything. It isn't required to behave the same way each time. Hopefully it will explode loudly, but it's possible everything will seem to be fine. In the worst case the undefined behaviour leads to a serious safety issue or security vulnerability.
Undefined behaviour is even permitted to 'time travel'. [0][2]
> If it's defined, how do you know it's wrong?
You use exceptions or some other well-defined means of detecting and handling runtime errors. For example, Java's NullPointerException and Ada's Constraint_Error.
[0] https://blog.regehr.org/archives/213
[1] https://blog.llvm.org/2011/05/what-every-c-programmer-should...
[2] https://devblogs.microsoft.com/oldnewthing/20140627-00/?p=63...
That's what I said. If your program is correct (it may-overflow like everything does, but doesn't dynamically overflow), then by undefining overflow, you can tell the compiler that it doesn't happen. That lets it reorder operations in a way it couldn't if every + potentially wrapped around.
> No, you absolutely don't.
You seem to have said "no" and then agreed with me. I was proposing trapping on all undefined behavior in debug mode!
> You use exceptions or some other well-defined means of detecting and handling runtime errors. For example, Java's NullPointerException and Ada's Constraint_Error.
Trapping is plausible, but that's not always how people want to fix undefined behavior. For instance some people want undefined memory reads to return 0, or want overflow to wrap. In that case it's hard to distinguish errors from intentional behavior.
I don't like exceptions very much either because control flow gets more complicated. Trapping like Swift does is fine, though.
I don't follow the distinction here. A correct program should never invoke signed overflow, regardless of input.
> by undefining overflow, you can tell the compiler that it doesn't happen
Right, that's essentially the effect of the standard saying it's undefined behaviour: it should never happen when the code runs.
> That lets it reorder operations in a way it couldn't if every + potentially wrapped around.
Right, or more generally, it enables various compiler optimisations.
> I was proposing trapping on all undefined behavior in debug mode!
Ok, I thought that by If something's undefined you know it's a bug every time you see it you were saying that UB always results in a loud explosion.
Unfortunately it's not easy to build a C compiler that traps whenever UB is encountered at runtime. An example: the compiler can't know the size of an array passed to your library. C uses 'thin pointers', unlike most languages where, whenever you pass an array, the callee can inspect the array's length.
> Trapping is plausible, but that's not always how people want to fix undefined behavior.
Ada's solution, of raising exceptions (roughly like Java), seems sensible. Of course, part of C's appeal is that it's very compact and lacks things like exceptions.
> some people want undefined memory reads to return 0
That doesn't sound reasonable. To implement that could be pretty burdensome.
> or want overflow to wrap.
This is something some compilers support as a non-standard feature. GCC supports it with the -fwrapv flag. I suppose it would be friendlier if there were a standard and portable #pragma to tell the compiler what you want, but I'm not sure it's a big enough problem to make it into the standard.
You can 'fake it' pretty well by converting to a unsigned integer type, doing the arithmetic, and then converting back to the original signed integer type. You could write a function to do this. You could use the preprocessor to defer to a compiler-specific intrinsic where one is available. I think GCC's __builtin_add_overflow would do the job but its definition isn't terribly explicit regarding wrapping behaviour.
I think this code would do the job portably, and I don't think it relies on anything platform specific. (I'm relying on the signed/unsigned conversions using two's-complement, I believe this is guaranteed by the C/C++ language specs. I've also used fixed-length integer types for good measure.) Godbolt tells me GCC can optimise it down to a single LEA instruction on AMD64.
#include <cstdint>
using std::int32_t;
using std::uint32_t;
/*inline*/ int32_t wrapping_add_int32t(int32_t num1, int32_t num2)
{
return (int32_t)((uint32_t)num1 + (uint32_t)num2); // Compiles down to LEA instruction
// Alternatively (also compiles down to an LEA instruction)
// int32_t ret;
// __builtin_add_overflow(num1, num2, &ret);
// return ret;
}
See also [0].> I don't like exceptions very much either because control flow gets more complicated. Trapping like Swift does is fine, though.
I agree it introduces action at a distance flow-control. I'm afraid I don't know Swift.
[0] https://stackoverflow.com/q/59307930/
Vaguely related fun: https://github.com/MaxBarraclough/IntegerAbsoluteDifferenceC...
The whole problem with undefined behavior is that it is faster to not check for undefined behavior and calling exit(1) (exit(0) would be a successful exit). Think about it in slightly higher terms: you implement a linked list that can search for an item and return a pointer to it once it finds it. Your implementation explicitly says that if you search for an item that isn’t in the list you will hit an infinite loop. I disregard the warning and let it search for an item not in the list. I hit an infinite loop. Could you have added a check for “if (current == head)” and bail then returning NULL? Sure you could but that introduces a branch and slows things down. Better label what can happen as undefined behavior because maybe on some future processor you’ll have that check because it’s cheap but on x86 it isn’t so you don’t. This is essentially the same thing.
In C's defence, you can easily get this behaviour by wrapping malloc in a safe_malloc function. Given that C lacks exceptions, it makes good sense to handle unable-to-allocate by returning NULL as this leaves the door open to all possible strategies.
This is wrong on two points.
Firstly, real world compilers very often do handle various kinds of undefined behaviour with immediate termination. On many platforms, dereferencing a null pointer will result in a segfault. Sometimes compilers generate code to trap if undefined behaviour would result. In the C++ standard, some errors are defined to result in a call to std::terminate, rather than undefined behaviour. [0]
Secondly, as IgorPartola indicates, doing this isn't irresponsible, it's the least bad way to handle undefined behaviour. If your loop has overrun the end of your array, you generally don't want the execution to silently proceed with invalid data, you want execution to end immediately.
Like I said, undefined behaviour is a problem even for Chromium and the Linux kernel. We've seen that just be diligent isn't a solution. We're way past that now.
> Every time you see these problems happening, the reason is that developers not only wrote incorrect (undefined) code, but also didn't test the code path that was affected.
Technically true, but not insightful. Testing can never be exhaustive, so this observation doesn't light the way to a solution.
> While I agree that a compiler should do a better job, ultimately the responsibility is on programmers
This statement is true of the C language, but it isn't a solution, it's the problem-statement.
It's not true of a language like Safe Rust, where the language itself closes the door on undefined behaviour, and where it is not the programmer's responsibility to avoid invoking undefined behaviour.
The major compilers already have ways to test for undefined behavior such as -fsanitize=undefined. Projects need to use these flags and test more.
Undefined behaviour is responsible for a non-trivial fraction of the security vulnerabilities of C and C++ codebases. Undefined behaviour in an application can be entirely eliminated by writing the application in a safe language. That's the point.
> The major compilers already have ways to test for undefined behavior such as -fsanitize=undefined. Projects need to use these flags and test more.
Do you really think the Chromium team isn't aware of that flag in GCC and Clang? If there were an easy fix to the problem of accidental invocation of undefined behaviour, the problem would have gone away years ago.
It's useful for a C/C++ compiler to offer to add runtime checks for a subset of the possible causes of undefined behaviour. We agree more projects should use such tools. As we're seeing, though, this isn't a silver bullet. Even extremely well-resourced and security-sensitive codebases end up with UB issues.
Memory safety problems are still a sizable proportion of the CVEs associated with C and C++ programs.
> The major compilers already have ways to test for undefined behavior such as -fsanitize=undefined. Projects need to use these flags and test more.
This is NOT a solution -- you cannot test every case. In fact, not even close: The possible state space to for signed addition of two 64-bit ints is 2^128. That is infeasible. A compiler for a 'safe' language CAN prove the absence of certain behaviors (UB being one of them).
It would be interesting to make a following experiment: * find and patch UBs in multiple opensource projects which are relatively popular (at least used not only by an author) * send pull requests wich UB fixes and see which fraction will be closed as "won't fix".
It requires a lot of time, but may show that many developers don't care about UB.
These points aren't wrong, but it understates how pernicious undefined behaviour can be. As I stated earlier, even well-resourced high-profile codebases like Chromium and the Linux kernel play cat-and-mouse with security issues originating in undefined behaviour. It's not just a matter of hiring diligent coders and using the usual quality-control processes (testing, code reviews, static analysis).
It's possible for C/C++ code to be entirely free of undefined behaviour, but in practice this only seems to happen when formal methods are used, or when people use code-generators/transpilers to avoid writing C/C++ by hand.
Anything can be secure (and conversely anything can be insecure). The theoretical potential doesn't matter because real life is never the theoretical best case. What matters is the overall risk (is liklihood * how bad < benefit?)
Exactly; real life C and C++ code that does real work tends to be insecure.
40 years ago, the case could be made. But C and C++ are not new languages, and the fact that just barely shy of no one can demonstrate the existence of secure C or C++ code bases without staggering levels of effort put into that process is data now, not just anecdote.
(And let me emphasize the effort as my yardstick. Writing truly secure code is arguably something nobody has ever done at scale in any language... but C and C++ are certainly unique in the sheer level of effort it takes poured in to them to even match what a number of other languages come with out of the box, let alone exceed them. If you aren't using some very high quality and fairly expensive tools like Coverity on a routine basis, you aren't even close.)
Agreed. That is why C and C++ should be replaced with more reliable alternatives.
"Many years later we asked our customers whether they wished us to provide an option to switch off these checks in the interests of efficiency on production runs. Unanimously, they urged us not to--they already knew how frequently subscript errors occur on production runs where failure to detect them could be disastrous. I note with fear and horror that even in 1980, language designers and users have not learned this lesson. In any respectable branch of engineering, failure to observe such elementary precautions would have long been against the law."
While Intel's MPX extensions were a failure, Solaris makes use of ADI on SPARC, iOS uses PAC, Google is researching adopting ARM MTE into Android, ARM sponsors CHERI with Morello project, Azure Sphere has Plutonium, and so forth.
However these are all pontual efforts, and their large scale effects will take years or decades to start making a difference.
https://clang.llvm.org/docs/UndefinedBehaviorSanitizer.html
UndefinedBehaviorSanitizer (UBSan) is a fast undefined behavior detector. UBSan modifies the program at compile-time to catch various kinds of undefined behavior during program execution, for example:
- Using misaligned or null pointer
- Signed integer overflow
- Conversion to, from, or between floating-point types which would overflow the destination
GCC has similar features.
Which makes it unusable with binary libraries.
Security is the process. It contains continuous risk assessment, penetration testing, fuzzing and using various other tools throughout the product development to eliminate attack vectors. Only then you could build a secured product. Just rewriting everything in Rust won't make it.
Yes, but do the programming languages / tools we choose make the security process easier?
There is no perfect solution, but some programming languages / tools are better than others for preventing unexpected behavior that can lead to insecure programs.
I sometimes work on the infosec side of the house. It's easy to point at vulnerabilities due to endless memory access problems in C code and fixate on that. And it's true, so it feels satisfying.
Just rewrite everything in not-C and this is fixed! And that's true as well. But remeber here the word "this" refers to the memory access bugs. It doesn't refer to vulnerabilities in that sentence.
Plenty of systems without a single line of C/C++ code and we have no shortage of ways to break in anyway. So in one sense, everything changed. But wearing the black hat, nothing really changed since I compromised the system anyway. A new language without all the engineering process parent post describes, won't magically get there.
For an existing system "just rewrite everything" will guarantee way more bugs for years to come simply because the old system has been battle-tested and reinforced for years. In the long haul the rewrite will converge to a better state but that haul is long (and assumes budget will remain in place long enough, which might not happen so you end up with a half-baked rewrite).
For new systems starting from scratch that may in any way ever be run in a security relevant context, sure, starting with not-C is a good idea today.
Unfortunately Rust seems to be the only alternative for use cases where C was actually needed (if it could be written in Java or Go or Python or ... then it didn't really need to be in C in the first place). And sadly rust is fairly user-hostile language so my guess is plenty of new projects will start with C for a long time due to lack of friendly alternatives. And tooling.
Could you elaborate on that? I started using Rust at the end of last year, and have found it to be one of the most user-friendly languages I've learned. There is clear documentation that is easy to generate with the standard toolchain, along with a very good language guide.
The compiler error messages are an absolute joy as somebody coming from C++. Pointing out exactly where the borrow checker found errors, or exactly what the type signature should be for a trait, is really useful.
For example, in the case of const generics, I've used the equivalent in C++ to make a class for a geometric vector in N dimensions. The size varies depending on the number of elements, but is known and can be checked at compile time. The existence of const generics is very closely matched to this particular use case.
Forcing it down to an overly simplistic setting, if I need to add 5 to a number I can increment 5 times, or add 2 and then 3, or just add 5. Which of those options should we remove?
Perhaps we mean, instead, that there should be one best way, but (sliding into metaphor, I hope) what about when we have a language without 5 and we are considering adding it? Is the use case already handled by 2+3? How are we to decide?
Language wise rust is better the python in every way except for compile times.
Language-wise, it feels like they are intended to solve different problems, so Python vs Rust is a weird comparison. Granted, that may because I started learning Rust as a replacement for C++, so that has been my point of comparison.
Rust seems to me to be user-frienndly, though for the cases where it's safety features are critical and at the fore, it front-loads pain that would be dealt with eventually, which can make it feel intimidating. But its forcing you to confront things that are likely to cause subtle bugs if dealt with sloppy (or overlooked), not actually adding unnecessary complexity.
The simple classes of bugs enforced by rust, are also caught fairly quickly with any kind of memory sanitizer (valgrind?), combined with static analysis tools (coverty?), and automated code quality standards (misra). Run a CI/Code coverage monitor while looking for for these kinds of errors, and I would bet the results are actually better than plain rust due to the maturity of some of these tools.
About valgrind or sanitizers: they're runtime, so just like tests they can only show the presence of errors, not their absence. Like dynamic type checking.
Valgrind might not catch 100% of errors, but at least what it catches are actual errors I care about.
If you're writing safety-critical software though, being forced to restructure your code to satisfy the type checker (which, in this case, is kind of a simple proof assistant) seems like a sane tradeoff.
IMHO, Rust simply isn't good enough at catching all types of bugs to justify rewrites at this point, and its likely when you look at some of the work being done at the processor manufacturing companies that they don't believe it either.
Consider: https://en.wikichip.org/wiki/arm/mte, https://en.wikipedia.org/wiki/Intel_MPX, and https://lwn.net/Articles/718888/
There are quite a number of these in the pipeline, which make some of what rust does redundant.
Rewriting C or C++ applications in Rust won't fix all security problems. But it would be forward progress. It's just like wearing seat belts won't stop people from being injured in car crashes. You could say that "safety isn't a product but a process," which is fine, but if your safety testing process finds that wearing seatbelts reduces injuries substantially it seems pretty obvious that using a seatbelt product is the way forward. At least until someone invents a better replacement for seat belts, or supplements them with other features like airbags, automatic braking systems, self-driving, etc...
(Caveat: Obviously 'unsafe' is relevant to this discussion, but at least you only have very small areas of code you need to check extra carefully.)
True, but if the process discovers that a lot of lives are lost because of the use of unsafe tools, it changes the tools to add safety features (https://en.wikipedia.org/wiki/Chainsaw_safety_features), or ditches them altogether (https://en.wikipedia.org/wiki/Hazard_substitution#Processes_...)
So, yes “rewrite in rust” isn’t the full answer, but that does not imply “don’t write it in C” isn’t part of the answer.
So you can easily make it secure by doing something not easy?
Something that requires discipline, willingness and unconventional practice to be secure is by definition not secure. A secure language is the opposite: The default behavior and conventions are mostly secure and you need to go out of your way to make it unsecure.
Otherwise, with enough time and enough people you are guaranteed to be doomed.
Anecdotally, I haven't seen an open source C++ code base that is uses type safety to that extent but it is certainly possible.
Is the kernel Linux not a real-world use case?
Now show us the list of replacement kernels written in Rust, Ada, Lisp, or whatever that are:
1) free software (or open source as a second best)
2) as widely distributed and as portable as Linux
3) supports as much hardware as Linux
4) supports the wide range of applications that Linux does
5) free of memory safety related CVE's as well
There is always a series of tradeoffs - Linux has adopted a classic but unsafe programming language as its basis which has in turn allowed it to cover an enormous amount of ground in a (relatively) short time frame, support massive amounts of new employment and accelerate the development of the web.
It is not clear that an alternative design or implementation in another (safer) language would have moved the world as far forward as Linux has in the same time frame, or whether its even possible, as nobody has demonstrated this.
Apparently they are still stuck at the starting line, discussing endlessly how perfect their system will be when they finally choose the right language to implement it in.
As a result, anyone working on Redox, Fuschia, or whatever popular contender now have a massive library of driver source code to refer to when working in their languages of choice. This is not a resource that the Linux and 386BSD communities had when starting out.
Rust isn't as standardized or at critical mass yet, as C was when Linux was launched. The language still lacks comparative mindshare, no matter how popular it is here on HN. That's why there aren't enough developers.
I never said "Linux is bad and insecure" either, and when making blanket assertions like that, it's helpful to have a point of reference.
So I ask again, which free software, widely deployed, battle tested and highly secure operating system are you comparing Linux to? That was the very obvious point of my comment, which you have apparently missed (ignored?) twice now.
I think this is a good question, and I was wondering if someone would ask that. To be more specific, I think points (1), (3) and (4) are of primary importance, so I would address those specifically:
1) commercial, proprietary and expensive systems were very fragmented and incompatible with each other, especially UNIX, and VMS, Windows etc. Nobody seems to want to go back to that. I think this is the biggest hurdle for any new commercial system to overcome, otherwise Linux (or other, better free systems) will continue to dominate until a better replacement comes along. Look what happened to Solaris. Better than Linux in some respects, worse in some, but it failed to retain its market share. Windows is not hugely used as a server platform either like it was in the early 2000's outside the infrastructure required to support Active Directory, Exchange and Sharepoint.
3) obviously popular server hardware needs to be supported by the system, eg. HP, Cisco UCS, DELL etc. This should be evident by the amount of effort that vendors like Redhat put into patching their kernel to ensure compatibility with hardware from major server vendors. If the proposed alternatives don't support these hardware platforms then they are going to find it hard to gain traction.
4) you aren't going to migrate people away en-masse from a C-based system like Linux without being able to support the major languages and ecosystems that rely on C and/or C++ for their runtimes. People will still need to run their Tomcat applications, Ruby on Rails, Python, NodeJS etc. all of which currently depend on C or C++ to some extent. I think at least Redox is addressing this somewhat with their "relibc" implementation, which would allow "legacy" C codebases to be hosted there in the future, but let's be pragmatic and admit that it's probably a long way off.
As somebody who builds/manages emergency services communications platform based on Linux that supports over 70,000 front line emergency services workers, who in turn provide criticial services to over 6 million people, those are the reasons why I see no currently viable platform worth migrating to in the short to medium term _for our particular application_. In addition to self-hosted environments like ours, Linux currently also has a very particular grip within cloud hosted environments as well (which aren't particularly relevant to me at this point in time).
In order for non-C based platforms to take off, you either need to support the C runtimes, or be prepared to wait a long time as other languages that don't depend on a C runtime become popular and the ecosystem grows enough.
There were no free software memory-safe languages suitable for implementing an operating system in during the 90s when a lot of the initial Linux development happened.
> In order for non-C based platforms to take off, you either need to support the C runtimes, or be prepared to wait a long time as other languages that don't depend on a C runtime become popular and the ecosystem grows enough.
C runtimes are everywhere at least partly because C runtimes are easy to implement. Most of the Lisp machines in the 80s had full support for C. And those were non-memory protected systems! Running a C program in a separate a process would let you use any existing C runtime; the C runtime doesn't care what language the system-calls are implemented in.
Do you honestly believe that there is a widely-deployed battle-tested free software operating system implemented in a memory safe language?
I assumed the answer was "no" from your original comment, but all of your replies can only be interpreted as if the answer is "yes"
It is already starting with MIT/BSD/Apache OSes for IoT.
Even if I indulge your fantasy for a moment - Linux (amongst other systems) has provided myself and many others with a solid career for the better part of the last 25 years, daily security patches from vendors and all, and I don't see any solid evidence of that changing before I retire.
Further I suggest most Linux people would be able to adapt if the world were to move on to an improved competitor with enough momentum to be adopted en masse - which hasn't happened.
The lack of formal verification features and easy-to-stray-into logic errors in Rust, Zig and Ada make it easy to create code that is difficult to fully grasp how it will behave, especially when large projects are concerned. The Rust, Zig and Ada code runs really fast and is usually memory safe but there are hidden dangers lurking.
I don't doubt that Rust, Zig and Ada will become more popular in time to come, but formally verifiable languages such as Verifiable C or Spark are actually "safe" in some meaningful sense of the term instead of giving everyone a false sense of "safety" in the form of memory safety features as if memory safety errors were somehow the only class of security critical error.
If a valuable tool or method is expressed in (or even only expressible in) a particular language, then so be it. Often, there are many more choices than people believe, and what is right for the person, so long as it serves the organization or need appropriately, is fine by me.
A language is a tool. Most languages do pretty much the same thing. Most languages' ecosystems, application adoption, and developers are much more important than the languages' seatbelts and headlights.
Perhaps: Language in-class variance is small (C# ~= Java, C++, C, RUST are close-ish). Cross-class variance is big (JS vs Rust). Therefore, recruiting a programmer from the same problem / language class is more important than the particular language.
Another example, one of our architects at work is a strong supporter of Apple, and wanted me to look into Swift. Well at the time you could get Swift for Ubuntu, but couldn't for any version of RHEL (that has finally changed now though). So again, writing my code in plain C was more of a win.
Not always. Google’s tcmalloc is actually written in C++.
I actually would have agreed with a lot of this in 2017. However, in the past several years, I think Rust has been a game changer. It can interface with the C ABI. It can access the same low-level abstract machine that C can. It doesn’t have garbage collection or virtual machines. And it provides memory safety out of the box. In addition, features such as strong types and pattern matching help less logical bugs as well (like the compiler checking that you did not forget one arm of an Enum).
This is born out now that a lot of security facing software is starting to do at least part of their internals in Rust, where they had been in C before.
I think Rust is and will be even more in the future a game changer in how we write the foundation programs and libraries that the computing world is built on.
I think it is more possible in C++ due to placement new, though tcmalloc is probably still relying on language extensions.
Implementing memory barriers (pthread_mutex_lock()) used to be strictly speaking impossible in C too, but it's possible now that it has a defined memory model.
The practical consequence of this is malloc is kept in a different library and not inlined into callers, and asan doesn't work with custom allocators unless you tell it they behave like malloc.
Is that due to aliasing rules?
https://en.wikipedia.org/wiki/Andrei_Alexandrescu
They've promoted it (holding annual conferences and such) but they don't have Mozilla or Google behind them, so resources are limited.
Edit: And here are some blog posts by Walter Bright about D as a Better C.
https://dlang.org/blog/2017/08/23/d-as-a-better-c/
https://dlang.org/blog/2018/02/07/vanquish-forever-these-bug...
https://dlang.org/blog/2018/06/11/dasbetterc-converting-make...
There are some rust features that I want in D, but we'll see about that.
https://run.dlang.io/is/TKOBgA
Once you move on to other goals (using the whole language, as is usually the case) there are numerous complaints. Some don't think the VS Code plugin is good enough and that sort of thing. Some argue that Dub, the package manager, is not good enough for their needs. I suppose like every language has people that try it and don't like it.
It doesn't take much to try it. You can use the online D editor and read the official tutorial. If you like it, you can dig in further to see if it has the ecosystem you need.
D has evolved (until the last few months basically) without any direct corporate backing and yet absolutely stomps quite a few very very well-funded languages from a design perspective (e.g. A template constraint in D is so simple it's one keyword, and you can even do your own error messages as a library and compose things, whereas C++ took almost my entire lifetime to standardise concepts).
D exists as a (large) group of very simple language design decisions - for example, we have
unitest { /* code here */ }
blocks that make writing unit tests much lower friction.We just had our monthly "Beerconf" online conference, and all I can say is that the knowledge per head in the D community is very high.
Some Were Meant for C (2017) [pdf] - https://news.ycombinator.com/item?id=19736214 - April 2019 (176 comments)
Some Were Meant for C: The Endurance of an Unmanageable Language [pdf] - https://news.ycombinator.com/item?id=15179188 - Sept 2017 (240 comments)
Do you know what subset of C NASA limits itself to? Or hw architecture? The rigour of their testing? Should all C developers follow the same restrictions as NASA?
https://andrewbanks.com/wp-content/uploads/2019/07/JPL_Codin...
https://nodis3.gsfc.nasa.gov/displayAll.cfm?Internal_ID=N_PR...
Hardware (according to Wikipedia) is a BAE Systems RAD750 radiation-hardened single board computer based on a ruggedized PowerPC G3 microprocessor (PowerPC 750). The computer contains 128 megabytes of volatile DRAM, and runs at 133 MHz.
https://en.wikipedia.org/wiki/Perseverance_(rover)
Testing sounds pretty rigorous, at least for large projects.
https://www.quora.com/What-does-a-software-engineer-do-at-th...
Personally I firmly believe that "all C developers" do not need to follow these regulations. It might even be counter-productive to slow down the development process for some clients. For safety-critical systems, these rules make sense. For little startups, they don't.
Developers are smart enough to learn these rules, so HR shouldn't ask for "5 years MISRA experience". It's really a choice of business model, time to market, and risk management. If you're a big company looking to cut costs, be careful about outsourcing firmware development to a little startup who might not follow these rules so strictly. I won't follow these rules for the stuff I throw together in my free time and put on Github, but I will be careful before committing code to master for medical device firmware.
So any indication on what Nasa would use on its Rovers has to be taken from projects that start from the point when Rust released 1.0
The types of analysis and programming practices used to send stuff to Mars is beyond what Rust, or D, or any other safer-systems-language tries to do. It's not that simple.
These types of projects effectively need to prove the absence of bugs using formal verification and very extensive testing. Surprise surprise, C makes it extremely expensive and theoretically difficult too.
For example: NASA wrote this project https://github.com/NASA-SW-VnV/ikos which uses abstract interpretation and would catch bugs in practically any language.
I have argued that C’s enduring popularity is wrongly ascribed to performance concerns; in reality one large component of it (the “application” component) owes to decades-old gaps in migration and integration support among proposed alternatives; another large component of it (the “systems”component) owes to a fundamental and distinctive property of the language which I have called its communicativity, and for which neither migration nor integration can be sufficient. I have also argued that the problems symptomatic of C code today are wrongly ascribed to the C language; in reality they relate to its implementations, and where for each problem the research literature presents compelling alternative implementation approaches. From this, many of the orthodox attitudes around C are ill-founded. There is no particular need to rewrite existing C code, provided the same benefit can be obtained more cheaply by alternative implementations of C. Nor is there a need to abandon C as a legitimate choice of language for new code, since C’s distinctive features offer unique value in some cases. The equivocation of “managed” with “safe”implementations, and indeed the confusion of languages with their implementations, have obscured these points. Rather than abandoning C and simply embracing new languages implemented along established, contemporary lines, I believe a more feasible path to our desired ends lies in both better and materially different implementations of both C and non-C languages alike. These implementations must subscribe to different principles, emphasising heterarchy, plurality and co-existence, placing higher premium on the concerns of (in application code) migration and interoperation, and (in the case of systems code) communicativity. My concrete suggestions—in particular, to implement a“safe C”, and to focus attention on communicativity issues in this and any proposed “better C”—remain unproven, and perhaps serve better as the beginning of a thought process than as a certain destination. C is far from sacred, and I look forward to its replacements—but they must not forget the importance of communicating with aliens.
It is quite more elaborate than other publications I’ve seen mentioned in those discussions.
I’ll quote section 6.2, “What is Safety Anyway?”:
> I have learned to enjoy provoking indignant incredulity by claiming that C can be implemented safely. It usually transpires that the audience have so strongly associated “safe” with “not like C” that certain knots need careful unpicking.
> In fact, the very “unsafety” of C is based on an unfortunate conflation of the language itself with how it is implemented.
> Working from first principles, it is not hard to imagine a safe C. As Krishnamurthi and Felleisen [1999] elaborated, safety is about catching errors immediately and cleanly rather than gradually and corruptingly. Ungar et al. [2005] echoed this by defining a safety property as “the behavior of any program, correct or not, can be easily understood in terms of the source-level language semantics”—that is, with a clean error report, not the arbitrary continuation of execution after the point of the error.
And the title, to me, evokes William Blake: "Some are Born to Endless Night".
Azure Sphere, Solaris and latest versions of iOS all rely on some variation of hardware memory tagging to tame C exploits.
It also uses a safe C dialect for iBoot.
https://support.apple.com/guide/security/memory-safe-iboot-i...
We know from 40 years of discovering memory-related vulnerabilities in even the most carefully written, rigorously tested C programs that writing safe C is intractable for real, human software engineers. So yes, C IS INHERENTLY UNSAFE. If you claim otherwise you clearly haven't been paying attention to what's going on.
What's wrong with memcpy here? As long as dst and src are both non-zero and the ranges of memory are not overlapping, the behavior of memcpy is well defined.
It's also insanely slow with gcc, compared to clang. Like factor 1000 with some compile-time constants.
1. Inertia. its already being used, so why not keep using it? I can't prove this, but it feels true
2. Lots of microcontroller vendors provide tooling for it. You are basically guaranteed to be able to run C on whatever micro controller you want. Even if you don't get a standard library, you can implement your own if you need to. Languages like C++ have large runtimes and are hard to port to lots of platforms
3. Its an "easy" language to use for people who work on these devices and the problems of an ECU are not the problems of a database for example. I've worked on large enterprise databases and now I work on cameras for a car-ish company. The scale is just different and you simply don't need as many features as in C++, haskell, rust, python. The hard part on a small embedded system is not the coding per se, its the algorithms. A lot of these devices work at fixed rates on a timer (or respond to a very periodic interrupt), have a well understood amount of work to do in a fixed window of time, and then go back to sleep until the next job is queued. Memory safety won't help you when critical failure is defined as being late to respond. Would I love to have total memory safety at the moment? Totally! But its the least of my concerns at the moment and if my toolchain doesn't support it, then oh well we just have to code to the standards of misra and follow standards like iso26262.
I should say, writing C on a large project sucks. I've seen it before and I personally hate it (large portions of the database were in C). Its just too complicated and while C++ isn't perfect, having things like a destructor and templates are really nice and just makes the problems tractable.
There is one other reason, and that's until recently, auto-qual MPUs with fancy floating-point units (or any floating point units) were very rare; hacks such as Qm.n notation were required to do anything semi-fancy with trig etc. This was true even 5 years ago, although I would hope by now 'decent' auto-qual parts that are cheap enough exist.
Lastly, there are a boatload of requirements for auto software; you have to fail-safe as your power is going away (e.g. a crash is happening). You need to have get reset to any value, and your CPU needs to detect something is wrong and reset itself. There are different failure tests for different subsystem; 'body' electronics isn't quite as stringent as propulsion.
There is also a misra C++ spec; I was unsuccessful getting even a pilot project with C++, as it's also substantially simplified in the 'legal' subset; and is rather nicer than C in many ways. But.... C is going to be with us for decades more, I think.
I don't think that's true. C is often slower than C++ because of how inlining works, plus some domains (e.g. compilers) it's best to just use a GC language from the get-go.
Hang on, you’re cutting a few corners there! It’s easy to use inline functions in C too.
Are you thinking of C++ templates, and the fact that e.g. std::sort is faster than qsort because it directly calls an inlined comparison function rather than a function pointer?
That’s true, although it’s possible to achieve similar performance levels in C via hand-rolled data structures or macro hackery. I’ll grant you that the efficient C++ code is more idiomatic and likely safer. On the other hand, idiomatic C code is likely smaller when compiled, which can be important for performance too.
I don’t believe “C is often slower than C++” is true in general.
And when it's not irrelevant, it's almost 100% certain that std::sort is not the right thing to use. Where it matters, it's probably possible to examine the context a little more closely and come up with a custom sort that runs in O(n) or at least faster than std::sort.
std::sort( begin( c ), end( c ), []( const auto& lhs, const auto& rhs ) { /* some comparison impl for the value_type */ } );
Same for everything that take a comparator in <algorithm>https://en.wikipedia.org/wiki/MISRA_C
Personally I don't agree with all the guidelines. This morning a colleague said "I'm so glad someone had the foresight to include [unused API function]". Strict MISRA compliance would require "no unused code". I also prefer having more comments. Generally though, I think the advice is helpful, particularly about specific types (uint32_t not int), avoiding malloc, and complexity limits.
This is a totally different style of coding to what I've been used to at startups in the past. It's "software engineering" instead of "move fast and break things". During my 20s, I'm glad I had more broad experience of creative solutions, new languages, janky code, late nights, pretty demos, shipping /something/ then iterating on that. Now I'm 31 and settling down, this slower-paced but much more careful approach seems more suitable.
I absolutely agree with that statement. When I do firmware for small MCUs I feel big fat zero need for any other language.
All the article's arguments for C apply substantially moreso to C++. Thus, the article leaves us with no objectively plausible reason ever to code in C, except where artificial constraints mandate it, or where merit doesn't matter. (I leave to the reader to decide where Linux and BSD kernels fit in that.)
There is especially no excuse for systemd to be coded in C.
So, in actual practice, no.
And of course no one who cares for their reputation is seen using shared_ptr.
So, the cost of moving a smart pointer onto the stack and, presumably, off again to someplace less ephemeral might matter, on a critical path. But it would be a mistake to exaggerate this: we are talking about a couple of unfortunate ordinary memory operations, not traps, function calls, or atomic synchronizations, when you might have hoped to be using just registers. If you are already touching memory -- something pretty common -- it is just more of that.
The most ugly aspect is that C++ has a lock-in effect like Whatsapp. As long as your codebase is in C, you can easily interface with nearly all other languages. Once important parts are in C++ though, only C++ can reasonably interface with it.
To the other point, you can write bad code in any language, Rust included. But you don't have to. A language can help by making good code easier to write than bad code. C fails so frequently by making good code much harder to write, instead.
I count 7, plus msvc, which adds another one. 8