Musings on the C Charter
blog.aaronballman.com
blog.aaronballman.com
i have been posting FOSS code since 1998, much of it in C, and not one line of it is something i would dare classify as "economically important." It was written for fun and much of the fun of writing in a language is lost when updates in the language break old code. (The reason i won't touch PHP anymore is because updates broke my long-working code too many times.)
C Charter folks: by all means, break long-standing features when folks build with -std=cNEW_VERSION, but keep in mind that -std=c89 or -std=c99 are holy contracts with decades of code behind them, all of which should still compile in 30 years when built in -std=c89/c99 modes, regardless of whether they're "economically important."
Now, nobody is stopping compilers from supporting old language modes, but it adds a lot of complexity, not for making the code compile, but for all the diagnostics which are different. If you look at the FE code of GCC, for example, a lot of code is dedicated this. If we could remove all this stuff at some point (which would still allow old code to compile, but one code remove diagnostics for things which were disallowed in earlier standards but are now allowed), then this reduce maintenance burden substantially. And if one really needs those old diagnostics, one could simply go back to an old version of the compiler.
Microsoft's C compiler did not support C99 until something like 15 years after the standard came out (and reportedly (according to colleagues who use that platform) still doesn't support certain features of it).
It happens with surprising frequency that someone reports to the sqlite project that The Latest Version no longer works on their pre-2005 environment and they'd like to see it patched to work there.
My point is: users of a given platform might not be able to use The Latest Stuff (or anywhere near it, for that matter) because their OS vendor is reticent, because they have to support an old platform which is not targeted by newer tools, or for whatever other reason.
> Now, nobody is stopping compilers from supporting old language modes, but it adds a lot of complexity, ...
i don't doubt that, and i do sympathize with the maintainers. i'm not saying they should never ever remove C89 support, but if they do then there needs to be an alternative, like a fork of the last compiler version which supported it, maintained at least to the extent that the rare genuine compiler bugs can be resolved.
https://herbsutter.com/2012/05/03/reader-qa-what-about-vc-an...
They only kept updating their C support to the extent required by ISO C++ requirements, and some key customers.
Anyone else still dependent on C was suggested to use clang, this is also why clang is part of Visual Studio compiler suite nowadays.
Around the time the Microsoft <3 Linux stuff started, they decided to backtrack on this matter, and started updating their C support, however since C11 made a couple of stuff from C99 optional, they decided to skip on those.
VLAs have anyway been proven a rich source of exploits, to the extent Google has sponsored the work to clean up the Linux kernel from all uses of them.
Also there were a couple of talks from Linux Plumbers Conference given by Kees Cook, if I recall correctly.
To some extent that's already possible with the "-std=..." option of course, but it would be better to also allow changing the standard version within the same compilation unit (at least if it turns out that including headers written in old C standards can't be supported, in that case I would like to wrap a block of #includes with a 'use c89 { ... }').
Random example for why it is important to allow older C versions to compile: I just made a mid-90's assembler for 8-bit CPUs written in C89 work in VSCode via WASM/WASI without requiring any code changes (and yes I've been looking long and hard for an alternative assembler tool which provides the same feature set written more recently without luck, I even considered writing my own assembler from scratch).
It would suck if I had to port the code base to a more recent C standard just to make it compile, and (if that would be possible) automatically formatting that code to a more recent C version would introduce a hard "before/after" split in version control.
I do not think it is feasible to maintain all language versions in parallel for eternity. What would make porting this code to a newer standard difficult in this specific case?
It's a good thing to only have declarations in headers, it simplifies parsing with 3rd party tools (to generate language bindings), and speeds up compilation (see the C++ stdlib headers for the canonical "why is implementation code in headers a bad idea" example). IMHO mixing interface declarations with implementation details was one of the cardinal sins of C++.
> What would make porting this code to a newer standard difficult in this specific case?
It's pointless busywork, and it's unclear how deep the changes should go (for instance: does it make sense to move C89 variable declaration from the start of scope blocks to their initialization point just because it might potentially be "safer" but touches half of all lines of code?).
The code demonstrably works as C89, and any bugs that matter most likely had already been fixed decades ago, also doing large scale syntax changes is just noise in version control which makes bug fixing harder (because that often involves diving into the change history).
Compiling with a new language mode does not force you to move variable declarations to the point of initialization, but you would now have the option to improve the code in this way when it makes sense. It is difficult to see this as an disadvantage. My point is that moving it to a new language version would usually not require changes to the code, except where the old modes were dangerous. So I would expect to find serious bugs when doing this.
That's actually a good point hmmm (not in case of this specific assembler, which seems to be quite robust and well-written), and even if there would be serious memory corruption errors lurking, they would be contained because the assemblers runs in a WASM VM.
For projects that are still maintained, a "moving standard" isn't that much of a problem, after all we've been conditioned that every new compiler update adds new warnings which "break" existing code ;)
...there is a certain value in taking a "finished/frozen" project and integrate that into a new project without requiring code changes though.
I'm sure there must be a good middle-ground for C, where obviously bad and outdated language features (starting with leftovers from the K&R era) can be removed without causing too much breakage even on most old code ... but then, why keep 'inline' ;P
...rather, doesn't need. C++ modules just (potentially at least) fix a problem that C++ created in the first place (exploding build times because of complex template code in stdlib headers).
Also, I have yet to see clear indications that C++ modules drastically improve build times in real world projects.
The trick to make them faster is the same as with C++, never compile everything from source unless required, this includes template code for common type parameters.
As for compile times see Microsoft Office modules migration, or that VC++ import std in C++23 is faster than a plain #include <iostream>.
I think this view is mistaken. Tools that only need small changes rarely are some of the most valuable, because they produce value at minimal cost.
> ...existing code that has been maintained to not use features marked deprecated, obsolescent, or removed in the past ten years is important; unmaintained code and existing implementations are not.
I understand that you can't support legacy code forever, but we need to get away from this idea that ten years is a long time for software to run, or that only software undergoing constant churn is worth anything.
> I understand that you can't support legacy code forever, but we need to get away from this idea that ten years is a long time for software to run, or that only software undergoing constant churn is worth anything.
The specific issue here is the code which lies at the intersection of the three categories:
1. Code that is N decades old and still in active use.
2. Code that needs to be compiled by the latest and greatest compiler versions.
3. Code that no one is willing to invest the resources in to make any changes to keep compiling.
That intersection is quite small, if not empty entirely. Consider the programming language with the longest pedigree, Fortran, nearing its 70th birthday, with a large amount of code very firmly meeting the first criterion... and yet I don't think any modern compiler supports anything pre-Fortran 77.
If the latest compilers were to stop supporting old code then some new code could be cut off from further compiler upgrades as well.
This though highlights a grave problem. Rather than code which everybody understands (which could easily be re-written for a newer language) this is code which nobody understands. Replacing this code is in fact urgent.
‘When two implementations support the same notional feature with slightly differing semantics, should the committee use undefined behavior to resolve the conflict so no users have to change code, or should the committee push for well-defined behavior despite knowing it will break some users?’
I lurched towards the second option for ‘well-defined behavior’. And I would answer that way not just despite knowing breakage, but I would say that is the correct choice even in the event of large breakages for some percentage of active code bases. I have a hard time figuring out who would choose to accept more undefined behavior for new semantic constructs.
While I may not be accepting of the authors stance and conclusions to the question of ‘Trust the Programmer’, I do agree with his thoughts about the inherent positives of full throated argument in favor of increasing semantic constructs being ‘well-‘ and ‘implementation-‘ defined.
I may even go a step further and reach for an ultimate position that if it merely degrades performance metrics by some percentage, then instances of undefined behavior should be eliminated if at all possible. At a minimum, the goal should be to move to, at the most liberal, implementation-defined behavior for any given semantic construct. This would force compiler writers to specify and particularize what syntactic and semantic constructions they were taking advantage of to generate performance gains and allow developers the ability to decide if they could adhere to the implementation’s semantic guarantees.
C99 "broke" implicit declarations, but few if any people were forced to use C99 and it never became the default in, say, GCC (-std=gnu11 became the default in GCC 5.5, released in 2017).
The above is because I would hope that one result of pushing a nearly fully ‘defined’ (well or implementation) standard would be a strong interconnectedness and compositionality of semantics between all semantic constructions. This should mean compiler implementers can not just fall back on a mish-mash of standards compliance and then claim undefined behavior lets them just omit certain semantic constructs. I would like to think having the language be very clearly defined would almost require a complete adoption of some given standard to ensure the compiler was compliant.
I am aware the possibility of such a radical realignment of C’s structure is nearly impossible, but if C can not or will not do it, there may be the option for an incredibly similar language to piggy back it’s way to common use. This may also satisfy some of the arguments/positions in TFA concerning ‘Trust the Programmer’, where this superseding language can ‘unsafe/non-conforming’ out to C directly in C syntax in the event a non-conforming semantic construction is needed or desired by a developer.
That's a false dichotomy. You could also use implementation defined behaviour. Or specify that these n behaviours would be valid.
'Undefined behaviour' is too big of a sledgehammer.
> I think it’s perfectly reasonable to expect “trust me, I know what I’m doing” users to have to step outside of the language to accomplish their goals and use facilities like inline assembly or implementation extensions.
This is an extremely good point: if you can "trust the programmer" to write C code which exploits undefined behavior or other fundamentally unsafe compiler-dependent features, then you should be able to trust them to write inline assembly or a compiler extension to accomplish the same goal. If they can't, then they shouldn't be mucking around with undefined behavior in C: they might understand the behavior of the compiler at a "ChatGPT level" - as a set of ad hoc if-A-then-B's - but I wouldn't trust them to make serious decisions about state and security.
To get the "just define it all" approach you need people on weird ecosystems to be okay with paying for it. You think it is worth paying for (and I do too, frankly). But a significant portion of both the C and C++ communities don't - and that makes this very very hard.
I'm actually more confident in that process than I would be of some insane plan to rewrite it all in Rust - which would take forever and Rust doesn't have any of the tooling, formally proven compilers, a language specification, more than one implementation, etc etc.
People really don't know what they're talking about when they claim C codebases will be rewritten or will "have to be rewritten" because of "regulations", when the regulators are on top of this stuff already.
[0] - A meme from the days they used to call programming with straightjacket regarding Modula-2 and Object Pascal, in Usenet flamewars.
But NSA is already advising against the use of C: https://www.nsa.gov/Press-Room/News-Highlights/Article/Artic...
And the EU is working on new liability rules. Yes, Airbus might be able to get around this. I am more worried about smaller companies or products including open source.
There is already a bill in the US (it's part of a funding bill so it's just a matter of time until it passes) that, after it passes, the DoD is going to be putting out a plan for moving towards memory safety for things purchased by the DoD. Now, as we all know, because this isn't something that will happen overnight, and so there will be signoffs for exceptions, I'm sure, so we'll see what actually comes of it.
There is also a similar one in the EU, but I know less about how things work there so I won't say more other than "I know it exists."
[1] https://en.wikipedia.org/wiki/Ada_(programming_language)
I agree that this is a great story about how even a government mandate does not mean something may come to pass.
However, I also think that the conditions are different enough that it does not guarantee that this will fall the same fate. There's a few reasons why: the first is that the goals are different. Ada was created by the DoD, because they thought that there were too many languages in use there, and that standardizing on a single, modern language would be far better. So then they set off to create Ada, and in 1980, the first version was done. But it wasn't seen as popular, for various reasons.
In the words of the mandate itself: https://web.archive.org/web/20160304073005/http://archive.ad...
> In March, 1987, the Deputy Secretary of Defense mandated use of Ada in DOD weapons systems and strongly recommended it for other DOD applications. This mandate has stimulated the development of commercially-available Ada compilers and support tools that are fully responsive to almost all DOD requirements. However, there are still too many other languages being used in the DOD, and thus the cost benefits of Ada are being substantially delayed. Therefore, the Committee has included a new general provision, Section 8084, that enforces the DOD policy to make use of Ada mandatory.
It didn't catch on enough naturally, and so therefore needed a push, and the mandate was supposed to accomplish that. But it backfired. First of all, there were a LOT of exceptions. (which I mention could easily happen in this situation as well). The mandate:
> "Notwithstanding any other provisions of law, where cost effective, all Department of Defense software shall be written in the programming language Ada, in the absence of special exemption by an official designated by the Secretary of Defense."
That "where cost effective" was a big loophole.
Second, well, take this article from 1997: https://www.militaryaerospace.com/communications/article/167...
> Chief complaints about Ada since it first became a military-wide standard in 1983 centered on the perception among industry software engineers that DOD officials were "shoving Ada down our throats."
Part of the idea of repealing the mandate was that it would be more palatable to people, and that they'd be more likely to use it if it were repealed.
But there were a lot of other parts to this story that led to the removal of the mandate, like an overall movement towards more off-the-shelf commercial components rather than making everything in-house. But this comment is already too long.
---------------------------------
Okay so why is this different? Well, first of all, because it's not actually a mandate: the language in the bill is
> SEC. 1613. POLICY AND GUIDANCE ON MEMORY-SAFE SOFT- WARE PROGRAMMING. > > (a) POLICY AND GUIDANCE.—Not later than 270 days after the date of the enactment of this Act, the Secretary of Defense shall develop a Department of Defense wide policy and guidance in the form of a directive memorandum to implement the recommendations of the National Security Agency contained in the Software Memory Safety Cybersecurity Information Sheet published by the Agency in November, 2022, regarding memory-safe software programming languages and testing to identify memory-related vulnerabilities in software developed, acquired by, and used by the Department of Defense."
That sheet is this one: https://media.defense.gov/2022/Nov/10/2003112742/-1/-1/0/CSI...
and the most salient part
> NSA advises organizations to consider making a strategic shift from programming languages that provide little or no inherent memory protection, such as C/C++, to a memory safe language when possible.
"when possible" feels like that "where cost effective" bit. I guess I've said this twice in this comment now. We'll see.
But moreover, this is not recommending "rewrite everything in Rust." It is not recommending rewriting every single thing in any single language, or move towards a language that is new, designed by them. It goes on to mention
> Some examples of memory safe languages are C#, Go, Java, Ruby™, and Swift®.
I suspect that this policy will not be as controversial as broadly as the Ada mandate was. It's just a different thing.
Time will tell, I guess.
All that time, I really wondered what the flip they are on about. Adding extensions that really, I really raise eyebrows at while ignoring complete elephants in the room, like the above feature.
Last year they were pontificating about something something adding some sort of (complicated) syntax to support destructor functions and were wondering about prior implementations, and I had to point to them the support of destructor function had been there as an extension of gcc for countless years.
It's like they aren't actually using the language, just speccing it. Kinda like linux maintainers who are more like some sort of priesthood and gatekeepers than actual users of the system.
With complicated syntax you mean the "defer" proposal?.
In any case, different people have different priorities and ideas. The best way to influence decisions is to contribute to the standardization process. For example, everybody can submit proposals.
But regarding case range values, there is now one: https://www.open-std.org/jtc1/sc22/wg14/www/docs/n3194.htm
But calling things "stupid" is relatively useless internet noise and the best way to be ignored.
Just standardize attribute((cleanup)) as it is already widely implemented and used, and stop messing around with defer nonsense.
cleanup is unlikely to be standardized exactly as implemented as it would violate the rule that standard attributes can be removed from a correct program.
> it would violate the rule
Good, let's change that rule for this case.
I think changing the rule would create a mess, so I am not in favor of this. I wonder if there is a way so that existing macro wrappers could be adapted...
Today if you are using C it's because you HAVE to (embedded toolchain, historical code, etc). In that case, you use the C you find, not the C in the next version of the standard. So I find the C standard committee's work... uninteresting.
Sorry, guys.
The standard is just the minimal supported common language core of a family of C dialects, with the actually interesting stuff happening outside the standard in language extensions and 3rd-party libraries.
In a way we can "thank" Microsoft for that. If the MSVC team hadn't boycotted the C standard for nearly two decades, adhering to the C standard would have been more important. But since MSVC didn't support anything past C89 anyway, the C world moved on without them by switching to different compilers.
If anything, I was saddned that they backtraced on their decision to focus on C++.
Then again, nowadays thanks to the ongoing security bills, Azure business unit has a roadamp to only use Rust for new systems programming projects, and C#/Go for when managed languages are not an issue, while Visual C++ team seems mostly focused on keeping game developers happy and little else.
If a couple open source devs can easily out-perform the MSVC team, Visual Studio really is in trouble ;)
Hobbyist retro-computing compiler for 8-bit and embedded CPUs is hardly anything to worry about for Microsoft.