C23: A Slightly Better C
lemire.me
lemire.me
struct bla_t bla = {};
...instead of struct bla_t bla = {0};
This was one of those pointless differences between C and C++ which caused a lot of grief when writing code in the common C/C++ subset.PS: also func() now actually is the same as func(void), and unnamed parameters are allowed.
The eighth line uses the static_assert keyword, which is a feature of C++11
C11 already had _Static_assert() and static_assert() in assert.h> The idea behind static_assert is great. You run a check that has no impact on the performance of the software, and may even help it. It is cheap and it can catch nasty bugs. It is not new to C, but adopting the C++ syntax is a good idea.
Nah those are easy to name, just annoying, it's types which literally don't have an external / public name, like lambdas, or locally defined types.
#include <iostream>
auto createVoldemortType(int value) {
struct Voldemort {
int value;
};
return Voldemort{value};
}
int main() {
auto voldemort = createVoldemortType(7);
std::cout << voldemort.value << std::endl; // output: 7
}For many compilers, as soon as the parser sees a left curly brace, it pushes a symbol table onto a stack, and when it sees the corresponding right curly brace, it pops the symbol table off the stack, and "forgets" any declarations that were made inside that scope. That is so things like this work as expected.
{
int x = 0;
{
int x = 1;
printf("%d\n", x); // prints 1
}
printf("%d\n", x); // prints 0
}
But, in C++, using the auto keyword, declarations can escape their scope with auto. I'll change the C++ code that OP wrote. The C++ compiler has to correctly resolve cases like this, which means it can't just forget all the declarations within the scope of the function after the definition is done. #include <iostream>
auto createVoldemortType(int value) {
struct Voldemort {
int value;
};
return Voldemort{value};
}
struct Voldemort {
std::string value;
};
int main() {
auto voldemort = createVoldemortType(7);
std::cout << voldemort.value << std::endl; // output: 7
}At least that’s how I think the parsers work I‘m familiar with.
I've read a book about a BLISS compiler [1] that does this, but still uses a stack like I described [2]. It implements a hash table that used linked list nodes for collision. A new declaration adds a new name to the table, and it attaches the node to uses of the name in expressions of the syntax tree.
When a scope is exited, the declarations from that scope are removed from the symbol table, but because they're still attached to the syntax tree, they can't just be freed. They're added to a linked list of "purged" nodes, so that the information they contain can be used later during code generation, and then freed.
One-pass compilers don't have this problem; they really can just free the memory for reuse, because after they exit a scope, they've already generated the assembly or machine code from the high-level language.
However, I don't know what LLVM or GCC, or any other remotely modern compiler, does. I haven't read the code much.
[1]: https://en.wikipedia.org/wiki/The_Design_of_an_Optimizing_Co...
[2]: Actually, it intertwines the stack and the symbol table in a complicated way, so there's only one hash table, and multiple stacks within it. It's explained by a diagram they include on page 13. You can find a PDF of it here: https://kilthub.cmu.edu/articles/journal_contribution/The_de...
I don't see any code ther that does that. Is it implicitly passed?
I don't know C++, though I did know C somewhat well earlier.
return Voldemort{value};
I guess that does it.
Ha ha, that reminds me of that phrase of yore, "the quality without a name (qwan)" (google it), which was heavily bandied about years ago, during the heyday of C++ and the software patterns movement (which continued a lot in the days of Java, of course). James Coplien (IIRC) and others of that time come to mind.
https://en.m.wikipedia.org/wiki/The_Timeless_Way_of_Building
https://en.m.wikipedia.org/wiki/Pattern_language
https://en.m.wikipedia.org/wiki/Jim_Coplien
Though I read a fair amount about that stuff, a lot of of it went over my head, but later, I did understand some of the patterns, after reading the design patterns book, and trying out some of them.
The template method pattern is my favourite pattern, because I understand it more well than many of the others :), and also because it is the basis of software frameworks (inversion of control, aka the Hollywood principle - "don't call me, I'll call you"). Other patterns that I like and understand are the command pattern, the interpreter pattern, the chain of responsibility pattern, and flyweight pattern, to name a few. Builder and Factory, not so much. Singleton is straightforward, or is it really? impls matter :)
And I have written a few toy frameworks, which is fun to do and use.
I agree that you don't need auto pointers in C the way you do in C++. C++ type names can get so cumbersome...much easier to let the compiler figure it out for you.
#define SWAP(var1, var2) do { \
auto tmp = var1; \
var1 = var2; \
var2 = tmp; \
} while (0)
Previously, you would have needed a third macro for the type, or you would have needed to do a byteswap to be generic. const typeof(var1) tmp = var1;SWAP(…);
Consider this:
if (e)
SWAP(a, b);
else
something_else();
Currently, that expands to this, which is still valid code: if (e)
do {
// ...macro...
} while(0);
else
something_else();
If it were just wrapped in braces, the code would be parsed like this: // One-armed if-statement
if (e) {
// ...macro...
}
// Empty statement
;
// Another statement
else something_else();
It's incorrect syntax to start a statement with else, so the compiler will say something like "Unexpected 'else' at line N.""What's the best way to write a multi-statement macro?"
I think they obfuscate things unnecessarily. In C++ it's understandable because with templates you have a tendency to have massively complicated type names.
Everything else looks great to me though.
even the newest clang does not support many of the c23 features.
gcc13 is much better, only a very few c23 are yet to be supported.
this is a long standing problem with clang/clang++: they're used in many linters and intellisense but they're lagging behind by a lot comparing to gcc.
Pure speculation, but I wouldn't be surprised if these two were directly related to each other. Clang/LLVM is more modular, which makes it easy to write new things that integrate it (like a linter), but this can slow down adding new things that change the data model and external interfaces. GCC is the opposite: one big blob that makes it easier to add things since the surface area is lower.
Almost all major compiler vendors have migrated to clang forks and if they contribute upstream, is on the backend side, for their platforms, not fronted changes.
Big contributors like Apple and Google, deciding to refocus on their own languages, and current language support being good enough for their LLVM use cases.
- Remove Trigraphs.
- Remove K&R function definitions/declarations (with no information about the function arguments)
This is horrible, I loved using `and` and `not` and `or` in C boolean expressions and to pretend I'm writing Python code. It's fun!
In C, "trigraph" specifically refers to nine special sequences beginning with "??": https://en.wikipedia.org/wiki/Digraphs_and_trigraphs#C
They were used to provide an alternate way of typing punctuation characters for keyboards that don't have them (e.g. "??(" is "[").
#define and &&
#define not !
/* ... and so on */
In C++, they're actually keywords that are built into the language, so you don't need a header then.When people talk about trigraphs in C, they're talking about the trigraphs listed on this page: https://en.cppreference.com/w/c/language/operator_alternativ...
Not used C for many years, but used it a lot earlier.
And had read both the first and second editions of the k&r c book (ed. 1 pre-ansi, ed. 2 ansi).
based on that, iirc, this:
>Remove K&R function definitions / declarations
should actually be:
function definitions / declarations as in the first edition of the k&R C book.
In my previous comment, I was going to say (from memory), that in the second edition, i.e. in ANSI C, both types of declarations are allowed, old style and new style.
Also, if you just wrote the declaration (return value, then function name followed by types with arguments in parentheses), followed by a semicolon, without a function body in braces, it was called a function prototype.
MS C (as in, some version of Visual Studio's command line C compiler), had a flag to generate the prototypes from the function definitions. I had used it some. /Zg, possibly.
name (argument list, if any)
argument declarations, if any
{
declarations
statements
}
Is this disallowed in C23?C is not a subset of C++. The simplest example:
> int new = 1;
Perfectly valid C, but not valid C++.
Now, modules would obviously be a bigger break than a few keyword incompatibilies, but at that point you'd also massively increase complexity and basically start creating C++ again, which is probably the actual reason it hasn't happened yet.
And no modules wouldn't make C more like C++. It's just make C a bit saner and faster to compile. Because modules are complete and don't depend on code compiled before the import. Which means you can compile your modules exactly once and only once. Well written C code bases you could probably just change #include to #import and reap the benefits.
In other languages you need IDE support for extracting the public API of a module.
There's more information about the proposal at https://gustedt.wordpress.com/2022/01/15/a-defer-feature-usi...
gcc 23? That doesn't sound right
the language being small enough that all basics can be grasped in a day.
the language being complete that it will remain the same in 10 years.
while at the same time being memory safe.
btw I don't think golang applies given the bad ffi story in go.
--- edit btw:: yeah this implies the use of a GC. though it must not have massive pauses or stop the world GC.
There's also Rust, but that's a little harder.
[x] nicely interfaces with c
[x] the language being small enough that all basics can be grasped in a day
[x] the language being complete that it will remain the same in 10 years (the 1.0 release was 23 years ago)
[x] while at the same time being memory safe.
> [x] nicely interfaces with c
Unfortunately, that's only if you can avoid D Strings. Otherwise, you'll need to use toStringZ which makes copies of each string to ensure they have a null terminator.
> [x] the language being small enough that all basics can be grasped in a day
I'm still learning new stuff after several weeks using D. The basics are indeed simple, but there's A LOT of stuff in D.
> [x] the language being complete that it will remain the same in 10 years (the 1.0 release was 23 years ago)
D is evolving slowly but evolving. With the new ideas about making parts of the stdlib GC-free and the borrowing concepts being slowly introduced, the language is changing... people are making a lot of pressure to add new "cool features" from other languages, like the recently accepted string interpolation proposal. It will not be the same in 10 years, but it's true that most of it will be unchanged.
> [x] while at the same time being memory safe.
D is not memory safe by default, you need to use `@safe` which is annoying because currently , a lot of the stdlib is not `@safe` (but it could be!). It's true it's much, much harder to mess up in D than in C, but compared to Rust, I think it's quite unsafe (which is why it's introducing borrowing, to catch more unsafety bugs).
What features of D are deep?
A feature like string interpolation is modest syntax sugar. The kind of change that takes 40 seconds to understand.
> Rust
Bwhahahaha if D is not simple or simple, IDK what to call Rust.
There's nothing that would prevent them from using C strings. Since they want a safe language, I doubt this would be a reason they don't want to use D.
> I'm still learning new stuff after several weeks using D. The basics are indeed simple, but there's A LOT of stuff in D.
This is a bad question, because there really is no language in 2024 that you can learn everything in a day.
> D is evolving slowly but evolving.
Again, a bad question. Even C is evolving. The main complaint about D is that it isn't evolving fast enough, with too much emphasis on avoiding breaking changes. If they implement editions, it actually would work as OP wants, because code that compiles today will compile forever in the future.
In other words: Memory safe, easy to learn, easy C-FFI? Pick two.
Interestingly this approach is also somewhat popular in Rust to workaround borrow checker restrictions.
For instance see:
https://floooh.github.io/2018/06/17/handles-vs-pointers.html
[0] https://livebook.manning.com/book/nim-in-action/chapter-8/60
The C FFI Nim library lineage goes c2nim --> nimterop --> something i forgot --> futhark.
You can have all the memory safety in the world within the bounds of your own language, but it mostly gives you an illusion of security if the common pattern in the community is to just wrap C libraries. Having FFI be a bit of a hassle can actually go a long way towards shaping the community towards stronger memory safety.
One caveat is still being heavily developed.
lua? luajit's got great ffi.
If you meant a compiled language then you can write code that is memory safe in C. You can even run tools against your code to measure this in various ways.
All "memory safe" languages that compile to lowest level ISA code are just memory safe "by default" and all of them necessarily offer escape hatches that turn it all off.
No "memory safe" compiled languages offers memory protections beyond what the operating system provides. If it's within the memory space of the process you can access it without limitation. "Memory safety" can reduce your exploit surface but it can't eliminate it out of an incorrectly designed program.
The problem with your point is that it's single ended, because the scale and scope of deployed Rust or Go software has not matched that of C/C++ software. We also don't have particularly good data on how many memory safety bugs are in a project with good deployment controls versus ones that aren't.
We also don't know how well tested any of the vulnerable Microsoft code actually was and so I'd be wary of drawing any broad conclusions across languages from that simple statistic. It's also likely self reported and not likely to be rigorously gathered for this type of analysis.
The fact that there's such a large difference between C/C++ projects with respect to historically discovered vulnerabilities to me suggests that it can't be down to the language but how the project deployments are engineered.
You're one unnoticed checkin of an "unsafe" construct in any of these languages away from having the dreaded memory safety vulnerability introduce itself into your project. Even worse, you could have a crate that has an unsafe block you didn't previously call, but a new checkin now calls this extant and disregarded method. So, what do you do? The language hasn't done anything for you here. Use an analysis and/or fuzzing tool? So, how are we anywhere different because of the language?
Even for Go, a language I love quite a bit, if you forget to synchronize shared maps with simultaneous reads and writes you're in for a panic, and possibly real safety bugs. The GC and the fact that "unsafe.Pointer" are "slightly hard" to use isn't a huge attribute as it leaves entire classes of bugs on the floor with the tines pointed straight up.
So I suppose it's not literally just the Rust language itself, but given the context of Rust's development, there seems to be an intertwined culture that was more likely to arise than not.
Gambit is also reputed as the second fastest Scheme compiler out there; only Chez Scheme produces faster code.
Zig is aiming for what you're talking about, but it's not yet stable. They're interested in memory safety, but they don't want to add a borrow checker and they certainly don't want to add a GC, so my outsider guess is that they'll end up tolerating memory unsafety and trying to make up for that with debug modes and tooling. I don't really have a sense of what the end result will feel like, and the problems discussed in https://youtu.be/dEIsJPpCZYg make it sound like the basic semantics still have a ways to go.
Go could've been the language you're talking about if they'd given up on goroutines and just used the C stack, but goroutines are arguably the most important feature in the entire language, and it's not clear to me that there's a market for "Go but worse for network services and better for FFI". It would be hard to carve out a niche as a systems programming language that's great at making syscalls but can't realistically implement a syscall.
I vigorously deny this claim :-)
There's an easy-to-learn and powerful language implementation, which I mentioned above, that has seamless interop with C (to the point that you can include snippets of actual C code in the source code).
memory safety doesn't mean just one thing, but probably it requires either a lot of rust-like features, a tracing garbage collector, or automatic reference counting.
the language being small enough that all basics can be grasped in a day
that disqualifies taking the rust-like path.
able to use all c libraries through ffi. without loss of performance or extra fu.
that disqualifies most (all?) advanced tracing gc strategies
it must not have massive pauses or stop the world GC.
that disqualifies simpler tracing gc strategies
depending on what precisely you're looking for, it's possible it might be a pipe dream. but it's also possible you'll find what you want in one of D, Nim or Swift. Swift is probably the closest to what you want on technical merit, but obviously extremely tied to Apple. D and Nim strap you with their particular flavor of tracing gc, which may or may not be suited to your needs.
> the language being small enough that all basics can be grasped in a day.
That would be Zig. You can directly import C headers and compile C code with the Zig compiler and also cross-compile C code without requiring a separate compiler toolchain.
Currently it's also possible to compile C++ and ObjC (but not import C++ or ObjC headers), but that functionality will probably be delegated to a separate Clang toolchain in the future.
> the language being complete that it will remain the same in 10 years.
...that will take a while (also depending on whether you consider the stdlib part of the language or not).
> while at the same time being memory safe.
...that wouldn't be Zig then ;) (TBF, Zig is much stricter than C or C++, which helps to avoid some typical memory corruption problems in C or C++, but it's by far not as watertight as Rust when it comes to static memory safety - Zig does have a couple of runtime checks though, like array bounds checks - dangling pointers are still a problem though and are only caught at runtime via a special allocator.
Woha! Well, maybe... Ada?
https://learn.adacore.com/courses/intro-to-ada/chapters/inte...
Edit: Also not yet mentioned: Julia: https://docs.julialang.org/en/v1/manual/calling-c-and-fortra...
Or guile: https://www.gnu.org/software/guile/manual/html_node/Dynamic-...
... Or ruby! (But by now we're solidly in the land of "all languages connect with C"):
Yes. https://ecl.common-lisp.dev/static/files/manual/current-manu...
> without loss of performance or extra fu.
Maybe a small loss, compared with fine-tuned C.
> the language being small enough that all basics can be grasped in a day.
The basics, certainly. You can learn the basics of Lisp syntax in about twenty minutes, the basics of looping, conditionals, datatype definitions, function definitions and all basic stuff in about a day.
> the language being complete that it will remain the same in 10 years.
Mostly, yes. ECL has mostly been the same for the previous 20 years. I see no reason that it would change substantially in the next 20.
> while at the same time being memory safe.
Caveats apply here, due to how deeply ECL can hook into C code. Even if you're doing weird things in the C code, it's unlikely you'd accidentally run into problems.
> btw I don't think golang applies given the bad ffi story in go.
What bad ffi story? I've never tried to use the FFI in go, but I haven't heard particularly bad things about it. Of course, that could be because most Go programmers aren't using the FFI anyway.
What was wrong with "decltype"?
Seems utterly bizarre given that the rest is verbatim copypasta from C++, even the attribute syntax that sticks out like a sore thumb in C, and also given that auto and decl.. I mean, typeof, serve no useful purpose in C.
People who casually write C don't really care for C23 since it fixes all the wrong things. Nobody really wanted C+ (i.e. something slightly closer to C++ than before) which is basically all this standard achieves.
Nobody wants that ? Everybody who knows it exist wants that.
Thank you for the correction
So much so that new revisions of C are just backporting features at this point.
All kinds of major projects switched to C++ already, for example GCC. All the major new projects, for example LLVM, are also in C++ from the get-go.
Even the Linux kernel is considering switching to C++ now.
Source?
considering doesn't mean it will happen in near future, because toolchains/ecosystem is not there, the same is applicable for many other projects.
Thankfully the Linux kernel development relies on testing instead.
the only sense in which the relationship between C & C++ is like that of a cybertruck to a bicycle is the one in which the latter are "transportation devices" and the former are "programming languages". there is no particular feature of a bicycle represented by a cybertruck other than "it gets you somewhere".
Maybe the c++ killer will turn out to be just some future, modern version of C...