C’s Biggest Mistake (2009)
digitalmars.com
digitalmars.com
C is finished if it doesn't address the buffer overflow problem, and this proposal is a simple, easy, backwards compatible way to do it. It is simply too expensive to deal with buffer overflow bugs anymore.
This one addition will revolutionize C programming like adding function prototypes did.
if C dies then what replaces it?
Are you saying both are doomed? Or is there some scenario where C++ survives without C?
C++ is actually in a slightly better spot ironically because it’s harder to integrate with. If you have a C program you can pretty easily start replacing parts with Rust. You can’t do the same with C++ which insulates it better in that sense.
So not in my lifetime.
People have been looking at that for years now. I’m convinced it’s going to happen one day. It might not be Rust, but it’s going to happen that we will have different models for writing these kinds of things
https://thenewstack.io/linus-torvalds-on-diversity-longevity...
Thanks to LLVM and GCC you can happily write embedded code in a higher level language, but the vendors don't bother supporting it because a lot of embedded coding isn't really what we would call software (no tests etc.)
I don't think anyone was going to write their fridge's code in Haskell anyway.
I've been seeing those exact words for decades now, and C is still going strong. Every few year a new language comes, somes writes something in it, that was written in C before, someone might even write a basic OS in it, and after a few years, that language is almost forgotten, a new one is here, and again, someone is writing something in it, but in the end, we still use C for the things we used it 10, 20, for some, even 30 years ago.
This is where C is superior to virtually every other language. It has K&R to start with [1], a wealth of examples to progress from there, man pages, autotools, cmake, static and shared libraries.
> good standard library.
It should have hash tables at least, but it isn't bad.
[1] Which is still the best language book ever written (yes, it has some anti patterns, you unlearn them quickly).
> Those functions are completely non-standard and should be avoided in portable programs.
If you want I could have brought up hcreate, hdestroy, and hsearch:
> The functions hcreate(), hsearch(), and hdestroy() are from SVr4, and are described in POSIX.1-2001 and POSIX.1-2008.
Happier?
Personally I do not mind using libraries typically installed by the Linux distribution's package manager anyways.
If the question is whether or not I think the C standard library could be improved, then yes, I would say it could, but I do not want it to have a hash table and all sorts of stuff like that, because there are lots and lots of ways to implement them, and they might not suit my needs. C is great, because you can build it from the ground up (if you want to) to make it specifically for your use case. It gives you the building blocks. I believe I have a comment regarding this somewhere, that I like C because it does not implement stuff for you that is in some ways "generalized", which is often a bad thing. This is my problem with "it should have hash tables at least". You cannot implement it in such a way that it suits everyone's needs.
As far as Rust goes, yes, I do not like that crates are full of one-liners, and so forth. It shares the same problems that npm has. I ran cargo build on many Rust projects before. No way.
As for build systems, autotools is a hilarious clown car of a disaster. You write a script (automake) to generate a huge, slow script (configure) to generate a makefile to finally invoke gcc? It is shockingly convoluted. It seems more like a code generation art project than something people should use. CMake papers over it about as well as it can, but I think cmake is (maybe by necessity) more complex than some other entire programming languages. In comparison, in rust “cargo build” will build my project correctly on any platform, any time, with usually no effort on my part beyond writing the names and versions of my dependencies in a file.
And as for package management, C is stuck in the 80s. It limps by, but it doesn’t have a package manager as we know them today. There’s no cargo, gems, npm, etc equivalent. Apt is no solution if you want your software to work on multiple distros (which all have their own ideas about versioning). Let alone writing software that builds on windows, Mac and Linux.
So no, C is not superior to other modern languages in its docs, build system or package manager. It is vastly inferior. I still love it. But we’ve gotten much, much better at making tooling in the last few decades. And sadly that innovation hasn’t been ported back to C.
I've got a parts drawer full of controllers that says it won't.
When it comes to microcontrollers rust is currently at the mercy of LLVM support and vendors.
> There are only two kinds of languages: the ones people complain about and the ones nobody uses.
This unfortunately seems to mostly hold true.
"Porque" means Because. "Por que" means Why.
Just use nix or even apt. Both of them are MUCH better when compared to trash like npm or cargo which do not even check for signatures.
> common build system
Such as make? There is also Ninja/Meson if you prefer.
C survives.
Meanwhile C is running strong since the 70s.
> the lack of package manager
What do you call linux distro's package managers then? I mean, in distributions like Debian you can even download a package's source code with apt-get.
If you want to count them as package managers, they're by far the worst ones of all the well known languages (with some notable exceptions e.g. guix's and nixos's).
They're not portable between distributions or even different versions of the same distribution (!), since it's non-trivial to install older versions of libraries (or, hell, different versions of the same library at the same time). Not to mention that it's a very manual and tedious process in comparison to all the other language specific package manager. 'Dependency hell' is a problem virtually limited to distro package managers (and languages like C and C++ that depend on them).
Getting older, unmaintained C programs to run on Linux is an incredibly frustrating experience and I think a perfect demonstration of how the current distro package manager approach is wholly insufficient.
The have the only feature I care about: cross-language dependency management.
Unless you are suggesting to reimplement everything in each language and then make users install ten different XML parser, SSL implementations, etc. just because not-implemented-in-my-favorite-language syndrome.
You see, my OS already comes with those, and I expect to use them. I have the Debian package system: dpkg, apt, aptitude, and so on. It's a big mess when other software tries to steal that role. I have the traditional build systems and more: make, cmake, autoconf, scons, and so on. If I'm building a large project with multiple languages, I'm going to use one of those tools. If a language wants to fight me on that, I'm not interested in that language.
C's function prototypes, syntactic sugar added circa 1990, were transformative for C programming.
https://www.cl.cam.ac.uk/research/security/ctsrd/cheri/
It is fundamentally like the 80286 segments, but with all sorts of usability troubles solved. The 80286 segments were impractical because there were a small number available and because the OS couldn't safely hand over direct control. Every little segment adjustment required calling the OS.
> Google is committed to supporting MTE throughout the Android software stack. We are working with select Arm System On Chip (SoC) partners to test MTE support and look forward to wider deployment of MTE in the Android software and hardware ecosystem. Based on the current data points, MTE provides tremendous benefits at acceptable performance costs. We are considering MTE as a possible foundational requirement for certain tiers of Android devices.
https://security.googleblog.com/2019/08/adopting-arm-memory-...
> Starting in Android 11, for 64-bit processes, all heap allocations have an implementation defined tag set in the top byte of the pointer on devices with kernel support for ARM Top-byte Ignore (TBI). Any application that modifies this tag is terminated when the tag is checked during deallocation. This is necessary for future hardware with ARM Memory Tagging Extension (MTE) support.
>....
> This will disable the Pointer Tagging feature for your application. Please note that this does not address the underlying code health problem. This escape hatch will disappear in future versions of Android, because issues of this nature will be incompatible with MTE
https://source.android.com/devices/tech/debug/tagged-pointer...
So unless you have other official feedback from Google management, I will keep repeating myself.
I agree 100%. This also reminds me of this article:
https://nibblestew.blogspot.com/2020/03/its-not-what-program...
HN discussion: https://news.ycombinator.com/item?id=22696229
Promises/async functions in JS and C# do absolutely nothing that you couldn't do without them. But they've had a structural effect on the average developer's ability to write scalable code.
In some ways it would have worked much better. Header files can easily be wrong. The object files being linked are what really matter.
So, suppose we decided to implement this today, on a GNU toolchain. At the call site, we'd determine the parameter types based on the conventional promotions. This info gets put into the assembly output with an assembler directive. The assembler sees that, then encodes it in an ELF section. It might get a new section called ".calltype" or it is an extension to the symbol table or it involves DWARF. Similar information is produced for the function body. The linker comes along, compares the two, and accepts or rejects as appropriate.
It also isn't a requirement that C++ use mangled names. Other ways of carrying the type information are possible. I like the idea of a reference to DWARF debug info, which C++ is already using to support stack unwinding for exceptions.
Overloading requires the type information to be part of the symbol "name" (ie, whatever is used for symbol lookup and linking) wether that is a mangled string or more complex data structure.
If you're worried about performance test it and turn it off.
Just terminating the program won't be much better than OOB memory access in many cases.
Continuing but discarding OOB writes/use a dummy for OOB reads could lead to much worse behavior.
An exception (or setting errno since this is C) would need that exception to be handled somewhere in a sensible way in which case you could just as easily add manual bounds checking.
>extern void foo(size_t dim, char *a);
And the like assumes that I have the space to waste a native type on every array. So if I’m using a 10 length array, I need to provision a native 32 or 64 bit value for “10”.
In embedded system this wouldn’t happen. At least not mine, I’m running up on limits all over the place even being careful with bitfields and appropriately sized types.
He’s right of course that foo(array[]) is converted to a pointer but that’s why I think you should always use array as a pointer so YOU know not to rely on its automatic protections.
I get the point; but I just don’t see C making this change.
So, I go back to the idea that it seems unlikely this would ever be an official C change.
Currently, f(a[]) with declaration void f(int a[]) passes a pointer to the first element, with no additional overhead.
Under the proposal, f(a[]) with declaration void f(int a[]) passes a pointer to the first element, with no additional overhead.
Help me understand why "other people would care"? What is the negative impact on someone who would not use the a[..] functionality?
Yes, it's not a lot of effort to manually add a size_t argument. But it is far too tedious and error-prone to expect a programmer to add all the bounds checks. Being able to effortlessly tell the compiler "please check for me" is the huge win.
The second huge win is that the array is type-checked. So if you pass it to another function, the compiler enforces that it must again be passed with the size included. You don't get that by manually adding a size argument.
So those programmers might just use the sugared approach and avoid the problem of writing past the end of an array, without ever knowing how tedious and/or difficult debugging such problems can be. They might sort of never even realize that they dodged a bullet simply due to some sugar.
How do you see it?
More specifically, at the moment, it evaluates to the size of the pointer itself, which is useless. On the other hand;
static void
foo(int a[..])
{
for (size_t i = 0; i < (sizeof a / sizeof int); i++)
{
// ...
}
}
... would be very useful, as it's the same syntax you can already use inside the function where the array is declared, which makes refactoring code into separate functions easier, as you don't have to replace instances of sizeof with your new size_t parameter name.The only thing I'd like to see is compatibility with the static keyword; so that you can declare it as a sized-array but still indicate a compile-time minimum number of array elements. At the moment, in C99, this does not compile without serious diagnostics which would immediately highlight the problem:
#include <stdio.h>
static void
foo(int a[static 4])
{
for (size_t i = 0; i < 4; i++)
printf("%d\n", a[i]);
}
int
main(void)
{
int a[] = { 1, 2, 3 };
foo(a); // Passing an array with 3 elements to a function that requires at least 4 elements
foo(NULL); // Passing no array to a function that requires an array with at least 4 elements
return 0;
}
demo.c:14:3: warning: array argument is too small; contains 3 elements, callee requires at least 4 [-Warray-bounds]
demo.c:15:3: warning: null passed to a callee that requires a non-null argument [-Wnonnull]Granted, it's not static analysis, but it should catch most aliasing related errors, no?
Edit: But to be clear, C really ought to have first class arrays. If you truly don't want bounds checks in a specific scenario for some arcane reason, you could still explicitly pass a raw pointer and index on that. (The same as you would in any sane systems language.)
Note that I'm using "platform" to refer to the CPU instruction set.
A good example of this is calling functions with the wrong parameter types. UB in C, but practically allowed by every compiler. No machine would care if you do this... until WASM came along and suddenly every function call is checked at module instantiation time for exactly this behavior. This is because all WASM embedders are fundamentally optimizing compilers. And what is the mother of all optimizations? Inlining: the process of copypasting code from a function into wherever it is called. If a function is being called with the wrong arguments, how do you practically do that? You can't.
It is meaningless to talk about UB without also talking about optimizations. If you do not optimize code, then you do not have UB. You have behavior that is defined by something - if not the language spec, then the implementation of that spec, or a particular version of a compiler. There are plenty of systems with undocumented behavior that is nonetheless still defined, deterministic, and accessible. Saying that something is UB goes one step beyond that: it is saying that regardless of your mental model of the underlying machine, the language does not work that way, and the optimizer is free to delete or misinterpret any code that relies on UB.
That's what it used to mean. But at some points compilers people decided that since UB means literally "anything can happen" they can make optimizers optimize the shit out of the code assuming that UB can't be there.
C code that used to work 20 years ago, because the UB in it resulted in some weird but non-catastrophic behavior, doesn't work at all compiled modern compilers.
Other commenters already responded to this, but I thought I'd link an article I came across a while back that gives a concrete and easy to understand example of how UB can be leveraged for optimization by modern compilers. (https://devblogs.microsoft.com/oldnewthing/20140627-00/?p=63...)
Sorry? I’m unsure what you mean here, because there are plenty of ways to use globals in ways I would call “safe”: no undefined behavior, correct output, …
int x;
and in another: int* x;
and you'll be mixing pointers and ints and it won't be detected. int x = 7;
void f() { /* do things using x */ }
void insidious_function(int *p) { *p = 3; }
now, inside f you cannot be sure that x equals 7, even if you never write into it. You may call some functions, that in turn call the insidious function that receives the address of x as a parameter. There's no way to be sure that the value of x is not changed, just by looking at your code.Any production-quality C code will already use a (pointer + count) combo when passing arrays to a function, which is something that will still be needed under your proposal because the vast majority of arrays is dynamically sized. So unless all arrays in C are given the fat pointer treatment, I don't really see how what you suggest would make much of a difference. That is, if fat pointers are made the first class language construct, then, yes, that can be useful... though I disagree if it's not done, it will cause a demise of C.
Any or some? I'm not sure if I've seen that in the wild.
unless you have a team of incredibly diligent coders, people are going to read past the end of bare arrays over and over again. one specific mistake I keep seeing is where people misinterpret the meaning of a variable named `size`. is it the number of elements or the size in bytes? who knows, but it's probably UB either way if you're wrong.
I don't code c full time, but it's what I have always done when needing to use c via ffi to get a speed up in a dynamic language.
the problem with this approach is that you are still relying on the application programmer to provide the correct size at the beginning and not to mess it up by directly accessing the struct member later. private/public does not really exist in c, so it is a lot harder to enforce invariants within an object. the library could make the struct layout a private implementation detail (ie, not fully define the struct in the header provided to the client and take a pointer to the struct as arguments in the API) to at least discourage this. you could combine this approach with a my_array_struct_init function that returns a pointer to an empty array object. this is a common approach taken in c libraries (eg, libcurl) where the author really doesn't want you messing with their structs.
- Statically check the code, since the static analysis tool knows for certain which value is the size and can check that you're using correctly.
- Initialize the size correctly, since you don't have to enter it twice, or more crucially, remember to change it twice (or create a #define in another part of the file, name it, and document it)
You also make an excellent point yourself about the meaning of 'size'. If this was standardized, it would be the same everywhere, minimizing the risk of ambiguity.
I don't interpret Walter's suggestion that way. Of course, I might be wrong. Since the compiler must know the size of the array at the time it's declared, my thinking is that the compiler is smart enough to pass the size without the programmer having to even think about it.
Quite right. I use, and highly recommend, the convention that `size` is for number of bytes, `length` is for number of elements, and `capacity` for the allocated number of elements.
assert(length * sizeof(element) == size);
assert(length <= capacity);You should keep making this prediction ... one day you might be right! :)
One of the key areas C is used that Rust cannot be used easily is in limited embedded devices. That looks like it'll be the case for at least 10 years and probably much longer than that.
The main limitation is going to be if your MCU is supported by LLVM. If you're targetting ARM, RISC-V, MSP430 or Xtensa you might get further than you'd expect.
There's a hell of a lot of tooling built around C though. So nomatter what happens I don't think we'll ever be rid of it.
A viable C replacement is Ada, although it is not for people who dislike the "code is documentation" bit.
Maybe he meant outside of the education system. I think the reason for that would be hype and peer pressure, and the feeling of novelty, with a hint of FOMO. I do not see any languages being pushed/hyped as hard as Rust.
That is a mighty good showing for C.
(We learned C++ in university in New York which was basically C with occasional help from C++'s standard library).
For every young programmer learning Rust there are probably 10000 learning C.
C is still the only language you can count on anyone with a programming-related education to have knowledge of. (That doesn't translate into being able to program is C, but still.)
not saying you're necessarily wrong (c is definitely simpler than rust), but I think most people would write a similar comment for whatever language they feel most comfortable with. I write c++ most days. even though it's the most verbose and complicated language I've ever used, I can still probably get stuff done way faster than in another language I happen to pick up.
if I'm just throwing something together really fast, I do mostly use the old school C functions though. scanf is way nicer than streams.
I don’t go out of my way to teach or suggest Perl to my younger colleagues. I don’t know why that is. They don’t usually reach for Python, which I think would probably be their best choice.
Maybe I’m just a curmudgeon ... and by the way, can you please stay off my lawn?
:)
Nothing wrong with learning multiple languages, of course. C was my first professional language, and I spend most of my days in rust now. No shortage of rust devs who are big on C too.
But I was responding to "C code is being replaced by Rust fast.".
The syntax is also currently REALLY unstable. As in: The hello world has changed almost every major version. Hopefully that too will be squashed with 1.0
To be fair though, Zig is probably the lest egregious and most flexible of the modern "C killers". I can see it's really trying to innovate low level programming. I really like it's flexible malloc systems and support for dynamic linking at runtime. It's compile time code execution is excellent too. The fact that they're actually trying to support obscure platforms like the z80 is a good indicator that they're staying true to C's "code anywhere" mantra. That's why I'm mostly focusing on linting issues of all things.
I think Rust has been very quickly fading into obscurity. What Rust hast brought to the tables was nearly the same what was brought by 100+ other programming languages in attempts to "fix C."
> What Rust hast brought to the tables was nearly the same what was brought by 100+ other programming languages in attempts to "fix C."
> Oh... can you show me those 100+ other languages that have opt-out memory safety, explicit lifetime annotation, a borrow checker and no runtime?
Maybe I can make it even clearer:
> What Rust has brought to the tables was nearly the same as 100+ other languages
> Oh... can you show me those 100+ other languages that have (some unique Rust features)
Because Google wants people to search using the Google search engine, and see ads while they're at it. That would happen less often if Chrome were properly capable of searching the history.
Opera did full text history search a decade ago, but that browser doesn't exist anymore.
Also, history is expected not to be tied to the device in the cloud era. What more history data could you scalably sync other than URLs and page titles?
It also helps that C's standardization proceeds in ways that feel somewhat between sabotage and utter neglect.
Meanwhile, C is still the absolute best binary interop language devised by mankind.
This is not a random pet peeve, and WalterBright is as far as you can get from someone "who never wrote a line of code in C". This is the cause of numerous security bugs in the past and currently, and the reason most C material written in the 70s/80s is unsafe to be used today (mostly due to usage of strlen/etc vs strnlen/etc).
Fair warning - once you get accustomed to DasBetterC, you're not likely to want to go back to C :-)
I've been meaning to experiment with DasBetterC for a while, and I have a project C I've been wanting to migrate to something with proper strings (it's an converter for some binary file formats, but now I want it to import some obscure text formats too). Maybe that's the push I needed :)
After 20 minutes and about 250 out of 2098 lines converted, the error messages are very good and give very nice hints about what to change, I must say I prefer them to Rust's verbose messages.
DasBetterC's trial-by-fire was when I used it to convert DMD's backend from C to D.
I'm sure you already know this, but the trick to translating is to resist the urge to refactor and fix bugs while you're at it. Convert files one at a time, and after each run the test suite.
Only after it's all converted and passing the test suite can refactoring and bug fixing be considered.
C will still be used long after you and I and everyone here have returned to dust.
there are also people still riding horses. does not make it relevant in any way.
This is one of HN's comment guidelines. If you're not sure that someone is who you think they are, you can just ask, e.g.: "Hey, are you Walter Bright who did X and Y?"
Is it a rule at HN that you can't take someone else's name? Otherwise, there's no guarantee that you're talking to the "Real" Walter Bright...
... or that you're talking to that Walter Bright, come to think of it.
Nothing about Walter Bright in this statement, but some of the harshest criticisms from others I have seen of C are not from expert practitioners in C.
People who are experts and also critics seem to have a more practical, realistic, nuanced critique, that understands history and challenges to adoption, admits that the long history and difficulty of replacing C isn't exactly for no reason.
I also agree that what the standard committee has been doing for the last 20 years amounts to willful sabotage.
What about between c99 and c18? Is there anything you can think of? I think the _s() functions, advertised as security features, are a weak effort. Anything else come to mind?
Also the amount of UB descriptions just increased and are now well over 200.
Annex K was badly managed, a weak effort as you say, given that pointer and size were still handled separately, and in the end instead of coming up with a better design with struct based handles, like sds, everything was dropped.
ISO C drafts are freely available, I recommend everyone that thinks that they know C out of some book, or have only read K&R book, to actually read them.
But were they expert practitioners of C in the past? My experience is that most of the harshest criticisms of C come from former C experts who moved on to other languages because it became clear to them that C would never be fixed - Walter Bright included.
My point was that people were mistaking the comment for an attack on you, which I don't think was necessarily intended or needs to be without it being a valid point about a different set of critics.
You're mistaking the “C” ABI with the C language. The so-called C ABI should actually be called the UNIX-derived ABI, as (i) C doesn't define an ABI and (ii) C can perfectly produce binaries using another ABI (such as e.g. the “Pascal” one, common on the DOS platform).
By what measure would C be losing ground?
When we start seeing Rust or similar being used instead of C would be a good metric and / or major OS development.
However, the word "isoteric" is more correctly spelled (in non-phonetic spelling) as esoteric. The prefix "eso-" means "inside" in Greek, as in "esothermic", or "esophagus". The prefix "iso-" means "equal", as in "isomorphism", "isosceles", "isometric", etc.
The question is are people using C to program these hardware more? .. or are people gravitating towards safer compiled languages (Rust?). That's a valid question, even if the answer is "no C's usage is only increasing."
Nobody would do a project in C today with Go/Rust/C++ etc unless it's for a very specific situation
It's the one they know.
A C programmer isn't automatically going to switch to Rust for new projects that they would use C for, unless their goal is to use Rust.
And the history of Linux kernel and Android might come to an end if Zirkon ever replaces it, then it will be no more C on Android.
Android and Fuchsia source code are public available, just as their code reviews and ongoing work to port ART to Fuchsia.
Anyone skilful to use C without creating CVEs of their own, is surely able to find that information.
More usefully, if you want people to use D instead, what is stopping them, what reasons do they give? How can these be mitigated, cos while I like C I sure would like something better.
This seems like just another gimmick to push D tbh.
1. the judgement of people in the industry, who will always be biased
2. the judgement of people not in the industry, who don't know what they're talking about
As for me, I still sell a C and C++ compiler https://www.digitalmars.com/shop.html
Is this really "simple, easy, backwards compatible"?
I think Rust kind of counterexamples this.
While Rust can throw around slices [] (effectively runtime length), throwing around [u8; 8] and [u8; 9] (compile time length) to the same function gets nasty.
Perhaps all the constexpr work in Rust will make this a lot easier.
I’ve been pushing for this feature for many years and have been playing with it since it first landed in nightly. It works quite well.
My point was that: if you dump slices/fat pointers into C, how does it help?
Slices/fat pointers and the corresponding checks are a runtime thing, and that's absolutely anathema to a lot of C programmers. If it's not anathema to the C programmer, they probably aren't in C anyway.
So, now you need a way so that compiled slices/fat pointers mean something, and I'm not convinced that doesn't have a lot of ramifications that are being glossed over.
Most people using D leave the checks on.
Even without the checks, however, I can vouch that implicitly carrying around the length of the array with the array pointer is a vast improvement in the clarity of the written code.
The first thing you do to see if a new processor is ready for prime time is to port a c compiler for it.
The people who talk about rust replacing c are the type of people who thing that GTK is a reasonable C project.
"a pair consisting of a pointer to the start of the array, and a size_t of the array dimension"
No, that still doesn't fix the ABI. It's syntactic sugar. It is most definitely not passing an array.
Passing an array means exactly that, no more and no less. For example, suppose this is your array:
double foo[100][100];
The size is 80000 bytes. That is exactly how much data needs to be copied onto the stack, no more and no less.Getting the array dimensions is secondary. It would be nice to have them work. They could automatically get names. They could get size checks, so a function might declare itself compatible with a certain range of sizes. That's all a bonus, of much lower importance than the actual ability to pass an array.
The inability to pass an array impacts numerous other languages because they use the C ABI. If you can't put those 80000 bytes on the stack in C, then you can't do it in any language. The whole software ecosystem is thus impoverished.
Take the address, and pass a pointer, if that is what you want to do.
Maybe I want the callee to be able to modify the array without affecting the caller. Maybe I'm even telling the linker to put that array in ROM, but I want a writable copy in the callee.
Whatever... I have my reasons. The language shouldn't block me.
I'm the kind of person who optimizes with assembly, counts cache misses, counts TLB misses, and pays attention to pipeline stalls. I definitely understand the performance implications, and I definitely wouldn't be passing arrays around all the time.
That said, I want the ability. I want the language to let me do what I want, and on rare occasions I want to pass an array. Let me pass an array.
We can have giant structs. I've seen some over a megabyte in size. The default is that the callee gets a copy. (depending on the ABI it could be in the "wrong" stack frame, but it is a distinct copy)
Are we having huge problems with structs being passed by value? I don't think so. Normal people pass pointers, except when they actually want to pass by value. It works fine.
I have frequently seen beginners struggle with C arrays and pointers. Part of the trouble is that you can't pass an array. You can try, but the compiler quietly substitutes different code. It's a source of confusion, generating incorrect mental models of what is going on.
But all of this is way off-topic from the OP: the original point was, "Passing a pointer and length separately is error-prone; there should be a way to easily package the two together and this should be the default pattern for 90% of cases." Then you came in and said "No, instead C should support this totally orthogonal side-case that's a bad idea 90% of the time but has some niche uses." It's not a bad suggestion in itself, necessarily, but it's totally unrelated to the original proposal, much less an alternative to it.
"the inability to pass an array to a function as an array, even if it is declared to be an array. C will silently convert the array to be a pointer, and will rewrite the function declaration so it is semantically a pointer"
...and later, referring to the new syntax:
"an array is passed"
In no way is it so. It has nothing to do with passing arrays. It's passing a fat pointer, which is different.
> While [the first edition of K&R] foreshadowed the newer approach to structures, only after it was published did the language support assigning them, passing them to and from functions, [...]
Doing this frequently in any application means your profile will have a lot of memcpy in it.
I believe the same would be true if the language allowed passing arrays. Programmers would not generally pass them around. There is no need for the language to protect us from this by failing to implement the ability to pass arrays.
For example: If you had just put the length before the array buffer, you could've saved a stall in almost every use. That's a problem with out-of-order processing that's hard to fix. Maybe your compiler will get sufficiently smart, or maybe CPUs will collude with the memory controller (or something else amazing will happen), but those things are really hard. However we fixed it ourselves; we didn't need anyone to do it for us, because (due to laziness or luck) C gave us enough of the tools we needed to do what we needed to do.
I think that's a bigger deal than buffer overflows, as unpopular an opinion as that is.
Assuming that it is so: when? It seems that C—despite all its shortcomings—remains a very popular language in some problem domains. For some platforms it seems like it's really the only performant HLL option.
I have been a member of many programming languages communities, and every language has its own culture. C was always a language for the arrogant. "The real programmers" that can handle their memory, not afraid to work with pointers and that can get their code right.
I've been there, done that for many years, and became more humble with time. In a way I still love the brutal simplicity and low-level nature of C, but I would use it only if absolutely can't use any other language for technical reasons, and I would be really, really cautious.
With all due respect, but if someone says something invalid, the fact that they have authority on a subject does not mean that we should agree.
As far as I understand the article (And I'm not the great Walter Bright, so I may be wrong) - the author states that "void foo(char a[..])" is better syntax than "void foo(size_t s, char a[])" but does not provide any arguments for it. Furthermore, the author initially fails to mention that there has been an attempt to fix the array-to-pointer-decay issue, when discussing "C's Biggest Mistake".
So, yeah, the author may be right that this has been C's biggest mistake. I don't know whether that is true or not, I do not have his experience. It is certainly true that this mistake would be high on rankings of all mistakes that C did. Still, the initial "sleight of hand" move followed by unsubstantiated argument leads to a post with the quality similar to that of a twitter post. Maybe even worse, since, you know, it's posted on a place other than twitter, so we are actually talking about it as if it was something serious.
in reality, most production build systems at least use -Wall (or their compiler's equivalent) and possibly also have a list of specific warnings turned on/off for different parts of the code. it would be nice to have some saner defaults, but it just doesn't matter that much.
I have no idea if this is the "right" way to do things, but it seems to work.
Back when I was doing pure Windows C for a while, this helped quite a bit,
Basically you had to #include<windowsx.h> and define the STRICT macro.
There were also several utilities that made it much easier to deal with events, dialogs and callbacks.
https://docs.microsoft.com/en-us/windows/win32/api/windowsx/
https://jeffpar.github.io/kbarchive/kb/083/Q83456/
I got to learn it via the "Programmer's introduction to Windows 3.1" book,
You don't have to use all features of C++ - if you prefer a more C like style you can have that.
[0] https://www.boost.org/doc/libs/1_48_0/boost/strong_typedef.h...
1. http://blog.llvm.org/2020/04/the-new-clang-extint-feature-pr...
I would certainly hate(and would continue using them btw) that they remove normal pointers from C. It would be like removing s expressions from lisp.
I believe that the solution to this "mistake" is just not using c directly, using other languages to write C code for you, or use c primitives that are 100% well tested.
That is what we do, our c primitives-libraries-modules are written and tested by lisp and our own language.
It it then very easy to use that code in python or c++, swift or whatever as libraries or modules.
C++ gives you more tools for avoiding them than C, BTW.
A direct link, for the curious: https://en.cppreference.com/w/cpp/container/array
As an aside, the C’s Biggest Mistake article cropped up in HN discussion 8 days ago, https://news.ycombinator.com/item?id=24373728
This doesn't diminish the advantage of std::array, though, as it embeds the size of the array into the object, unlike when a raw array is passed and 'decays' to a pointer.
template <size_t N>
void foo(int bar[N])
And this gives you the size without an additional size parameter as you’d usually need in C (of course with the limitation that the parameter now has to be a compile-time sized array).The actual mistake is to don't pass size_t as a user. This is one kind of "premature optimization". We can safely say the language design doesn't encourage the user to write safe code, and succeror languages do that.
Don't get me wrong — I just try to do the point that C itself is not the point to blame. It's people using computers who write the million dollar bugs.
You would still be able to declare your function as taking a pointer (instead of an array, which in this world would be a far pointer) if you need to
He's saying to deprecate char[] as a parameter type, not char
struct string123 {
char data[123];
};
Then create functions that user pointers to these string123 structs.1. variable length buffers
2. every other piece of code you want to interface with uses `char*`
char (*data)[123]; // syntax is somewhat awkward
D allows passing both raw pointers as parameters and pointer/length pairs. It's up to the user to choose. In practice, people have simply moved away from using raw pointers into buffers.
As for performance, in C to determine the length of a string one uses strlen(). Over and over and over again on the same string. This can be a major performance problem, even not considering the memory cache effects. When I look at speeding up C code, often the first nuggets of gold is reviewing all the explicit and implicit uses of strlen(). (Implicit uses are functions like strcat()). It's also the first place I look for bugs when reviewing C code - anytime you see a sequence of strlen, strcat, strcpy, it's often broken (typically in neglecting somewhere to account for the extra 0 byte).
#1 and #2 are integer overflows and aliasing mistakes.
Yeah and all the string functions should have been marked as depreciated with C89 and fully depreciated with C99.
It's not unheard of to have millions of tiny little (< 10 character) strings, and not storing lengths alongside them can shave off a sizeable portion of space requirements.
Also, consider that the terminating 0 byte isn't just one byte. There's also the alignment of what malloc returns, which may be 4 or 8 or even 16 bytes.
Which is why one should never allocate a single (short) string from a generic allocator. Instead, one allocates a big chunk upfront (e.g. 4K bytes or more), and breaks from that, using a simple index that points to the first unused bytes.
In this way, the overhead of allocating a string is truly only the terminating zero byte - no alignment constraints. This scheme is easy to implement as long as strings don't need to be freed individually.
For one, strings are just chunks of memory like other arrays. So for almost any string that is not a literal in the source code, you just store offset/length as needed, like for any other array. I have sizeable projects (on the order of 10K lines) that have maybe 0 or 1 instances of strlen() in the code.
Very often though, strings are bounded and pretty short (since they are meant for human consumption), and in these cases, using strlen() when looking up a string is often sensible, since persisting the length might require more memory than the string itself.
Another case is when you're scanning the string from left to right anyway, so you just "stream" through it until you find the terminating 0. That's how printf() and friends work (they take a formatting string), and arguably this scheme works just fine.
Btw, the length of a string literal is (sizeof "Hello" - 1). You can also initialize char arrays using string literals and have the size available:
static const char name[] = "Foobar";
static const int nameLen = sizeof name - 1;Also, look at functions like sscanf(). It can be orders of magnitude slower than fscanf() because every invocation calls strlen while fscanf incrementally reads from the current file pointer. I don't know why sscanf doesn't also work incrementally but the implementations I've tested don't do that.
The main point here being: if strings had a size_t size plus data, that would change an O(n) scan to an O(1) length lookup, and that would have huge performance gains throughout the C library, not to mention your own code as well.
I figure it would be possible for musl to implement that stream using a function that scans for NUL and copies at the same time, but maybe that's not an improvement in the end.
It would be much simpler of course if sscanf() would take the length of the input string as an additional argument. But actually, I don't really care.
Because, does this even matter? Using sscanf() is far from ideal anyway. The stdio functions are not what you use if you're going for performance. Their conversions are probably not the fastest (being quite featureful), and they are even locale dependent which is a huge mess!
Heck, when we're going for performance to a degree where a strlen() matters (bear in mind that we have to read the input at least once anyway, so the waste is definitely bounded) we should certainly not be parsing text at all. That is much more wasteful in comparison.
Much if not most of libc is there to provide you a portable base to (comparatively) quickly get your project up and running, and to simply to keep old software going, but it's certainly not to help you achieve performance.
http://git.musl-libc.org/cgit/musl/commit/?id=18efeb320b763e...
That's not going away in the article's proposal. It's being complimented by an array syntax that makes the current size (in memory) of a non-static data structure bounded upfront.
void foo(size_t length, char (*x)[length]){
size_t size = sizeof(*x);
assert(size == length);
printf("sizeof(*x): %zu\n", sizeof(*x));
}There are actually two problems here. One is the absence of bounds checking. The other, which is related but technically orthogonal, is the hole in the type system: an array is not an object. It's been true since Unix v7 that you can pass or return a struct by value, but you can't pass an array by value unless you wrap it in a struct.
The type system also makes no distinction between a pointer to a single object and a pointer into an array of objects. I've worked on static analysis tools that try to find potential buffer overflows, and this turns out to be a surprisingly big problem. One has to do a global dataflow analysis just to discover which pointer variables could ever point at array elements.
About the singular vs array question: this is true but just a special case of the absence of bounds checking, is it not? C's approach of allowing e.g. "int i;" to be addressed as if it were an array of length 1, allowing construction not just of &i but also &i+1 as pointer values, is valid and sometimes useful, but you have to make sure you never access *(&i+1). That's the same problem as how given "int a[2];", accessing a[2] is not valid, as far as I can see.
But I guess my statement that "C doesn't have array expressions" was true before the advent of C99. And that's also why array decay made even more sense back then. (It still makes a lot of sense today IMO).
So determining the type of "&a" is not an issue, it's just one case in determining the type of a C expression (look up the object "a", is it an array? The type of the expression is a pointer to the array element type).
This is not a special case, at least not more special than how to determine the type of the expression "a", or "1".
No array type needed.
It's just too hard to "do things right" in C. Even people who have been "doing things right" for decades make mistakes.
int a[10];
you allocate 10 integers and "a" is the pointer to the first one of them.
Arrays are just memory, just like what you get wen calling malloc, and memory is accessed using pointers in C.
What you mean is maybe more aptly named "dynamic (memory) allocation".
Although C is usable on many types of hardware, interesting things could be done, e.g. on desktop OSes if certain hardware-specific extensions were made.
One of the things I wish we could do with our now-absurdly-large pointers (64-bit) is to reserve a handful of bits for other information such as the size. Sure it means we can’t store anything at location 2^64-1 but it wasn’t that long ago we only had 32-bit pointers and the 33rd bit is twice as many addresses all by itself so I think we can lose a few.
For example, if all allocations were rounded up to buckets of a certain size, the precise byte count would not need to be encoded in the pointer (just the number of buckets, requiring fewer bits). There could be a couple bits to give pointers a type for other interesting scenarios, e.g. perhaps a pointer identified as an “immediate value” that isn’t actually allocated at all, and it is “dereferenced” by treating its “address” as the “stored” value. There could even be a couple of bits to track use of common allocators (it would be so nice to simply know that a pointer was allocated by "malloc" vs. "new" for example).
In high-level languages, then, the syntax change would be not to identify arrays specifically but pointers with encodings that are “complete” (e.g. "char const complete*" or something), covering both stack arrays and dynamic buffers.
All that is needed is a mechanism for forming a fat pointer from a pointer and a length. In D this looks like:
int* p = cast(int*)malloc(length * sizeof(int));
if (!p) fatalError();
int[] a = p[0 .. length];
...
int x = a[length + 1]; // runtime error: buffer overflow
In C, this could be done via a macro with no additional core language changes. struct foo {
int a[10];
} f;
func(&f);In addition to unsafe/irregular buffer handling, I also constantly see poor data structure choice, presumably due to a lack of default choice of library. It is very common to see code scanning linked lists when they should be doing map look ups. (And often even the linked list operations are ad-hoc and repeated for every type of struct with an embedded next/prev pointer.) Everyone always defaults to linked listing it up because they never have to pay the up-front cost of finding a library or investing in re-inventing the wheel. I think this is also why you see so much sketchy buffer code - no one has bothered investing in safer buffer abstractions.
Perhaps some of this is caused by the difficulty of taking on dependencies in a portable way. (CMake/Autotools can make this better, but it is a far cry from NPM.)
Unfortunately the iPAX 432 was delayed, so Intel introduced the 8086 as a stopgap processor and computers have been using the x86 architecture ever since. It's interesting to think that if history had gone a bit differently, whole classes of security problems would not exist.
That's a strong claim for something as vaguely defined as "the length of an array."
What if I allocate a huge array of memory and emulate another processor using that memory? What if I subpartition an array? What if I overallocate an array to avoid reallocations?
This is C, not C++. Keep it simple.
Here, the idea is that there is no special type for "pointer+size" (what the author proposes as an array). Ok, let's add one and see the implications.
- How do I get the size, the number of elements? A "sizeof" like operator?
- Can I resize the array? If yes, how? If no, why?
- What happens if I overflow? Undefined behavior?
- A memcpy-like would be an obvious function to implement, what happens if sizes differ?
- What is the relationship between static arrays (ex: int a[5]) and "pointer+size" arrays? Are these completely different types? Is there an implicit cast between the two?
- About casting, how can I go from a separate pointer and size to an array and vice versa? If it is possible at all.
- What if I do a bit of pointer magic to access the internal representation of the array? Probably undefined behavior.
It is much more complex than "just add array[..]", I expect more tradeoffs, more undefined behaviors (C wouldn't be C without them). Adding complexity to the language can actually make things worse.
As for zero-terminated strings, they have advantages and drawbacks. They are preferred by the C library, but you can do pointer+size if you want by using mem* instead of str* , or %.*s instead of %s in printf (not sure about this one).
Here are it's answers:
- How do I get the size, the number of elements?
Builtin .len field operator. @sizeOf works too.
- Can I resize the array? If yes, how? If no, why?
No. Array lengths are comptime known; there is something called a slice which is runtime known and bounds checked in safe releases.
- What happens if I overflow? Undefined behavior?
Panic, in safe releases. UB in dangerous releases (small or fast)
- A memcpy-like would be an obvious function to implement, what happens if sizes differ?
Bounds checked at runtime for safe releases.
- What is the relationship between static arrays (ex: int a[5]) and "pointer+size" arrays? Are these completely different types? Is there an implicit cast between the two?
Yes, and yes. Arrays can implicitly be converted to slices at compile time; slicing into a slice with compile-time known index width yields an array.
- About casting, how can I go from a separate pointer and size to an array and vice versa? If it is possible at all.
There is an escape hatch function for this.
- What if I do a bit of pointer magic to access the internal representation of the array? Probably undefined behavior.
It's defined, but unsafe.
Do I haven't really worked too much in zig (it's not my daily driver) but I think it says something that all of these questions have answers to them, they are sensible, and very easy to remember.
sizeof(array)
> the number of elements? array.length
> Can I resize the array? If yes, how? If no, why?In D it would be:
T* p = realloc(array.ptr, newLength * sizeof(T));
array = p[0 .. newLength];
A C macro could handle that in C.> What happens if I overflow?
2's complement arithmetic happens.
> A memcpy-like would be an obvious function to implement, what happens if sizes differ?
That would be up to the implementor of the function. There are many ways.
> What is the relationship between static arrays (ex: int a[5]) and "pointer+size" arrays? Are these completely different types?
Yes, they are different types.
> Is there an implicit cast bet>ween the two?
Conversion between a static array and the dynamic array should be implicit. Can't go the other way.
> About casting, how can I go from a separate pointer and size to an array and vice versa? If it is possible at all.
In D, this would be:
a = p[0 .. size];
In C, I expect a macro can do it: a = ToArray(p, size);
> What if I do a bit of pointer magic to access the internal representation of the array?Then you'd need to be careful to do it right, as in any systems language when you go under the hood.
> using mem* instead of str*
That only works for some functions. Not fopen(), for example.
> or %.*s
I use that a lot, but it's annoying because the length argument must be an int, and so must be cast from size_t.
But even weirder is the total lack of a proper string library. Nobody but Microsoft uses whar, and they are the only ones with the proper whar_t size. Everybody else was wrong with size 4. But nowadays it should be clear that only u8 is the only way forward, C++ even adopted now char8_t for it. But they all still ignore the unicode problems with an overly simplistic, glorified null-terminated memory buffer library. These are not strings anymore nowadays. Strings have multiple representations of characters in unicode, strings need the unicode version to be exposed which changes every year. They need a proper fold case and norm API, otherwise you cannot compare them, so you cannot search for strings. Grep would be happy to find unicode strings, but it still cannot. coreutils still cannot do unicode in 2020.
Also the complete lack of security, esp with names, ie identifiers. Such as pathnames. Most filesystems just ignore security, spoofing, bidi changes, mixed scripts as if this problem does not exist at all. Strings are not normalized, not properly fold cased.
The _l locale mess, it still relies on global runtime state, which is not compile-time optimizable, in opposition to _l or simply just a new u8 API. Not reentrant. Not compile-time optimizable. It's a huge mess.
gcc cannot do compile-time constexprs checks, only clang can, leading to up to 200x faster libc code. gcc cannot do user-defined warnings of errors.
glibc, FreeBSD libc, musl, none of it fixes anything.
In particular you used to use structures in a weird way, essentially field names within structures lived in their own global name space, any pointer could be used with any structure - there weren't unions yet so this was used for good effect in Unix kernel drivers (there was a standard buffer queue header you could add your own stuff at the end of.
I think this was kind of descended from the BCPL/Bliss world view where explicit pointers were a relatively new thing in languages and their typing was pretty simple (there was a limit to the number of indirections allowed) - fully orthogonal typing systems were only just becoming a thing then.
Also I suspect that the idea that a[i] was the same as a+i was an idea with legs, this is still legal C:
char *x()
{
int i; char *p;
return &i[p];
}Wait I would just use `sizeof` but then I'm still doing pointer math then?
#define ARRAY_SIZE(x) ((sizeof x) / sizeof *(x))
But ideally there would be a builtin to do this, since this gives nonsense results if you pass a pointer instead of an array.1- Trying to find the proper style and methods to write few lines of code. The reason for this is because C is an old language that kept changing. Thus, you can read a book, yet find someone to tell you “you shouldn’t do it that way”.
2- compilers made C into different flavors. Microsoft C compiler provides scanf_s with the old scanf being deprecated. In the other hand, gcc has different approaches without the scanf_s that Microsoft has. This can be so annoying to use.
I for one would love to see this proposal become a reality.
If C did pass array lengths it still wouldn't matter since C doesn't (and in my opinion shouldn't) check for overflows.
Because it's a language, not an implementation. An implementation is free to do so (and there are such implementations after all).
void f(size_t n, int (*v)[n])
instead.I wonder if a better C can be made by just stripping-down the bloated C++ and introducing the "unsafe" keyword for dangerous features like directly using arrays, etc.
typedef struct fat_ptr_t {
size_t size;
void * start;
} fat_ptr_t;
extern void foo(fat_ptr_t a);
if it's a good idea to use a fat pointer, why do we need new syntax to sell it? What am I missing here?- A concise syntax for declaring, accessing and mutating them. Dealing with lists is such a common thing in programming languages, that it's simply crazy to not have a proper syntax for them.
- Generics/templating, so that you can use concrete types instead of 'void *'. Having that prevents mistakes and also tends to make code self-documenting.
"If it's such a good idea to use a safety belt and airbags, why do we need special devices for it? Why can't I just use a piece of rope I had in a drawer and some leftover balloons from my previous birthday party?"
void* was merely for example. Yes make it typed when you use it in your C code. Also used accessor functions to wrap array index so you can switch on & off a macro for bounds checking, absolutely do that. Does new syntax change anything if you do these things?
Don't much care for the seatbelt analogy there, remove your working, properly fitted seatbelts with our red ones because they're easier to see? These kinds of analogies always break down. Especially car analogies for programming and yes, I use them too.
I assure you, there are arrays in C.
int a[100]; // `a` is an array, not a pointer
int* p; // `p` is a pointer, not an array
a = p; // error, array is not a pointer
> The array-ish syntax that’s available is just a some syntactic sugar on top of pointers.Sorry, this is incorrect. In some circumstances, C will implicitly convert an array to a pointer, which is what the article is about, but don't mistake a conversion with identity.
Conflating pointers with arrays.
I don’t mean them using the same syntax, or the implicit conversion of arrays to pointers. I mean the inability to pass an array to a function as an array, even if it is declared to be an array. C will silently convert the array to be a pointer, and will rewrite the function declaration so it is semantically a pointer:
[...]
This seemingly innocuous convenience feature is the root of endless evil. It means that once arrays leave the scope in which they are defined, they become pointers, and lose the information which gives the extent of the array — the array dimension. What are the consequences of losing this information?
An alternative must be used.
For strings, it’s the whole reason for the 0 terminator.
For other arrays, it is inferred programmatically from the context. Naturally, every situation is different, and so an endless array (!) of bugs ensues.
The trainwreck just unfolds in slow motion from there.
The galaxy of C string functions, from the unsafe strcpy() to sprintf() onwards, is a direct result. There are various attempts at fixing this, such as the Safe C Library. Then there are all the buffer overflows, because functions handed a pointer have no idea what the limits are, and no array bounds checking is possible."
PDS: The root of all of this -- is that C, being a low-level, close-to-the-hardware, designed in the 1970's programming language (some in academia pejoratively call it a "glorified assembler"), was not designed with a proper string storage class as we know them in programming languages today; instead, arrays of characters were substituted for this purpose, and those arrays were not implemented containing total size (length) and dimensionality information.
Basically an array in C -- is a set of contiguous memory, which has a starting address (the pointer passed), and a stated element size that the compiler knows about, but not the length (aka, total size, element count, etc.) of that array, nor its dimensionality.
Observation: C's arrays need length information signalled in an out-of-band fashion (that is, this information cannot exist as a zero (0) -- somewhere in the array).
The irony of all of this is that C was invented at AT&T, and AT&T for the longest time had difficulty with phreakers exploiting 2600hz signals to gain access to its long distance trunk lines, from which they could call to anywhere in AT&T's system for free.
But, that's what the engineering error of in-band signaling generates...
C, by using arrays to implement strings, and letting the zero terminator (information about string length) exist in the memory space of the string, made exactly the same engineering mistake -- just in software -- and that is the mistake of in-band signaling.
Now, that being said, hindsight is 2020, and it couldn't be expected that Dennis Ritchie, who invented C in the 1970's would have foreseen the consequences of that engineering "mistake" (AKA, "act which generated quite the education for a future populace". <g>).
Such is the price of being an innovator.
On the one hand, he advanced computer technology greatly -- far beyond the technology advancements created by most of his contemporaries of his day...
On the other, that advancement gave us this highly educational engineering "mistake" -- that we can all learn from!
Such is the price of being an innovator -- and pressing the "bleeding edge" of what is possible...
Humanity could not be advanced without such innovators, and the occasional future flaws (and the wisdom that comes from examining them in hindsight!) that their innovations generate...
AND Strings.
FTFY.