C’s Biggest Mistake (2009)
digitalmars.com
digitalmars.com
That said, maybe new C standard should come up with all-batteries included standard library upgrades. On second thought, imagine the compilers snafu.
I love C and I love Rust, for different reasons.
A proper import system would massively reduce the amount of wheel reinventing because then it would not be fiendishly difficult to actually reuse code in a coherent way. Only after that does it make sense to start quibbling about what is included where and by default.
Tagged union types would also be really, really nice.
Also the standard lib and POSIX reminds me of a homeless persons shopping cart full of precious stuff they picked off the street.
It is almost a meme here, that, on any mention on C and C++, there will be people contributing comments: 'but Rust ..'.
C supports modularity since the 70s. You do not need a standard library with everything and the kitchen sink. The standard library is not the only component you're able to consume.
If there's a lesson from high level languages such as python and node.js, it's that being able to consume third party libraries is all anyone needs to get stuff done, and standard libraries bring convenience to the table but they aren't critical to get stuff done.
I think today we understand more about what a language needs to provide to enable modularity. One issue that comes up over and over with interoperating libraries is confusion around ownership. If you pass a char* to me, do I own it now? Is it my responsibility to free it? Or is it just a short-lived reference to data that's still yours? Garbage collected languages can kinda sorta sidestep these questions, at least with memory-use-after-free, and they're more likely to tolerate the costs of immutable data structures and defensive copies. But in C this question is critical, and we screw it up constantly, especially at the interface boundary between libraries. C++11 does much better, with first class support for moving ownership through a function call, and Rust goes further with lifetimes and the borrow checker.
Another point that's kind of in the weeds is the atomic memory model, which C didn't have until 2011 (and I think C++ gets most of the credit here for formalizing it). I'm still waiting for proper support for C11 atomics in MSVC. Multithreading has been ubiquitous for a while, and while it's possible for an application to decide that everything's going to be single-threaded, that's not an option for a portable library. Without standard atomics, you have to resort to platform-specific extensions to do basic stuff like initializing globals.
But likely neither of these is as big as letting void * represent any arbitrary construct, be it pointer or data structure or the universe itself. Finding references to void * in a codebase is akin to crossing the event horizon of a black hole.
What's the difference between returning a two element struct and a pair of value and error?
I never understood why "multiple return values" is a thing anyone should care about.
2. Without first-class support, it's very difficult to assert that the error is actually being checked. Consider:
int someVal = getSomeVal(...).v;
doSomethingWith(someVal);
vs. (int someVal, error e) = getSomeVal(...);
doSomethingWith(someVal);
Which is easier to catch at code-review time, or even at linter time?In most languages that don't support multiple return values, both problems are solvable with templates or local equivalent - see, for instance, absl::StatusOr for C++ (https://abseil.io/docs/cpp/guides/status).
one trivial possibility/option is to define in-out parameters, and just return errors:
<error-type/int/…> foo(T*, <params>)
will that not be ok ?I wonder if typeof in c23 has changed this at all. Previously there was no sense in defining an anonymous struct as a function's return type. You could do it, but those structs would not be compatible with anything. With typeof maybe that's no longer the case.
e.g. with clang 16 and gcc 13 at least this compiles with no warning and g() returns 3. But I'm not sure if this is intended by the standard or just happens to work.
struct { int a; int b; } f() {
return (typeof(f())){1,2};
}
int g() {
typeof(f()) x = f();
return x.a + x.b;
}
edit: though I suppose this just pushes the problem onto callers, since every function that does this now has a distinct return type that can only be referenced using typeof(yourfn).> In most languages that don't support multiple return values, both problems are solvable with templates or local equivalent
I don't really see the difference between tuples in C++ and multiple return values in Go (since you cite C++ as a language that doesn't have multiple return values).
val, err := do_something()
or auto [val, err] = do_something();
What's the real difference there? Syntactic sugar? package hello
func good() (int, error) {
return 0, nil
}
func bad() *struct {
int
error
} {
return nil
}For the same reason people care about multiple parameters. There's no reason that the data coming out of a function should be in any way more constrained than the data going into the function.
Go literally hardcoded the same error-prone errno-like error handling, it just has some syntax sugar now. It still doesn’t compose, is unreadable and gives only the impression of properly handled error cases. It even fails to return a proper sum type, so now you get a possible important value and an error value as well.
That article was from 2009. C11 didn't fix it. C18 didn't fix it. C23 didn't fix it.
Local variables are stack frame dependant. It makes perfect sense to make an array a pointer when using it out of the stack frame. It forces the programmer to actually think about what is going on underneath.
Edit: I just realised the article was written by D's Author! So maybe that is C's biggest mistake.
So ... yeah ... Walter's right, IMO.
C++ adds a number of its own flaws, but that's beside the point when discussing C's flaws from a perspective of knowing C++.
C has weaknesses. But it also has strengths, and those strengths mattered. It turns out they mattered more than the set of strengths that JOVIAL had.
Turns out people will adopt anything that is free, without regards to consequences.
Had UNIX been sold under a commercial licence from the get go, it would have been as successful as Plan 9.
And thus the world would never suffered from C's existence.
The reality is that your approach didn't meet the needs of the real world as well as the C approach did. You may dislike it, you may resent it, but your way didn't work very well.
It's not only because Unix was free, and C came along for the ride. It's also that Multics took several times as long to write, and required bigger hardware. (So Unix would have been cheaper even if the OS was not free, both because of lower cost to write, and because of lower hardware cost for a system to run it on.) And less cost opened up far more possible uses, so it spread quickly and widely. (How many Multics installations were there, ever? Wikipedia says about 80.) Which is better, the language and OS that have flaws, or the language and OS that run on hardware you don't have and can't afford?
And Unix was far more portable. Didn't have a PDP-11 either? No worries; it was almost entirely written in C. You could port it to your machine with a C compiler and a little bit of assembly (a very little bit, compared to any other OS). Didn't have C either? It was a small language; it wasn't that hard to write a compiler compared to many other languages. If you were a university, you could implement C and port Unix on your own. But once on university had done so for a particular kind of computer, everyone else could usually use their implementation.
And the final nail in the coffin of your style: Programmers of that era preferred that languages get out of their way far more than that languages hold their hand. Your preferred kind of languages failed, because they didn't work with people. It may have worked with the people you wanted to have, but it didn't work with the people that actually existed as programmers at the time. A tool that doesn't fit the people using it is a bad tool.
But all is not lost for people like you. There are other languages that fit your preferences, and you can still use them. Just stop expecting them to become mainstream languages. The languages that the majority of programmers found better (for their use) are the languages that won. And that majority was not composed entirely of fools and sheep.
I have not had any reason to touch C since 2001, other than to keep up with WG14 work, exploit analysis, and security mitigations from SecDevOps point of view.
History lesson, there were other OSes other than Multics.
Also UNIX only became written in C on version V, and contrary to urban myths spread by UNIX Church, it wasn't the only OS written in a portable systems programming language.
It was the only one that came with free beer source tapes.
Now with governments looking into cybersecurity bills, let's see how great C is all about.
It will be great for lawyers, that's for certain.
If not, then you didn't really answer what I said. You made an argument that used some of the same words, but didn't actually answer any of the substance.
The security complaints... yeah. A safer language would give you a smaller footprint of vulnerabilities, which would have made your life easier for last two decades. (Maybe you earned being bitter about C.) That's unrelated to the spread of C and Unix in the 1970s and 80s, though.
Yes, they were as easy to port as UNIX was, for those that owned the code, Unisys still sells Burroughs as ClearPath MCP, nowadays running perfectly fine in modern hardware.
1980's he says,
"Many years later we asked our customers whether they wished us to provide an option to switch off these checks in the interests of efficiency on production runs. Unanimously, they urged us not to--they already knew how frequently subscript errors occur on production runs where failure to detect them could be disastrous. I note with fear and horror that even in 1980, language designers and users have not learned this lesson. In any respectable branch of engineering, failure to observe such elementary precautions would have long been against the law."
C.A.R Hoare in 1980, during his Turing Award speech.
"Oh, it was quite a while ago. I kind of stopped when C came out. That was a big blow. We were making so much good progress on optimizations and transformations. We were getting rid of just one nice problem after another. When C came out, at one of the SIGPLAN compiler conferences, there was a debate between Steve Johnson from Bell Labs, who was supporting C, and one of our people, Bill Harrison, who was working on a project that I had at that time supporting automatic optimization...The nubbin of the debate was Steve's defense of not having to build optimizers anymore because the programmer would take care of it. That it was really a programmer's issue.... Seibel: Do you think C is a reasonable language if they had restricted its use to operating-system kernels? Allen: Oh, yeah. That would have been fine. And, in fact, you need to have something like that, something where experts can really fine-tune without big bottlenecks because those are key problems to solve. By 1960, we had a long list of amazing languages: Lisp, APL, Fortran, COBOL, Algol 60. These are higher-level than C. We have seriously regressed, since C developed. C has destroyed our ability to advance the state of the art in automatic optimization, automatic parallelization, automatic mapping of a high-level language to the machine. This is one of the reasons compilers are ... basically not taught much anymore in the colleges and universities."
-- Fran Allen interview, Excerpted from: Peter Seibel. Coders at Work: Reflections on the Craft of Programming
" We really are using a 1970s era operating system well past its sell-by date. We get a lot done, and we have fun, but let's face it, the fundamental design of Unix is older than many of the readers of Slashdot, while lots of different, great ideas about computing and networks have been developed in the last 30 years. Using Unix is the computing equivalent of listening only to music by David Cassidy. "
-- Rob Pike on his Slashdot interview
And to finish on security,
> The combination of BASED and REFER leaves the compiler to do the error prone pointer arithmetic while having the same innate efficiency as the clumsy equivalent in C. Add to this that PL/1 (like most contemporary languages) included bounds checking and the result is significantly superior to C.
https://www.schneier.com/blog/archives/2007/09/the_multics_o...
"Multics B2 Security Evaluation"
https://multicians.org/b2.html
> Although we entertained occasional thoughts about implementing one of the major languages of the time like Fortran, PL/I, or Algol 68, such a project seemed hopelessly large for our resources: much simpler and smaller tools were called for. All these languages influenced our work, but it was more fun to do things on our own.
https://www.bell-labs.com/usr/dmr/www/chist.html
More fun did they have indeed.
C compiler development has forgotten this. C compilers must not be too clever in optimizing. They should generate tidy code that allocates registers well, and puts local variables into registers well, and peepholes away poor instruction sequences. But all the memory accesses written in the program should happen.
Maybe Steve and all the early C people were philosophically wrong, but that's the language they designed; it should have been respected. In C, how you write the code is supposed to matter, like in assembly language.
First, "such a project seemed hopelessly large for our resources: much simpler and smaller tools were called for." This actually is what I was arguing - you could do things (like write or port an OS) with a much smaller group, which opened doors that were closed by more "advanced" tools/languages. The barriers to entry were lower, so lot more people were able to do a lot more things. There was massive value in that.
Second, it's not like the legislature made it illegal to work on all these other approaches. They were abandoned because the people working on them abandoned them. Nobody put a gun to their head.
This kind of goes back to the first point. C opened a lot of doors. Research stopped in some directions, because new areas opened up that people found more interesting.
In fact, this whole thing has a bit of a flavor of the elites mourning because the common people can now read and write, and are deciding what they're going to read and write, and it's not what the elites think they should. You (and the people you quote) are the elites that the people left behind. You think that yours is the right path; but the people disagree, and they don't care what you think.
(The point about hardware from 1958 being less than a PDP11 I will concede.)
I wonder if in several thousands year, mankind will have managed to get rid of C.
Zig in particular could scarcely do more to scream "I'm a replacement for C" without maybe asking WG14 to write "Just use Zig instead" in their next document.
As does Rust, though it makes for a bigger change.
This is possible for fixed-length arrays using 'myType (*myVar)[size]' as the argument type.
The annoying thing about this is that the array is just sitting there on the stack of the calling function. The size is completely known at compile time. But it gets thrown away as soon as you call a function.
If the syntax is the problem, yeah, if you want first class data structures you might not like C.
The title of the article is “C’s biggest mistake.” If you’re going to rebut the article, do so. Otherwise your reply just comes off as “this is the way it is, deal with it!” which is a pretty shallow dismissal.
That's what an array is. Fixed size. If we want to talk about a slice of some runtime determined amount of a thing then we need fat pointers, and C doesn't provide any fat pointer types so it's not a surprise to find it can't do slices.
There is no reason this shouldn’t be possible. It should not require fat pointers at all because the size information is known at compile time.
This is exactly the solution proposed in TFA..
> using the int array1[2]; like this: arraytest(array1); causes array1 to automatically decay into an int .
> HOWEVER, if you take the address of array1 instead and call arraytest(&array1), you get completely different behavior!
> Now, it does NOT decay into an int
package main
func foo(a []byte) {
println(len(a))
}
func main() {
var b []byte
// cannot use &b (value of type *[]byte) as []byte value in argument to foo
foo(&b)
}https://lemire.me/blog/2020/09/03/sentinels-can-be-faster/?a...
Results:
N = 10000
range 1218.1
sentinel 937.258
mine 672.227
ratio 1.29964
(ratio is range/sentinel, not mine/whatever)You can get even more crazy with SIMD. But for all that you need to know the length beforehand.
Edit: The b4 variable should actually be called b8. That's a remnant of a previous version, where I used 32-bit chunks.
https://lemire.me/blog/2020/09/03/sentinels-can-be-faster/?a...
package main
import "fmt"
func main() {
fmt.Printf("%q\n", "hello \x00 world")
}So, the desire and need for that flexibility was entirely warranted.
So C was [mostly] great, but today we have better options.