What "old ideal" ? Where do you believe this "ideal" was expressed? Should I expect to find it in the First Edition K&R perhaps? Or in the documentation for the original C compiler? In the ANSI standard ?
C was never this mythical "portable assembly language" that's just something people say about it, mostly to ridicule it, you are engaging in Nostalgia.
"The language is also widely used as an intermediate representation (essentially, as a portable assembly language) for a wide variety of compilers, both for direct descendents like C++, and independent languages like Modula 3 [Nelson 91] and Eiffel [Meyer 88]. "
If they are right or not, it is another matter.
For Virgil I chose to settle on 2's complement because it's what hardware gives and has no overhead. It comes with the full compliment of fixed-width integers of widths 1-64 and odd-width ones come with zero-extension or sign-extension as necessary to make the underlying hardware width unobservable.
Yes, and which is exactly what I hinted at with my question to your comment. With the HW we have today, implementing such semantics without an extra hit is not possible and which is why I thought your comment wasn't completely fair but coming from more theoretical stance.
I suppose range analysis can eliminate many overflow checks, e.g. in counted loops, but it's not completely zero cost.
Yes, but now as well you have to return a value to communicate the overflow to the call site. And call site has to check for that value and that's yet another and another ... and another branch. Depends how deeply you want to propagate that error and decide how (?) to deal with it, this will grow the code size, which can contribute to the higher frequency of I-cache misses and page-faults, but it can also inhibit compiler optimizations.
On the CPU level, I think there also could be an attached cost as well in case some of those branches end up as entries either in branch-target (BTB) or branch-order (BOB) buffers or both. Sizes of these buffers are quite scarce so ending up with the unfavorable ratio of check-for-overflow entries vs entries occupied by other type of branches found in the code is something that will put more pressure to our branch-prediction unit. More "important" branches will now more frequently start to lack their entry in the branch history simply because of the fact that we started sprinkling check-for-overflow branches. And yet we know that branch misprediction is the costliest operation (15-20 cycles) we can encounter in the CPU pipeline.
Also, I think a bigger picture must be observed in this context. E.g. what is the percentage of arithmetic operations some big real-world sized binaries contain? I'd figure that in average it would be a sizeable amount, and in ones with a lot of math even more so. And then I wonder what we could observe if we applied the check-for-overflow transformation to all such signed-arithmetic operations.
I'm aware that there are some artificial benchmarks showing that there's no cost attached to branches which are essentially never taken but it makes me wonder if that cost would really be zero if we exercised that change on the actual code instead. For at least the reasons from above.
Reality though is that we don't have such hardware and we have to deal with it in software, so, instead of a single instruction emitted there will be a bunch of them, and of which there will be certainly branches involved.
Only if you ignore
1. variadic arguments as in printf,
2. function pointers which are a very elementary (type-unsafe) form of closures,
3. a nice way of casting to void
4. setjmp and longjmp goodies (?) which allow you to code up co-routine libraries and exception handling mechanisms
I'm sure there are more such facets I am missing.
The elegance of the specification may be questionable, but the scope of what C tried to achieve is breathtaking. It actually is superior to most of its improvements.
Function pointers as an elementary form of closures, come on, what’s next? Closures are defined by capture. It’s nearly as fun as pretending C as coroutines because of setjmp.
How can a statement like C being semantically poor even be seen as controversial? For god sake, we are talking about a language which semantically doesn’t even have proper arrays.
> The elegance of the specification may be questionable, but the scope of what C tried to achieve is breathtaking. It actually is superior to most of its improvements.
Seriously? It wasn’t even a good language when it was released. Lisp and Pascal were far better. It won because of compiler availability and adequate performance on limited platforms.
HN really is a joke sometimes.
A lot of what I'm seeing from your comments here is you can't handle people who have something positive to say about C.
They're not saying it's the one true way or something. Just that they like it in some respect.
And your attempts to dismiss that and call "HN" a joke for harboring someone who thinks this way look kind of childish to me.
Considering I was having interesting discussion about the subtleties of the Hindley-Milner type system on this same website a decade ago, yes, I do think HN is becoming a joke. The joke is on me however because apparently I keep commenting for reasons which are not always apparent to me I must confess.
Small example, linked lists. I don't think non-C linked list code tends to be as straightforward as I've seen in C.
Or the character-at-a-time style of string processing. It's kind of unique to C.
You can say there is stuff about that you don't like. That's fine. Linked lists suck with modern CPU caches anyway. C strings have lots of misadventures in terms of buffer overflows. But it's unique and interesting. Lots of elegant things have been written this way. Your unfamiliarity with it doesn't make it "a joke" to point this out.
I have my own gripes with HN. Try mentioning politics and it brings out all sorts of fascist-sympathizing crazies. But saying good to neutral things about C (while not even universally praising it) is not one of those issues.
You are confusing what the language can do and what is semantic is.
C strings are just a contiguous allocation of byte and a bunch of functions which interprets it as ascii characters and stop on a specific value. That’s literally the worst representation of a string you can have.
const char *my_strchr(const char *str, int ch)
{
if (*str == 0)
return NULL;
if (*str == ch)
return str;
return my_strchr(str + 1);
}
It's a good format for storage and communication.Name any other string data structure and I will cite you all the disadvantages compared to the C string.
- If the format contains pointers, you have to marshal it to a flat representation to communicate or store it. The C string is already marshaled.
- If the format contains size fields, their own size matters: are they 16, 32, 64. What byte order? Again, needs marshaling. You might think not if going between two processes on the same machine. But, oops, one is 32 bit the other 64 ...
I shudder to think of what FFIs would look like, if C strings were something else. It's the one easy thing in foreign interfacing.
C programs themselves sometimes come up with alternative string representations, for good reasons; but those representations don't interoperate with anything outside of those programs, or groups of closely related programs.
You are basically arguing that C strings are a good default because they are the default. If they were something else, well, we would have saner FFI in a lot of place.
C strings are a terrible default. They are fundamentally unsafe. Mishandle the null char for any reason and you now face a serious security issue.
Not char/byte C strings. You might be thinking of wchar_t strings.
> Mishandle the null char for any reason and you now face a serious security issue.
With what string data structure can applications peak into and mishandle an implementation detail and not risk creating a security problem?
Same you have a structure with the length and a pointer. Mishandle the length field or pointer and you have a problem.
It certainly is a problem and it's common for C applications to take on responsibilities for manipulating the internals of the string structure.
No I’m talking about the classic byte arrays C pretends are strings. If the receiver you are talking to is not expecting a chain of bytes where each byte is a character and one special value means the chain is over, well, you will have to transform what you are sending to what your receiver is expecting.
"I think C has a lot of features that are very important. The way C handles pointers, for example, was a brilliant innovation; it solved a lot of problems that we had before in data structuring and made the programs look good afterwards. C isn't the perfect language, no language is, but I think it has a lot of virtues, and you can avoid the parts you don't like. I do like C as a language, especially because it blends in with the operating system (if you're using UNIX, for example).
All through my life, I've always used the programming language that blended best with the debugging system and operating system that I'm using. If I had a better debugger for language X, and if X went well with the operating system, I would be using that."
Dec 7, 1993, Computer Literacy Bookshops Interview.
Nobody "worships" C and people who have positive experiences with C often have exposure to higher level languages.
You can find strengths or upsides in C and still acknowledge faults, and still acknowledge merits elsewhere.
We can make a more complicated example that nevertheless uses recursion and essentially the same way, which looks for a UTF-8 character. We can examine and decode a prefix of the string as a UTF-8. If there are bad bytes or the character doesn't match then we recurse.
We can also write a wide character (wchar_t) version of the function which looks the same. That will handle all of Unicode on sane platforms, and the basic multilingual plane (BMP) on Windows.
To be fair, the most straightforward definition of the linked list is generic, and C completely lacks such facility.
> fascist-
Sounds familiar.
Template mechanisms cause code bloat and encapsulation violations: having to reveal the entire implementation to the clients so they can instantiate it for whatever type they want.
Some generic mechanisms have dumb restrictions, like you can make a "list of integer", whose elements, sadly, all have to be integers.
In C you can envision what you want genericity to look like at the detailed bit-and-byte memory level, and how you'd like to be able to work with it at the language level, pick some compromises where the two are at odds, and make it happen.
It is false to say that C completely lacks any generic facility because it has union types. Using a union we can put several types into a structure along with a type field which indicates which. We can overlay an integer, character, floating-point number, or pointer to something outside of the node.
Generic linked lists can be done with pointer casts. See the way the Linux kernel does it.
The list manipulation routine takes a pointer to a member, and the more specific type is opaque. The code that knows what the actual structure is can derive it from a node pointer.
A structure can even have multiple node pointers in this scheme.
I cannot decide whether it's an honor or an insult to be called a _fascist sympathizer_ for opposing the authoritarian suppression of speech.
As proven by books like "A book on C" from 1984 (Robert Edward Berry and B. A. E. Meekings).
The question is about C semantics. Given the reply I get it’s pretty obvious that some here don’t understand what language semantics are. It’s about the amount of concepts you can express in the language. Haskell - a language I personally despise - is semantically very rich. So is modern C++ for what it’s worth. C simply isn’t.
It’s even deceiving sometimes because it has the apparence of having some semantic elements (arrays for exemple) which are not there in reality and are really only syntactic sugar on top of other semantic constructions(pointers).
Also you complain about semantic richness of C, but then only point to semantically rich languages you despise. A curious reader would wonder: why care about semantically rich languages then? An observant reader would wonder: so aren’t there times where semantic richness is not useful?
I’m not even complaining about the semantic richness of C. I’m just stating the fact that it doesn’t really have semantic richness. Then again don’t get me wrong. I do think C is a terrible language and I say that having worked on a C static analyser. It’s full of avoidable undefined behaviours and silly sharp edges.
Thankfully, nice languages with rich semantic exist. I was just pointing some I don’t like to separate the issue of semantic richness from likability. I enjoyed working in Ada a lot. Ocaml is awesome. I have never used Rust but from what I have seen that seems nice.
Semantic richness is useful because it helps programmers express what they want to do in way which are clearer and therefore more likely to be correct.
The successor to Pascal, which fixed it's problems, was Modula-2, which is vastly superior to C. Unfortunately by the time it came out and started becoming available, it was too late for it to compete fairly with C. Ada is comparable to Modula-2, especially before Ada added OOP, but Ada was too late to compete with C as well.
It's taken 30 years but finally we are starting to see safe languages, eg. Rust, compete with C and C++. I don't know whether Rust will be the language that overtakes C and C++, but if not, I believe a successor to it will and C and C++ will become like COBOL, that no ones uses to write any new code.
Many of the so-called issues with Pascal were resolved way back in the mid to late 1980s. In regards to Pascal's competition with C, this arguably evolves around AT&T and Unix. It's AT&T and their huge government, industry, and business influence that pushed C over the top.
One small thing I like about Modula-2 is an improvement in the syntax where BEGIN is not required after conditionals and loops.
The nice thing about C is that if you're opinionated about what you want your ideal language to look like, you can use C in that pursuit. You can write your run-time in C and parts of the implementation. C is good for interfacing with operating system platforms.
You can fully bootstrap yourself off C entirely, or stick with C for the run-time and perhaps other parts. If your language isn't very good at expressing low-level manipulations, you can keep C in there as an option for writing those kinds of modules.
Some languages have used C as a target language for compilation.
C can be an enabler in the pursuit of other languages.
I wish that C had a more rich way to define struct layouts and low-level representations for integral types.
I really like how Ada does it:
https://en.wikibooks.org/wiki/Ada_Programming/Representation...
If you need to read or write specialized hardware registers, being able to define a data structure with a custom representation is very nice and can save significant time and effort.
I identify with @antirez's sentiment. I know both C and C++ very well. I find C definitely more artistic. C++ is utilitarian.
Prolog is the other language I code in "artistically". It too is quite simple and constrained, though it is on the opposite end of the high/low-level spectrum as C.
/unpopular-hot-take
I've never used Ruby but from what I've seen I understand it to be similar to Perl in that regard.
Of course good naming and abstractions are of key importance, and comments and inline documentation are important finishing touches, which so many programmers don't seem to have the time for. It doesn't hurt that Larry Wall, the designer of perl, was a linguist. Perl was meant to be flexible and expressive.
And everyone knows that there's nothing worse to read than obfuscated perl. Perhaps some think that is artistic in a different way, to make the shortest and most unreadable program possible. Not my thing, personally.
I would argue that C++ can be dramatically more deceiving than C --- see inheritance and operator overloading, just to name two.
If you can just extend the types then you can provide operators that way and there's less opportunity for ambiguity. See Rust.
Also so many of the C++ operator overloads are broken instead of just not existing, which tempts you to overload things instead of saying "Nope, that's a bad idea" and just walking away altogether.
Example, boolean short-circuiting AND and OR. If we write
if (this() || that()) foo();
in C++ then that OR is short-circuiting, so that() won't get called if this() is true.But if the return type used overloads the boolean OR operator, short-circuiting is disabled, and now both this() and that() are always called...
Note how terrible doing any amount of math is in C.
I don’t understand that. You can have functions that return promises that then can get passed to other functions returning promises, and leave it either to the compiler or to an expression evaluator you write (that ideally runs at compilation time as much as possible) to optimize away anything not needed. For example (pseudo-code)
vector3D a = …
vector3D b = …
promise<vector3D> c = addLazily(a,b)
print c.x
in the end, could do the equivalent of print a.x + b.x
I think you could make a modern C++ compiler do that for this simple example.Also, there’s a simple bijection between expressions with operators and two-argument function calls:
a + b * c
+(a, *(b,c))
plus(a, times(b,c))
Because of that, I don’t understand why operator overloading should give better optimization opportunities than function calls.The books you have are excellent ways to learn the language, and this one complements them in that it lists common practices around, for example, error handling, and other common topics. It shows different approaches and mentions practices from real-world open-source C projects that follow them.
It seems like a good book to continue exploring C after learning the language.