I still like C and strongly dislike C++
codecs.multimedia.cx
codecs.multimedia.cx
There's an argument to be made that C is a simple language compared to some, I'll grant that. But implicit casts and the integer promotion rules really aren't a good example of simplicity. Terseness, sure, but not simplicity. For example, from https://www.cryptologie.net/article/418/integer-promotion-in...
uint8_t start = -1;
printf("%x\n", start); // prints 0xff
uint64_t result = start << 24;
printf("%llx\n", result); // should print 00000000ff000000, but will print ffffffffff000000
result = (uint64_t)start << 24;
printf("%llx\n", result); // prints 00000000ff000000
There are good reasons why e.g. Go, which prioritizes simplicity very highly, does not allow implicit integer casts.If you write code like that in C frequently (e.g. working with system registers on embedded CPUs), you do eventually get used to explicitly casting defensively, and I agree that it's nice that you get saner defaults in e.g. Rust.
Seriously, are you talking about C, a language where you have to write code to access any matrix optimally by hand always... (or an advanced library to do it, better than BLAS, which ffmpeg does not even use)
I feel the same whenever somebody shows up talking about their favorite functional language. Sure it can be less code, but it looks like an explosion at the unicode factory.
Really though, I'd much rather have a few language "gotchas" that you need to learn once, over countless gotchas in other developers code because they are using a language the lends itself to ungrokable code. C really does have the properly where you can look at the code and just know what it does.
That said, I don't think C scales very well for big projects and I personally develop in C/C++/Python/ environment as a robotics system SWE. C for embedded stuff, C++ for nearly all system code, and python for the data/Deep learning stuff.
If you can spot undefined and other platform defined behavior always, someone probably would want to hire you as a security specialist.
uint8_t start = -1;
uint64_t result = start << 24;
however there's some hope of defining it once C requires 2's-complement integers. (this wasn't reasonable in the past but is now)There are worse problems with `a << b` because the shift instruction on CPUs doesn't actually treat b as an 'int', it reads the lower X bits and uses that. C can't define what X is because not only is it different in x86 and ARM but x86 doesn't even match itself - it's different for scalar and SSE ints.
uint8_t start = -1; // Your code
uint8_t start = 0xFF; // Absolutely equivalent
As for the expression `start << 24`, `start` will always be promoted to `signed int` because `int` is guaranteed to be at least 16 bits wide, so it can hold all `uint8_t` values.If the width of `int` is 32 bits or narrower, then `start << 24` (which is 0xFF000000) will put a binary 1 into the sign bit, which is indeed undefined behavior.
Otherwise if `int` is 33 bits or wider, then `start << 24` will fit in `int` nicely and be fully defined behavior.
Reasoning about these complicated cases involving implicit conversion (-1 -> 0xFF), promotion (uint8_t -> signed int), and implementation-defined bit widths (how wide is int?) is exactly why I dislike dealing with C and C++.
The idea that C programs can survive having the size of their most common type changing is not realistic, though luckily people rarely try.
As for writing code that behaves correctly on any legal C implementation, it is certainly possible and I practiced this for years. Go ahead and audit my published programs. It is not easy though, and I have made many mistakes along the way before I figured out how to do it consistently. Although there are many rules that I need to follow, the two most important ones are: Obey the fact that built-in integer types have minimum widths (e.g. char is at least 8, int at least 16, long at least 32); Use size_t for array lengths and indexes.
In my experience, I run into "undefined behavior" in C-land extremely rarely. I can't remember a single instance in 15 years of embedded programming (unsigned/type promotion sure everybody does that at least once sometime, some weird architecture defined thing, nope). Maybe that's because I'm some super amazing C programmer, or that I read the docs, or maybe it's because the chance is low?
Regardless, people very often become hyper-focused in this types of debates to be point of being completely unbraced by reality. The valid way to do these sorts of language comparisons is to look at the the pros and cons and give them a weight. Something like this pro is (very good), but (rarely comes up in practical programming), or (This con is very bad) but basically only shows up in (interview questions), etc. Then take the sum of the weight of each side, rather than making statements like, "language has/doesn't have x feature/drawbrack thus I'm right about every thing yay!". I.e. if you are looking at a 10d problem in 1d, you're going to be wrong.
int32 a;
int16 b;
//...
b = (int16)a;
Later you decide to change b to int32. int32 a;
int32 b;
//...
b = (int16)a;
Now you have a bug. If you didn't have that explicit cast the program would have worked as expected. b = (int16)a;
This contains two casts... one explicit cast to int16, and one implicit cast back to int32. In languages without implicit casts, this would be an error.I don't think it makes sense to ban them in a language with powerful enough checks, like dependent types. If you have two variables and one is typed "int 1-10" and the other is typed "int 1-5", and it fails at runtime if it doesn't fit in the destination range, why do you need to type something that says you really meant to assign one to the other? Just assign it.
I'm reading this as, "Many red boxes are blue circles." If the language allows implicit widening casts, it has implicit casts.
I don't think more powerful checks are necessary. It's just that the implicit conversions in C are a bit wild, and they result in unexpected behavior and surprise programmers, and from experience, making all casts explicit is not such a burden (except for stuff like char ptr -> const char ptr).
I think the fantasy here is simple... as much as we want to explore new ways to make safe programs with better type systems and runtime checks, there is still some design space near C which is safer but not really more complicated.
There seems to be an uncanny valley effect with C-like languages. The best you can do is have non-standard compiler extensions on top of ISO C.
We're not talking about making an entirely new language which is incompatible with C. You adopt a style guide, an accompanying static analysis tool, and maybe add some annotations to your source code. Your code is still C-compatible and you can still use the same compilers.
Static analysis of general C programs is quite difficult, and turning on an analyzer for an existing codebase typically results in obscene numbers of false positives. However, if you pair a static analyzer with a strict style guide which restricts the language, you can make the static analyzer much more powerful and useful.
This could be MISRA, it could be a formal verification system, or it could be something else entirely.
The problem is that people state things like "Java doesn't have implicit casts for correctness." But then it does have implicit casts, so now is there still correctness?
Another case is Haskell, where the wiki tells you:
> Conversion between numerical types in Haskell must be done explicitly. This is unlike many traditional languages (such as C or Java) that automatically coerce between numerical types.
But what this actually means is instead of `a = b` you do `a = fromInteger b`. This is obviously also an implicit cast, because Haskell has type inference. So again, you're not writing a proof that your conversion is correct. You're just writing a different, longer "yes I really mean to assign this" statement.
I am a bit too tired to engage with this kind of outright insanity. If you have a problem with what "people state", but not specifically with what I state or what has been stated in this conversation, then go have arguments with "people".
The idea that "Java has correctness" or "Java does not have correctness" does not make any sense. Java is a language, it's the programs that you write with it that are correct or incorrect.
> This is obviously also an implicit cast, because Haskell has type inference.
Incorrect, it is an explicit conversion. There is no such thing as a "cast" in Haskell, there are only functions which convert values of one type to values of another type. (well, in FFI code, there are functions which "cast" pointers and the like, but that's FFI.)
And yes, you're not writing a proof that the conversion is correct. This is... blindingly obvious, so I have no idea what kind of point you are trying to make. "Ordinary Haskell code is not formally verified" is not news.
Touchy! I am claiming the people who want always explicit casts have not actually used a system where they're all explicit much (because like in Java, some are still implicit) and would find it annoying if they did.
That could result in performance compromises, like using int arrays everywhere instead of smaller types. Which is fine for scalar values, but for arrays it wastes memory.
> Java is a language, it's the programs that you write with it that are correct or incorrect.
Surely this language feature is intended to promote correctness. What else are errors for?
> Incorrect, it is an explicit conversion.
But `a = (uint16_t)b` and `a = fromInteger b` aren't the same thing - in one you have to name the destination type and in the other you don't. One is more explicit than the other. Is one of them bad?
> "Ordinary Haskell code is not formally verified" is not news.
It is news to people who don't do numeric programming. "If it compiles it probably works" / "if it compiles it's probably correct" is a real thing said about the language said on this forum.
C clearly has issues (the implicit unsigned short->int promotion is totally wrong) but I think trapping on overflows at runtime would be a much better improvement than adding compile errors.
Yes. "Probably", as you said. "Probably" is much, much weaker than "formally verified".
Hey, I'm only human. You made some inane comments.
> I am claiming the people who want always explicit casts have not actually used a system where they're all explicit much (because like in Java, some are still implicit) and would find it annoying if they did.
That claim is trivially false... Go and Rust have explicit casts, and they're reasonably popular. For example, in Go,
var x int16
var y int32
x = y // ERROR!
x = int32(y) // ok
The same is true in Rust. let x: i16 = 1;
let y: i32 = x; // ERROR!
let y = i32::from(x); // ok
let z = i16::from(y); // ERROR!
> That could result in performance compromises, like using int arrays everywhere instead of smaller types. Which is fine for scalar values, but for arrays it wastes memory.This is clearly false in practice, just look at extant Go and Rust code.
> But `a = (uint16_t)b` and `a = fromInteger b` aren't the same thing - in one you have to name the destination type and in the other you don't. One is more explicit than the other. Is one of them bad?
If that's a rhetorical question, just make your point.
The fromInteger function is an explicit conversion. The conversion is explicit, but the destination type isn't explicit. The source type isn't explicit either, but apparently that's okay... consider the C code:
int x = (int)y;
Would you call the cast "implicit" because you don't know the type of y just by looking at the code? No, it's still an explicit cast. Maybe a "super duper mega explicit" cast would look like this: int x = (float->int)y;
Reminds me of Scheme. You would want super duper explicit conversions in dynamic languages, because neither the source nor destination type are annotated. You want the destination type annotated in traditional type systems because otherwise the compiler would not be able to figure out what you are doing. Haskell does not need the source or destination type annotated, this is ok.> It is news to people who don't do numeric programming.
I'd say that it's blindingly obvious to people with a rudimentary understanding of programming.
> C clearly has issues (the implicit unsigned short->int promotion is totally wrong) but I think trapping on overflows at runtime would be a much better improvement than adding compile errors.
You can already -ftrapv if you like, but "trapping on overflow" is a contributing factor to the Ariane 5 disaster, and the lesson there is that trapping on overflow is not necessarily a safe default. The cost of runtime overflow checks is surprisingly large, too, which is why people so often turn it off in languages that support it out of the box (like C# and Ada). The mitigations for these problems will involve some combination of run-time checks, compile-time checks, and testing.
Can you link some examples that weren't obviously at least 50% tongue in cheek?
Weeell, maybe not. I agree that whenever I see a cast, I think "WTF?", but maybe you want what the cast does?
b = static_cast<decltype(b)>(a);
And as other commenters have pointed out, if implicit widening is disallowed that would also prevent the potential bug.If what you want to express is "I know b and a have different storage sizes", that could be useful, but it doesn't do that because eg if you typo `b` for `a`, then `b = static_cast<decltype(b)>(b);` is still accepted.
Sure, but at least the problem is back to "should conversions/casts be implicit or explicit?", rather than "I cannot do what I want with explicit casts".
> If what you want to express is "I know b and a have different storage sizes"
You may not even know this (e.g., in generic code), and if you did want to ensure different storage sizes specifically, then you want a (static) assert.
> eg if you typo `b` for `a`, then `b = static_cast<decltype(b)>(b);` is still accepted.
That would also apply to the version with implicit casts/conversions, wouldn't it?
(On balance, I personally think that not having the implicit conversions, as Rust and Go do, is generally the better decision, but I certainly acknowledge that C is easier in this regard.)
Really? I think there must be a better way - I also don't like implicit conversions, but the rust approach gets extremely verbose, especially if you want to check for truncation etc.
What I believe most people would expect is that 0xff would get shifted up, and that's it, but instead you end up with that plus a sign extension to fill out the "new" 32 bits at the top of the 64-bit value.
How it should work is that type of "start << 24" should be synthesized from the operands, without any idiotic "promotion" rule: it should therefore type uint8_t.
The shift should then require a diagnostic that it exceeds the width of the type. No undefined behavior nonsense; if the shift amount is a constant then it is statically diagnosable against the width of the type, and therefore should be.
This diagnostic will inform the programmer that their intent isn't being expressed.
An alternative safer model called as-if-infinitely ranged integers solves this problem:
https://resources.sei.cmu.edu/library/asset-view.cfm?assetid...
SHL 1, AL ;; x86 assembly
will not affect any part of the EAX register other than the low 8 bits specified by AL. This is very simple and obvious.An operation involving uint8_t staying entirely in that type is a complete no-brainer.
If you want to shift bits out of the uint8_t, you cast the value to something wider, and shift that.
And note that this is what you have to do when working with types that are as wide as int, or wider, which do not promote. uint32_t << 1 will lose the top bits!
Losing promotion would make everything consistent: uint8_t << 1 loses the top bit the same way like uint16_t << 1 and uint32_t << 1.
"As if infinitely ranged" (AIIR) is an obvious idea but it has flaws. One is the run-time checking that it requires, which is unattractive in a C-like language. The other is semantic issues. In a language like C, the values of calculations eventually have to settle into typed boxes. Under AIIR, these are not equivalent:
int x = 42;
int y = x * x + 3;
int z = 2 * y / x;
int z = 2 * (x * x + 3) / x;
In the second version, we have not bound the (x * x + 3) term into a variable, but inlined it into the z calculation. Therefore, its result of that subexpression permitted to be outside of the range of int by AIIR.This is completely unacceptable, though; simply introducing a variable of the precisely correct type to hold the value of an intermediate calculation changes the semantics of the calculation.
Only the values being juggled in the AIIR, so to speak, benefit; not values that have landed.
It seems to be acceptable in Swift so far. (Swift traps on overflow on all expressions though, it doesn't use AIR.)
It is unpleasant if you don't like to distinguish between statements and expressions, but it might be usable. Anything is as long as it's well defined.
The main advantage is that it makes it easier to fold away intermediate steps like (x + 1 - 1). This is the kind of thing that looks pointless, but is needed to remove abstractions like in macros or C++y generic code.
If you'd like your functions total (not trapping), then there could be a set of wrapping/saturating operations to go with the default infinitely-ranged ones. Or you could declare x/y/z to be unsigned.
It's odd for the author to complain about the commingling of C and C++ in compilers and then complain about Microsoft specifically not doing that.
That is the official excuse to only support C on Azure Sphere, despite the whole security sales pitch.
Scientific and Engineering C++ by Barton and Nackman.
Multi-paradigm design for C++ by James Coplien.
I have not as yet found any equivalents for "Modern C++".
I've been working primarily in PHP for a while now (and honestly put a lot less value on language comparison) but from what I've seen the expressiveness and power of C++ (at least in the ways I like it) have mostly been superseded by Rust and I think that'd be where I started if I ever need to work on something where in application logic would actually be a product bottleneck (instead of data source interaction).
Frostbite or Dawn Engine with its bad editors and performance problems?
Crytek used by few games? (This one has a chance to challenge UE) Unigine which is also rather rare?
For example if I wanted to make a 3D platformer in unity, I could probably commission an artist to create the character, download a starter kit of some sort, and be done with it in about a week.
However this means I'm tying in so much code, and if anything doesn't work exactly the way I expect it I can complain. I might complain about the starter kit I purchased, I might complain about Unity crashing when the real reason for the crashes are my own crappy code.
Particularly when I was younger, I would often try to do way too much at once and this leads to frustration. But, theoretically you could create the next gears of war with three other people in about a year. Odds are you're going to run into tons of problems though just because it's so hard to do that
this is simply not true
It's why I like zig so much, It really is the language that understands what was wrong with C and fix's those parts and doesn't do too much more.
Macros suck -> comptime is just zig code Pointer or array is ambiguous most of the time -> array and ptr notation Error handling things with a lot of possible errors leads to goto's or messy clean up -> errdefer and defer make that clean.
Here is my take. I like C when I work on firmware for less powerful microcontrollers. I like modern C++ when I work on backend servers / middleware. I like Delphi / Lazarus when I work on desktop software. I like Java Script when writing web front-end libraries / apps. I like Python when shell script gets a bit too complex. Etc. Etc.
The real truth - I do not really like any of those. They're just practical tools that help me build my product. Product designed and created by me is what I like.
ALGOL dialects for systems programming had it 10 years before C was born.
CPL the language that BCPL subset was designed as bootstrapping compiler, had it.
PL/I the language used on Multics had it.
PL.8, language created by IBM for their RISC project and LLVM like compiler toolchain, had it.
As did plenty of others.
> And in most cases you know what will compiler produce—what would be memory representation of the object and how you can reinterpret it differently (I blame C++ for making it harder in newer C standard editions but I’ll talk about it later), what happens on the function calls and such. C is called portable assembly language for a reason, and I like it because of that reason.
When the target CPU is an old 8 or 16 bits CPU, and the code gets compiled with optimizations turned off.
> Avoid repeating tired maxims like “C is a portable assembly language” and “trust the programmer.” Unfortunately, C and C++ are mostly taught the old way, as if programming in them isn’t like walking in a minefield.
I don't care about RAII, so I don't care about the ownership tracking that Rust provides. I like to think that memory management is actually part of the problem, rather than a responsibility for the language's runtime.
I want to generate hundreds of datastructures or algorithms at compile time that are optimal in specific situations using introspection of types.
I want everything else in the language to get out of the way and know that I'm writing code that actually runs on a real computer, not an abstract machine.
We could use more safety even, like automatic race condition analysis in language and more.
Conservative extensions of C and C++ exist, and they're not very popular, just check how popular D is...
Which features do you recommend we deprecate?
As a comparison, has Python decided what it wanted to be? A scripting language like Perl, data processing language like R, tool command language like TCL or web backend programming language like PHP or RoR? But nobody in their right mind will ever say Python need to decide on his direction when it grows old.
As you probably have known when Python is around 20 years old (like D today), it played second or third fiddle to PHP, TCL, Ruby (RoR), Perl and R in their respective domains, and at the time the growing pain into Python 3 is yet to happen. But looks where is it now?
Personally I think the D language foundation is already solid [1]. It just need a killer application for it to be more popular and well-known just like what RoR did to Ruby. And if you have a solid foundation, it will probably just a matter of time for it to properly take off.
While they don't know where it should go, C#, C++, Swift keep adopting D like features, with their rich ecosystem.
D could have been C# on Unity, given Remedy experience with the language, instead not.
D's Metaprogramming was great 10 years ago, not when placed against C++20 metaprogramming, or the features arriving in C++23.
Programming languages used to stick to one programming concept for examples either imperative, functional or object oriented. But now most modern languages support at least more than two concepts to stay relevant.
Regarding other languages later adopting D like features, sure you can do that but the end results will probably be sub-optimal and clunky. For example, Python adopting array processing in Numpy library and becoming very useful and popular, but it will not be as seamless, intuitive and natural as R or Fortran array processing capabilities.
Phobos still doesn't fully work with @nogc, DIP1000 is undocumented and now there is @life getting the spotlight, meanwhile the GC is stuck in early 2000 design, the std.allocators library is in experimental limbo for years.
Even if the results are cluky, they can double down on an eco-system in libraries, IDE tooling and OS support, that D lacks.
Currently the motto is Jack of all trades, master of none.
Also DIP1000 isn't that badly documented these days, not brilliantly, but frankly I find that a lot of complaints come from people who weren't ready (this doesn't necessarily apply to you in particular, just that these attitudes spread) to understand it in the first place.
And what do you mean by OS support? I find that not many other languages take it as seriously as us?
Taking C# as example, I mean being able to do embedded stuff (Meadows, IoT Core), regular desktops, IBM mainframes, game consoles, and mobile OSes, without having to create my own compiler toolchain and self made druntime.
Plus, inheriting abstract interface classes is more convenient than writing a vtable by hand when such a thing makes sense. You don't have to participate in Java-style class voodoo if you don't want to.
I am sure all the other languages mentioned like Rust, Zig, and D are very good too, but C++ is where all the high-performance libraries live, and whatever productivity gains I would get by switching languages is utterly dwarfed by that.
https://github.com/attractivechaos/klib/blob/master/khash.h
https://github.com/attractivechaos/klib/blob/master/kvec.h
FWIW writing a generic vector like that ^^^ is going to be faster than almost any std::vector implementation you can find. And that is just one way to do it, there are other tricks.
The actual logic of a B*-tree would be equally hard in any language. The macroized template part is easy.
Even if it were simple, programs coded in C get no benefit from that: the language's limitations mean that the program has to provide whatever the language does not.
Thus, a program in C++ will be much simpler than the corresponding C program because you get to lean on what C++ provides. And, let's face it, practically all the complexity we encounter is in program code. The less there is of it to read, the less there is to understand.
This seem to be arguing that C++ is accumulating cruft as it grows and it's bad, but I like the fact that C++ tries hard to maintain backward compatibility. The greatest asset of a language is all the stuff that is already written in it.
That said, some outdated stuff do get thrown out, such as trigraphs.
In fact, I'm not confused with simple function types. It is just that sometimes, complex function types are convolutions for me.
#define a stdio
#define header <a.h> // five tokens: <, a, ., h, > (otherwise preprocessor wouldn't substitue "a")
#include header // one token after substitution: <stdio.h>
int main() {
struct {
int h;
} stdio;
a.h = 0;
printf("%d\n", stdio.h);
}
What happened here is:1. In #include directives a thing like <stdio.h> is one token; everywhere else (including other preprocessor directives) it is five tokens;
2. But an #include directive can have multiple tokens on the right hand side of "#include" before macro substitution. After macro substitution the strings are lexed again - into ONE token.
3. In main() "a.h" is substituted with "stdio.h" - this time it's THREE tokens.
(Oh, and writing "#include <a.h>" will produce an error, because here we have a single token <a.h>, so no "a" to be substituted, and normally there is no "a.h" header file in include path.)
So the same string gets lexed into a different number of tokens in different places; and that happens even after it got lexed first into the same number of tokens before macro substitution.
Now have an example of a C file that compiles with GCC, but not with Clang:
#define CMPS x /*
*/ <y>
#define ALMOST_SHIFT<CMPS
#define FWD_FST(x, y) x
#define CONCAT(x, y) x ## y
#define x
# /*
LOL
*/ define y stdio.h
#include CMPS
#undef x
#undef y
#include FWD_FST(<stdint.h>, y)
// btw. comma is a valid character in a header name. but this will work anyway,
// because lexing in C works in mysterious ways
#include CONCAT(<std,def.h>)
int main() {
int x = 0;
int y = 0;
int z = 0;
printf("%d\n", CMPS z);
// printf("%d\n", z<ALMOST_SHIFT); // doesn't compile
// but the rest works fine...
}
So clearly the C syntax isn't simple. It's so not simple, that compiler writers can't agree on what exactly it is.(I've been writing such programs for the purpose of using them as edgecases to test my C compiler.)
Then of course you have the type declaration fiasco: <https://blog.golang.org/declaration-syntax>
Then there if of course the awful ternary operator.
Oh, and I almost forgot: spaces are insignificant... unless you're in a "#define" directive. Talk about language consistency...
Those things not only make life hard for compiler writers, but also for language users, because, I don't know about you, but when I write a program the smallest unit that I use to think about its source code is the token. When I can't easily predict tokenization, then something went really wrong. Same with other syntax complications. Guy Steele once said that when he designs a language he verifies that its grammar is LL(1) with a parser generator, which helps make sure that it's not only easy to write the parser and that parsing can be very efficient, but also that it's easy to understand by humans. C fails here big time.
The preprocessor is part of the language. Both in literal sense and in real-world. All C programs use the preprocessor. You need to deal with it. Also, the grammar of the preprocessor is intertwined with C proper due to "#if" directives. The expression on the right hand side of an "#if" is parsed according to the grammar of C proper.
> I'm not sure that anything you say about it is true (the preprocessor is complicated).
You can run the examples under compilers and verify. Also, the part about it being context-sensitive (with regard to tokenization in #include directives vs everywhere else) is explicitly noted in the standard. Nothing controversial here.
> Certainly, your first example won't compile without warnings for any sensible GCC options.
Compiles just fine without any warnings with GCC with options -Wall -Wextra.
I've written several C-style preprocessors. They can be fiddly to write, but you shouldn't think of them as lexing to the same lexemes as a C parser needs. You can take a lot of shortcuts there.
Things like '<' and '>' being overloaded is quite common, even in lexers. You see it with nested templates / generics in C++, Java, C# etc. There are simple tricks for dealing with the ambiguity between '>>' vs '>' '>' in 'T<U<V>>'.
(FWIW, having written some of those preprocessors, I think they're a bad idea. It's notable that other languages haven't copied the idea, other than C++ which inherited it.)
> Unlike most modern languages, which have a context-free grammar, it has a context-sensitive grammar.
Implies it C does not have simple syntax.
Is there a law a language with "simple syntax" cannot have context-sensitive grammar?
The specific problem is this has two parses:
foo * bar;
It's either an expression multiplying foo by bar; or, if foo is a type, it's a declaration of a variable foo pointing to values of type bar.The normal way this is handled is to inform the lexer of known type names (e.g. by letting it peek at the symbol table) whenever it parses an identifier, so it can produce a type name token instead.
But generally I disagree that preprocessor is separate from C for the following reasons:
1. It is specified by the C standard.
2. Different C compilers implement it differently, just like they implement other parts of C differently.
3. There is some truth in the claim that it is a separate text transformation, but here I also do not agree entirely, because the C standard doesn't say that the preprocessor outputs text. It says that it outputs tokens and those tokens are then converted to tokens of C proper[1]. It is also clear that this model of "text transformation" is also not entirely what happens in real world compilers, because if it did we wouldn't get good backtraces for tokens that went out of the preprocessor after substitution. Maybe GCC serializes it to text at some point, but clearly it gets more information out of it than what a naive interpretation of the phrase "text transformation" would suggest.
[1]: Translation phase 7 of the C11 standard:
> Each preprocessing token is converted into a token.
$ brew install cling
$ cling
****************** CLING ******************
* Type C++ code and press enter to run it *
* Type .q to exit *
*******************************************
[cling]$ #include <stdio.h>
[cling]$ #include <sys/utsname.h>
[cling]$ struct utsname u;
[cling]$ uname(&u);
[cling]$ printf("%s %s %s\n", u.sysname, u.release, u.machine);
Darwin 20.5.0 arm64 [cling]$ int i=21
(int) 21
[cling]$ i*2
(int) 42 [cling]$ u.sysname
(char [256]) "Darwin\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0\0" [cling]$ u.sysname
(char [256])The solutions are either to use a C++ std::string as you mentioned, or a typedef struct with a length property.
I should finish writing a big article title "Why I still don't like C++" for my yet unpublished blog...
Even if I tend to use orthodox C++ for convenience.
But I've been impressed lately by some newcomers, like Zig.
There are situations that you will not interpret correctly no matter how much time you spend learning the language.
It's a serious flaw in the language allowed because it makes writing optimizers and compilers easier.
I've never run into a bug in Java wherein a look at the compiler or runtim error, runtime stack and some look at StackExchange I couldn't solve.
But I'm constantly fighting CMake, linker issues etc..
I avoid huge swaths of C++ because I consider them to risky to deploy without mastery.
For bonus points we can include the Java implementations outside OpenJDK.
I'm having difficulty compiling a Qt app right now due to complicated and arcane marco directives. The compiler / linker give me no information to work on, and there's no documentation.
It takes 2x as long to develop in C++ it's grinding.
C++ is also explained and documented as much as Java is.
As for building, try to use Makefiles, Ant, Maven, Gradle when only being comfortable with one of them.
Or configure the builds across all major IDEs.
E.g. Gcc and Gdb are nowadays compiled with G++, and new code in them is C++. It is very easy to start this process in just about any C program, so it happens many times every day, all over the world, with little notice because the bump is unfelt.
If it is meaningful at all to talk about a language going away, it is about a decline in writing new code in it. Obviously all the code written in badly obsolete languages still exists, somewhere, but the demand for people to write any more of it falls at an increasing rate.
If you like C, you should really try Zig. Same simplicity and control, but really nice improvements to make code easier to read and maintain.
[0]: https://blog.thecybershadow.net/2018/11/18/d-compilation-is-...
[1]: https://forum.dlang.org/thread/pvseqkfkgaopsnhqecdb@forum.dl...
I used to think that using C++ was acceptable in some cases. I no longer do.
The faults isn't even C++ the language, per se. It's the ecosystem--or more appropriately--the lack thereof.
A lot of the modern languages integrate with C, properly. They integrate on Linux, Windows, OS X/macOS, iOS, Android, etc.
Nothing integrates with C++ well, and it's the fault on the C++ side. That means dumping C++. Perhaps the C++ compiler writers will finally have some incentive to fix their crap.
Or just integrate with C++ code by exposing it through C linkage with `extern "C" {}`. Sure, C is the de facto lingua franca of programming languages, but that doesn't seem like a pragmatic reason to choose C over a different language that may be better for the job. Also, C++ has a massive ecosystem.
Nope. Been there. Done that. Got the scars.
For example, what happens when an exception gets thrown on the other side of that "extern C"? Yeah, undefined.
The problem is that if you're using C++, you get the C++ machinery--allocators, exceptions, etc.--and you can't dodge it.
Not even wrong !