Initialization in C++ is Seriously Bonkers
mikelui.io
mikelui.io
No, it's because i lives in an area allocated at the OS level, and OS has to have initialized that memory with something (otherwise you'd have a security bug). The .bss segment has been zeroed, for obvious reasons, since the advent of modern linkage.
The explanation is backwards. The reason that stack variables are UNinitialized is (contra the article which thinks it's because the programmer didn't put an initializer in the source code) that memory is on the stack, which is allocated internal to your program and in practice was used previously by some other call for some other purpose.
The fact that stack variables are uninitialized by default is actually an intentional feature, not a bug, as it correctly expresses the behavior of the runtime environment. Now, it may not be a good feature, and modern language designs have omitted it (by, let's remember, relying on much better modern compilers to elide all the extra code that would otherwise have been needed to zero out a stack frame!).
But it's got nothing to do with the syntax of the language. You can rely on .bss being zero in assembly, and similarly be surprised that stack memory is going to have junk in it.
Anyways, the problem C++ has for beginners is that the language tries to obscure the already hard to understand low-level behavior with high-level concepts (to be fair, that is the language's ultimate goal: to provide a high-level abstraction for low-level code.) So maybe good for your second/third language, but definitely not your first.
Some believe that it's important to teach those things about computer systems first. To anyone who knows how the machine works it's pretty obvious why a primitive stack variable is uninitialized: initializing them would cost a store which would in most cases be pointless. People use C and C++ because they want all of the performance the machine has to offer, and usually more. Stuffing the program full of immediate value loads and stores would make it huge and slow. C++ is already cursed with enough pointless stores, for example initializing the buffer of a string just after resize and just before copying something else into the same memory. We don't need more of these things.
The author then goes on to complain at length about {} initializers in C++11, which everyone recognizes as problematic, and which has been partially fixed in C++14. This has basically nothing to do with the first part of the article.
But for a first language I'd want to use something like Python or Javascript or even a toy language. It simply needs to be something where some basic, high level concepts can be taught without worrying about the details.
It's not even really hard, just crushingly tedious and error-prone. You could also write all your programs on 80-column punch cards or with (1)ed, but why bother if there isn't some technical reason why a text editor isn't available?
And the big advantage is that when you understand those concept and pointers, there's no magic in other languages (value/reference parameters, objects, functions as first-class citizens)
C and C++ are horrible academic languages for other reasons - the original design is just too much of a hack to begin with, and then there's decades of legacy backwards compatibility with it in all newer developments. Syntax is ugly and inconsistent, doing things right is often harder than doing them wrong, there are many features and exceptions that are only there for historical reasons, and standard library is very eclectic in terms of what is and isn't included (from a modern perspective). But take something like Modula-2, and pointers aren't a problem.
For your amusement: in the Linux kernel, at least on x86, .bss isn’t initialized when the kernel is loaded. Instead, the kernel memsets it to zero a little later. I don’t know why.
Also, because all this stuff predates any concept of security, there is no read-only equivalent of .bss in most systems, meaning that you get suboptimal code for:
const int i = 0;
On OS it's done by the OS for security and performance reasons. One you don't want people to snoop on memory freed from other processes. Two the OS can use the MMU to map in previously zero's pages of memory on demand, so you don't need to actually zero the entire .BSS section on startup.
On a bare metal system usually it's done either in assembly (or more cheezy in C) + linker magic.
I think you can in fact realize the idea in PECOFF, by the way. I still need to figure out some aspects to it these days (it seems a bit arcane and maybe Windows doesn't follow the spec very well). But in any case PECOFF has this notion "VirtualSize" (size of section in running program) and "SizeInFile" (size of the prefix of the section that should be filled with contents from the executable file). If SizeInFile is smaller than VirtualSize then the rest of the image gets filled with zeroes I think.
You can express it in ELF too, with a NOBITS segment with ALLOC but no WRITE flag. Whether that works or exercises bugs in the dynamic loader is an open question.
Correct. A common and low-hanging-(more like "lying on the ground")fruit optimisation for PE files is to reorder and realign the sections such that all the 0s are at the end, in which case they can be "cut off" by setting those header fields appropriately and not waste storage space.
;-)
What OS? I don't have one. I have to link in init code to zero sections of ram before my embedded code runs.
Sure, it got into the standard because some popular OSes like Unix zero initialized pages, but it's not a universal truth. The reason it's guaranteed in C is because of the spec.
C89 only requires that static values be initialized.
A modern C standard (section: 6.7.8(10)) requires static values be initialized, but what the value is initialized too be _technically_ indeterminate.
There is the guide line given that integers must be zero, and pointers be NULL. But if a static storage class isn't consisting of purely integers, pointers, or (fixed sized) structures, arrays, and unions who's elements can recursively reduced to integers or pointers. Then the standard says the initialized value is indeterminate.
While relatively straightforward, there is a few gotcha's.
``` If an object that has static or thread storage duration is not initialized explicitly, then:
- if it has pointer type, it is initialized to a null pointer;
- if it has arithmetic type, it is initialized to (positive or unsigned) zero;
- if it is an aggregate, every member is initialized (recursively) according to these rules, and any padding is initialized to zero bits;
- if it is a union, the first named member is initialized (recursively) according to these rules, and any padding is initialized to zero bits; '''
I think that covers every possible value you can create, and it seems pretty non-indeterminate by my reading.
Yeah, I know. I live in that world too. I don't know that it's particularly relevant. The C standard is written to a norm of a Unix userspace environment, and that's clearly where the linked article is working.
(Edit to correct: obviously glibc "touches" .bss because it has its own static variables. But there's no "zero .bss" step in crt0.o)
And there are other parts of the C standard that would just declare 'implementation defined', so the standard doesn't have to say static duration is initialized just because that's what multi process systems do.
.L_bss_init:
! clear BSS, this process can be 4 time faster if data is 4 byte aligned
! if so, use swi.p instead of sbi.p
! the related stuff are defined in linker script
la $r0, _edata ! get the starting addr of bss
la $r2, _end ! get ending addr of bss
beq $r0, $r2, .L_call_main ! if no bss just do nothing
movi $r1, 0 ! should be cleared to 0
.L_clear_bss:
sbi.p $r1, [$r0], 1 ! Set 0 to bss
bne $r0, $r2, .L_clear_bss ! Still bytes left to setNow C might want to do it too, but bare metal stuff is not necessarily compliant C so that doesn’t matter much.
It's possible to post-rationalize those decisions, but as I understand it, there's no particular reason why .bss has to be initialized -- it just happens to be the way Unix has worked since forever.
With modern static analysis, there's probably no performance benefit to not initializing stack variables that are read before they're written?
Not really relevant because that's a bug.
But the common thinking "let's always default-initialize stack variables because the compiler will make it efficient anyway" is clearly "sufficiently smart compiler" thinking. Implicit default initialization will never be 100% as efficient as simply not requiring initialization.
Plus, I like the error messages I can get, sometimes, with no default initialization. The compiler can statically detect some uninitialized reads. It cannot do that if all variables are implicitly initialized. In many other cases at least one can be quite certain to experience random crashes that let one track down the bad read rather quickly.
In other words, implicitly initializing to zero normally just hides logic bugs. It doesn't make them go away.
It seems like a workaround around insanity with respect to UB (where compilers fail to notify the programmer that they statically detected a logic bug, and instead go on with compiling, making crazy optimizations based on the wrong assumption that the logic bug was never there).
Not at all. I also believe that we shouldn't hide logic bugs, but the problem with this class of bugs is how they are hard to catch: forcing the initialization can at least make the behavior more consistent across execution of the same program (avoid the rare sequence of condition that will clobber the value the right way before you use it). For some specific application this may be an OK tradeoff for release builds.
But to clarify why I posted this, I was just trying add one piece of data to this part of the thread of discussion
>> " With modern static analysis, there's probably no performance benefit to not initializing stack variables that are read before they're written?" >"Implicit default initialization will never be 100% as efficient as simply not requiring initialization"
Actually we don't know the exact impact, it is likely codebase dependent, but this patch in clang will allow to experiment with various tradeoffs.
> where compilers fail to notify the programmer that they statically detected a logic bug
Compilers (at least clang) won't detect a logic bug without noticing the programmer. You have warnings for this.
The optimize just "assumes" that there is no logic bug but can't reason about the logic:
{ int a = 1; foo(&a); } int b; bar(&b);
Can I optimize toward:
int a = 1; foo(&a); bar(&a);
If we assume that there is no logic bug, then it seems like a valid transformation to me. But if bar reads it parameter before writing to it, then this optimization makes foo impacting bar.
int x = _;
The vast majority of variables can and should be initialized when they're declared (and often don't ever change after, so const is a good idea too). This is especially so in C++, which is less likely to return values via references/pointers these days (with tuples and uniform init for returning structs).Personally, I enjoy that the C standard makes a bunch of guarantees that I can use for my reasoning about correctness. Of course I can only get so far reasoning in abstract, language-defined terms and ignoring the execution environment, but where it is sufficient, being able to forget about operating system minutiae is certainly a relief.
Since the main thrust of the article is that C++'s language-imposed rules are overly complicated, restricting the perspective to the language specification seems reasonable.
It is because the programmer didn't put an initializer into the source code. That's how the language is defined. What you are discussing is the underlying reason the language is defined that way. You are speaking past the author, not pointing out a mistake.
The blog post does contain a genuine error relating to this point, though: Any C programmer worth anything knows that this initializes i to an indeterminate value. This is wrong. Reading the variable does not simply give you an indeterminate value, it gives undefined behaviour (something the article never mentions, surprisingly).
> that stack variables are uninitialized by default is actually an intentional feature, not a bug, as it correctly expresses the behavior of the runtime environment
But it doesn't. Reading an uninitialized local variable gives undefined behaviour, which might not correspond to the behaviour of using any particular garbage value.
Anyway, the language is defined that way simply for performance reasons. The risk of undefined behaviour is a considerable downside but was deemed acceptable.
> But it's got nothing to do with the syntax of the language.
We're talking about the semantics of C. If C were defined such that static integer variables be initialized with 1 rather than 0, that's what would happen, no?
(Aside: I believe ++i++; is no longer UB in the latest version of the language, but used to be.)
On the particular topic of initialization in C/C++, I go with the rule of thumb of be as explicit as possible.
Beyond that, I'm not afraid to assign a marker value to a local only to overwrite it soon afterwards. I save myself endless trouble in case my code is buggy, and if it's not, the compiler is likely to elide the first assignment entirely.
Clang in particular is quite aggressive about this, and is prone to producing surprising results. GCC has started doing this too, but it is much more conservative,
> variables must always be initialized before they’re used.
Also I believe I’m quoting close to the standard there about indeterminate value.
I stayed away from using the term “undefined behavior” because I’ve had a lot of problems in the past with students believing undefined behavior means “but it works if they ran a test with a particular compiler version and flags and didn’t see any problems”. No matter how I try to equivocate undefined behavior with invalid code, it has trouble sticking. This is a <edit>understandable</edit> perspective built from starting at interpreted languages that immediately notify of any runtime errors and otherwise don’t have undefined behavior.
> I stayed away from using the term “undefined behavior” because I’ve had a lot of problems in the past with students believing undefined behavior means “but it works [...]
Seems to me that developing a due terror of UB is vital to having a basic understanding of C as a programming language * , and beyond that, it's vital to appreciating what C is.
As well as being a considerable practical headache (it's not a compiler bug, it's UB going haywire), it's one of the major differences between C and, say, Java (alongside the way C types are not portable, etc). I'd say it deserves some emphasis as a basic principle.
* I suspect some theorists would contend that, strictly speaking, C programs aren't programs, and C isn't a programming language. Programs are, theoretically, meant to unambiguously map inputs to outputs. C is not unambiguous. But, needless to say, I digress.
Try telling students to add -fsanitize=undefined to their compilation flags. That might help make it clear that their code is buggy.
Undefined behavior in the technical sense is not acceptable in a language at all, let alone a feature. If you take the idea of undefined behavior seriously, then it is valid to blow up the world in response to an error and the compiler cannot protect you. Nobody would use C or C++ if they actually accepted the meaning of undefined behavior, so to program in these languages requires embracing doublethink.
That's not its job. The compiler cannot protect you from everything. (You shouldn't be using a computer that has the ability to blow up the world.) The compiler merely translates code from language A into language B. If you pass it some code in some third language which could parse as mal-formed A code, then how is to detect that?
If I write a language with the same syntax as C but entirely different semantics, and I pass it to a C compiler, what do you think should happen? Undefined behavior is what happens.
The compiler can protect us from many things, such as uninitialised variables. It’s just that we choose to define the semantics of the C compiler such that it doesn’t.
Also I don’t think GP meant literally, physically destroy the world. It was a metaphor for catastrophic consequences of program misbehaviour. They meant that literally anything could happen, having dire consequences for reliability and security.
It's the equivocation between inconsistent ideas that frustrates me. Of course, practically speaking you assume really bad things don't follow from undefined behavior. But then why the constant refrain about how we shouldn't rely on what actually happens?
I think I understand where the motivation for declaring undefined behavior comes from - people who set standards don't want responsibility for situations they don't completely control. But this disclaiming of responsibility puts the users of the standards in an impossible situation as a result.
I'm coming from a perspective of someone who programmed in C as my second language after BASIC, back in the 80s, before modern standards and before I knew anything about language standards. The philosophy and attitude of people who talk about standards and undefined behavior is something I first encountered on Usenet in the 90s, but I still am disturbed by it and haven't been "educated" to accept it.
This 'other language' approach doesn't strike me as a good way of thinking about it. It misses the point that C is pretty unique in its broad use of undefined behaviour. Unlike Java, where everything has an unambiguous definition. (Well, ignoring plenty of platform-specific variation points in the standard library, such as file-path syntax.)
> Undefined behavior in the technical sense is not acceptable in a language at all
The C committee disagrees, and they have their reasons. They're aiming to maximise performance and support for various weird and wonderful platforms.
C's philosophy is not like that of Java where, say, an int is defined to be 32 bit and your machine just has to make it happen, even if it's a peculiar 48 bit machine or something.
> If you take the idea of undefined behavior seriously, then it is valid to blow up the world in response to an error and the compiler cannot protect you
Well sure. UB means you screwed up. C isn't a hand-holding language. Explode the whole process. Perfectly valid. Not without precedent in the C++ spec, incidentally; C++ is defined to explode your process if you screw up in a certain way with exceptions https://stackoverflow.com/a/43675980/
Your program can't literally blow up the world, of course, but that's beyond the scope of the language. If UB could result in nuclear apocalypse, that would mean a compiler bug could generate a binary that did the same thing, regardless of source language.
> Nobody would use C or C++ if they actually accepted the meaning of undefined behavior, so to program in these languages requires embracing doublethink.
Plenty of C/C++ programmers have a very sloppy attitude to undefined behaviour, and sometimes they get bitten by it. Aggressive optimising compilers like gcc, really do depend on you taking responsibility for writing code with defined behaviour. Break that contract, and you can see nightmare intermittent bugs that only appear when using certain compiler flags. See: http://blog.llvm.org/2011/05/what-every-c-programmer-should-...
Of course, this is all made worse by that it's impossible for static analysis to detect all instances of UB. If you want a language that doesn't take this attitude, you might try Ada. It's a pity more people don't.
#include <iostream>
struct A {
A(std::initializer_list<int> l) : i(2) {}
A(int i = 1) : i(i) {}
int i;
};
int main() {
A a1;
A a2{};
A a3(3);
A a4 = {5};
A a5{4, 3, 2};
std::cout << a1.i << " "
<< a2.i << " "
<< a3.i << " "
<< a4.i << " "
<< a5.i << std::endl;
}
which outputs: 1 1 3 2 2
he claims to be mysterious but is actually pretty reasonable.In the case of:
a1: There is no initializer list in the variable declaration, so ctor 2 is called.
a2: An empty initializer list should reasonably behave like a default constructor, and a default constructor should be more efficient than processing an initializer_list, so ctor 2 is called. In general, the constructor with the matching number of arguments and correct types is preferred over initializer_list constructors. Makes sense, being able to specialize on number and type of arguments is more powerful than a design where a single initializer_list ctor invalidates all other ctors.
a3: No list, ctor 2 is called.
a4: Non-empty init list and no specialized non-init-list ctor, so ctor 1 is called.
a5: Several-variable init list and no overriding 3-variable constructor with matching types, therefore ctor 1 is called.
Complex? Maybe, but that's what you get with C++: very fine control of program semantics, benefit being expressive libraries. No other language fills this niche that I'm aware of.
I agree that C++ isn't a good language to teach in a CS 101 class. But no other single language is good either. The goal of CS 101 is to not make students give up before they get hooked, and that can happen because the material is too challenging or not challenging enough. For people in the former category, give them a scripting language and visual feedback, like Lua + Garrysmod. For the latter, give them assembly, haskell, C, C++ (teaching it like "C with templates" not "C with classes").
Congrats on successfully summarizing the complaint.
It's not mysterious, but it's gratuitously complex. And this is only variable initialization.
If just setting integers is this complicated, what do you expect to happen when you are trying to solve real problems?
struct A {
int i = 0;
}
int main() {
A a;
std::cout << a.i << std::endl;
}
Hey, look, done. And you only had to teach a single thing - default initialization. Which has an obvious & simple syntax. The int i is always initialized to 0, as intended, and it's in a single spot at the point of declaration.Now try doing that in C. Oh, wait, you can't. The closest you can get is this mess:
struct A {
int i;
} const default_A = {0};
int main() {
struct A a5 = default_A;
printf("%d\n", a5.i);
}
Which requires you to know that you can declare a type & an instance of that type in a single statement, why that const is where it is and why that's important here, how braced initialization rules work, that you need to always remember to manually initialize to your default_A, and that %d means 'int'. And the compiler won't help you with any of this except for the %d part if you get it wrong. struct A {
int i;
};
const struct A default_A = { 0 };
Why didn't you try the obvious simple thing first?The whole undertaking is pointless in any case. Why would you need default values (i.e. templates for constructors) baked in? The only reasonable default value, sometimes, is all-zeroes (or all-ones...).
Now, constant data (i.e. things that contain more useful information than just "it's the default because a real value is missing" -- and that aren't copied around pointlessly) is another case, and C's plain old value initializer syntax (as demonstrated above) serves it perfectly well.
That's very false. Consider a basic string container with a small-size optimization. The default value for capacity is neither 0 nor 1, but the size of the inline array.
Similarly it could be an enum value, and the default for a given class isn't whatever the 0 value happened to line up with because that's arbitrary anyway. An example being a basic type id of a fixed number of types.
> that's what you get with C++: very fine control of program semantics, benefit being expressive libraries.
It's not gratuitous because you can't remove much of it without losing power or expressiveness (and still remain a C).
Also, by complex I didn't mean complicated.
Yes, initializer list with empty initializer list should behave same as empty constructor, and if it doesn’t, it’s really bad style. However, if both are defined, then if you are constructing an object with empty initializer list, the initializer list constructor should be called with an empty list as an argument, simple as that. That’s very logical and sensible semantics, so obviously C++ had to choose something different.
Unfortunately, std::initializer_list created a new ambiguity, so it's only an initializer list if it doesn't match an existing constructor.
My biggest goal for a CS 101 class would just to build computational and critical thinking skills. I would love to just start off with peanut butter and jelly sandwiches[1]. As an engineering department, our students start with an engineering programming class that needs to also serve chemical, mechanical, civil, material, and biomedical engineers along with electrical and computer engineers(I hope I didn't leave anyone out).
[1]: https://edtechmagazine.com/k12/article/2008/07/programming-a...
https://www.youtube.com/watch?v=7DTlWPgX6zs
C++ has a rigorous standard. There is no rule that wasn't added for a reason. Nor are there lines of code with which you can specify which rule it must follow.
The path to hell is paved with good intentions. Just because each step is logical and defensible doesn't mean the end result isn't fire and brimstone.
I believe that good things are more than the sum of their parts. You often hear this with respect to creative works like movies or games. I think it also applies to source code, language design, and much more. The flipside is that bad things are less than the sum of their parts. I've shipped games that fall into this category. C++ initialization has a LOT of parts. And I would strongly argue the end result is less than the sum of those parts.
Rust gives you too much control when you ask for it, with unsafe blocks, cells, etc. C++ will stab you in the gut at random because you looked at it the wrong way.
> But no other single language is good either.
I had Python for my CS101, Java for 102, C for 201, and C++ for ~202. This was a decade ago now, but even then I thought it was a pretty good track. You learn fundamentals of computation in Python fully apart from the hardware warts, you get to taste what the corporate monotony many of you will be faced with for decades to weed out the chafe, then you get a dose of cold hard reality when its too late to turn back that it gets even worse.
Of course, those warts were necessary for C++ to be successful back when it was introduced and competing against others, and therefore to its popularity today. And this popularity is just as much a part of C++ appeal as its power. But we can call them out for what they are, without trying to justify them.
True. But an empty initializer list should also behave like, you know, the constructor which uses an initializer list. C++ is in the unfortunate position of having to choose one behaviour or the other. Either choice is reasonable, but both can be confusing.
This is a ridiculous argument. The entire point of the post was, as the author even noted, purely to deep dive into a rabbit hole, get super picky & pedantic about standards wording so that you can act surprised that copy constructors exist in C++ while it was completely glossed over in the identical C example of a struct copy initialization, and almost none of this is useful to know.
So don't fucking teach it and you'd have plenty of time to cover all that other stuff. Bam, problem solved. Ignore pre-C++11 entirely, and purely teach & use the new stuff which fixes all the complexity, and leave the rabbit hole for people that care about exploring the past.
Oh, and don't use standards wording because those are for compilers to use to implement the language, not for programmers to understand how to use it effectively. The only time it's ever useful to know the full definition of an aggregate type or how it has changed is when you want to write a blog post calling them crazy or complex. It's never useful to know when using the language. And if you have an object that cares to enforce it just
static_assert(std::is_aggregate_v<A>, "A isn't an aggregate type");
Tada, now the compiler will tell you if/when you violated the rule and you don't need to try and understand the full scope of the ruleset.It doesn't, unfortunately it just adds to the complexity.
Take for example, `auto`, added in C++11. If you don't have to support C++98 or C++03, the use of `auto` significantly simplifies many workflows and is often strictly better with less mental overhead.
When writing code, yes, but in my experience, it can also make interpreting what code that overuses auto is doing (especially what variables are and what functions return) an awful lot harder in many situations, and you end up jumping all over the place to work out that "auto result = ..." is actually "uint32_t result = ..."
That's one of my personal pet peeves about more modern C++ (and other languages to some extent like Rust): they seem to be more optimised for writing code quickly (which generally only happens once), not understanding it later (i.e. you didn't write the code someone else did) and maintaining / altering it in the future, which generally happens a lot more over code's lifetime.
But knowing if something is a base type, a reference, a (smart) pointer, or an expensive-to-copy class/struct is very important when writing high-performance efficient code. Being able to see this at-a-glance by the type in the code in my experience helps tremendously with understanding what the code's doing and the implications in terms of data passing / transfer and understanding how that section of code interacts with other parts or could be changed to do other things.
I also think it allows people to be a bit sloppy and not care what's going on (i.e. with regards to whether it's by-value or by reference, etc) as they don't fully need to understand the code, and again, in high-performance computing where you seriously care about processor cycles and memory allocations, this can make a big difference if you're not careful.
I could have easily used Go or Swift as examples instead.
And to be clear, these new languages are in general improving things a lot, but I just worry there's an over-emphasis on "being able to do a lot of complex stuff which very few syntax / characters", which I don't really fully agree with (at least for large complex long-term projects).
That's not to say I'm right, but in my experience of programming over 15 years in everything from ADA, C, C++ to Java and Python, at least on large projects, the speed at which code was created was very rarely that important. Getting it right, bug-free and performant was generally much more important. In some cases newer more-condensed syntax can definitely help, but in others, it can cause trouble.
Code is read much more often than written, and most newer languages optimize for convenience rather than readability.
In my view Rust is actually pretty verbose compared to other new languages, and feels a bit cumbersome to write. Technically it could be made quite a bit leaner syntax-wise.
I just wanted to point out that Rust and the Rust community often share your view and want to optimize for readability as well. There actually was a lot of heated discussion last year around proposals that suggested reducing readability for convenience.
That's a lot better to go on than "auto".
That's what auto fixes. It avoids coupling the concept of iterating from the specific type of the iterator which isn't important.
Similarly auto helps you achieve DRY. 'auto myFoo = std::make_unique<Foo>();'. The type wasn't removed, it just wasn't repeated twice.
But at this point auto, or things like auto, exist in nearly every major language. So you'll need to teach best practices for working with it at some point. That's a general thing that's everywhere.
Please keep it civil. I'm sure you could refute the GP without expressing contempt.
Things like strings and pointers are nightmarish in C. Segfaults aplenty.
I've never encountered a segmentation fault since switching to modern C++. Just because you use C++ doesn't mean you have to use the entirety of it.
Welp, time to go back to BASIC everyone. Or Go
Powerful tools can be complex, who would have thought
- use C++, given that you stick to some convenient subset of C++ and can use stl
- stick with C and force them to do the basics, like linked lists and string abstractions over and over again.
I guess in the olden days really good students used to develop their own library of C abstractions, and reuse them with several courses; but you can't quite do that if you have C in just one course and what's the point anyway in this day and age ?This could happen for example if the data structure course was in python or java and later courses are in C.
Also you can do the darn list in many ways: single linked list, double, ring, with counter/without counter. All very important in this day and age when you should know to avoid them altogether because linked lists fucks up cache behavior.
Also, standards definitely shouldn't be inscrutable to programmers. Look at the ECMAScript and Go specifications, they're clear as day. Sometimes you just need to know exactly what your code is doing, and this only becomes more important in the low-level scenarios that C++ is often used for. I haven't read the C++ specification, but if it's as hard to read as everyone says it is, then there needs to be some other way to understand exactly what's going on. How else am I meant to precisely understand how my structs will be laid out, or which casts are guaranteed to be valid, or whether the language allows data pointers and function pointers to be interchangeable?
https://en.wikipedia.org/wiki/Most_vexing_parse
And it's one of the motivations for the {} syntax.
The post was mostly written to point my students to, so I don't have to keep repeating myself. I get a not-insignificant number of 1st, 2nd, and 3rd years (in a 5-year program) believing C is some antiquated language and believing that they're getting held back in some way by learning C vs C++. One even suggested the department was incompetent for not teaching C++. Because I work in an engineering department, many students have not had a great deal of time learning programming fundamentals early on and regardless, they are eager to learn more advanced tools than they are ready to use. It is a bit rambling for the purpose--I have a tendency to...erm, overwhelm...with information to make my point.
I haven't blogged much so I originally submitted it at lobste.rs[1] for any advice on the writing and visual style of the site. I welcome any constructive feedback. E.g. "I hate that side-nav! It keeps popping in and out!"
Also to clarify, I am not anti-C++ is anyway. I am a firm practitioner of Chesterton's fence[2] and believe in nuance. That cuts both ways. C++ is the way it is because it filled a specific need. It's greatest flaw is trying to appease everyone and, recently, trying to catch up quickly to recent QoL features in other languages. This fortunately gives it a lot of features other languages don't have, and it unfortunately gives it a lot of features other languages don't have.
Once again, the greater point being a warning, that C++ can easily become a time sink in language-specific knowledge instead of domain-specific knowledge. Of course sometimes that's what you want, e.g. when trying to optimize for performance.
[1]: https://lobste.rs/s/tul188/initialization_c_is_seriously_bon...
[2]: https://en.wikipedia.org/wiki/Wikipedia:Chesterton%27s_fence
I personally like rust and hope it does well. I see it somewhat orthogonal to both C and C++
This is obviously a personal opinion, but sometimes I like being able to create bugs in my code. Not from an industrial or business standpoint, but from a greater understanding POV. It's easier to reason about the underlying machine (yes, yes I know C/C++ models abstract machines) when I can actually break that machine with the tools at hand. There's unsafe Rust which I have not looked at, but in terms of getting my hands dirty, sometimes C/C++ just feels better. The primitiveness, even of template programming vs Haskell's typeclasses, or constexpr vs D's CTFE. Something about that raw primitiveness is attractive. This is absolutely positively probably just experience bias. I'm not sure if anyone else can relate to this.
So in that regard, I see Rust as orthogonal. If you want that feeling like, "hey, I'm just directly fiddling raw virtual memory addresses", that's not Rust's target. Rust markets itself as a safe language that hides all those bits by default.