Deep C and C++ (2011)
slideshare.net
slideshare.net
In an interview, I would much rather hear a person say something about a statically declared variable with no initialization being poor code to leave behind for the next person than some arcana about the standard.
Why are you forcing those two choices? One could be well-versed with the arcana AS WELL AS point out that bit about statically declared variable ..
The problem IMO is most programmers cargo-cult/copy-paste/stackoverflow their way into programming jobs. No, that does not mean you never ask for help. No, that does not mean you never copy-paste. (Phew!)
Its sort of like when you're learning math. Great mathematicians can understand the theory and just apply it to whatever problem they come across. Its because they have a very solid foundation underneath them. They wield their knowledge like tools and can just build anything with those tools because they understand those tools very well. Most students however just learn the patterns of the problems. And once they know enough patterns they can solve problems which fit into one of those pre-understood patterns.
The people who are deeply knowledgeable about the language are good at knowing the boundaries of the language, knowing when you're using constructs that are not valid-syntax (which still compile), etc. The social aspect of 'good comments' , 'readable code' or 'maintainable code' is also important. You can have programmers that do both. Ofcource those programmers will never work for 'you' (not you, personally..) because most programming jobs are shitty and do not require programmers of that skill. You'll find them toiling away in anonymity, working in research labs, working on compiler optimizers or operating systems or some other domain with challenging technical problems.
As someone who's both a mathematician and a programmer, I think that analogy glosses over an important difference, which is also relevant to why many people refuse to memorize certain things.
Good mathematical theories are internally consistent in a very strong sense. When you're learning them, you feel like you're learning something eternal that could not possibly turn out any other way. If you forget half of calculus, you can reconstruct it from the other half, and the reconstruction will be unique.
The details of programming languages, like C and especially C++, are not like that. For many mathematicians who would otherwise make excellent programmers, the details of C will feel like something not worth memorizing, because they don't make sense - they are not an inevitable, provably unique solution to any problem.
To a mathematically inclined mind, programming concepts form a hierarchy depending on how "inevitable" and worth memorizing they are. Concepts like computability, lambda calculus or big-O complexity are near the top, syntaxes of specific languages are lower down, and API details are at the bottom. That also reflects the rates of change: math doesn't change, languages change every decade, and APIs change every year.
If I want to hire a programmer for the long term, as opposed to getting quick help on my current project, I will be mostly asking them about concepts from the top of the hierarchy, not what static means in C. Of course some people will disagree with me and demand detailed knowledge about static or templates or whatnot, while dismissing mathematical questions as "puzzles". Maybe that's also a valid style of programming, I'm not sure...
No wonder everyone is always talking about the rapid pace of learning in high technology.
Have you implemented a compiler or a interpreter for a popular language? I could be wrong but I strongly suspect you haven't. If you have ever implemented a compiler you understand the language at a fundamental level and when stumped with a bug or a highly technical puzzle or what have you - you get that "aha" moment where its like "Ofcource ! How else could it be. This language feature must be using this construct because of X or Y reason and the compiler has to work this way because this part here doesnt allow for it to work any other way, etc ,etc and the optimizer has xyz amount of scratch registers so abc condition could never happen, and on and on."
Lawrence Kesteloot said it well at http://www.teamten.com/lawrence/writings/what_interests_me.h...:
On a planet far away they have universities. In those universities they have many of the same subjects we teach here, such as math, biology, philosophy, history, computer science, psychology, and literature. But although the subjects of the classes are the same, the contents are different. They teach a history class, but their history is different from ours. So are their biology and literature. I’m not interested in subjects where their content is different from ours.
In some classes, though, the content is the same. They’re teaching the same physics class (assuming they’re at the same level of understanding we’re at) and the same math classes (with a different base). In computer science most of what they teach is different, but their theory of computation is surely the same. Their biology is different, but their teaching of evolution is the same. Their philosophy is vastly different (think how different Eastern and Western philosophies are on earth), but their philosophy of science is probably very similar.
Even if this point of view doesn't make sense to some people, I subscribe to it because it's very inspiring to me. YMMV.
You're mis-applying things here. The language is not something to be theorized by itself. Like any (edit:spell) other body of knowledge, it has its own set of initial axioms and the 'C standard' which defines the language is simply a collection of rules other meta-knowledge which are based on those. In that sense it is internally consistent. Thats what I meant.
Compare with a formal theory like PA, which fits on a page and describes the natural numbers. Proving a theorem about PA has implications everywhere, because natural numbers are everywhere. Or if we want to talk about computation, there are many formal theories with short descriptions that have much more to say about computation in general.
You misunderstand the usage of consistency regarding mathematics. It is not the same. These rules which compose C are design decisions, and whether they are entirely artificial. Mathematical consistency arises from imposing axioms, following which there are no decisions to be made.
Well, no. You can deduce A language but it might not be C.
I sense we're going down a fruitless path here. In my first post I was only giving an analogy of someone understanding math at a fundamental level to someone understanding a programming language at a fundamental level and how they could inform similar abilities/prowess in people when it comes to dealing with complex problems/etc.
> You could make many different tiny changes to the rules of C and still end up with a workable programming language. (Many people have done that, there are tons of languages derived from C.)
Coincidentally, same is Math. I give you Euclid's Fifth postulate as an example.
>You could say that the rules of C are "random" - after learning the first n rules, you cannot use logic to predict the n+1st.
This depends on how do you enumerate these rules. If you learn that auto variables are allocated on the stack it's easy to predict that you cannot return their addresses and there are not going to be default initializes for them.
I'm not sure what you mean about the operator overloading thing.
Template types are also not really first-class members of the type system. The compiler can't really reason about them, it only type-checks them after expansion, when all the genericity disappears. This is completely different than the fail-fast approach in Haskell or even Java, where the code is type-checked at the generic level.
Don't get me wrong - templates is an awesome, extremely powerful feature of C++, but it doesn't play well with the rest of the language, and really feels tacked-on.
Function templates are not generic functions. Same as template classes are not types. They play really well with the rest of C++ but I agree, not so well with the rest of Java or Haskell :)
This is the major problem people nowadays have with C++ - they learn some "practical" high-level language in school (usually riddled with random stuff, judging by how much people here are pissed with my mentioning of PHP) and try to think about C++ in concepts of such a language. Then blame C++ for not being their favorite language.
Don't get me wrong. I am quite happy with this making my skills very rare and expensive, giving me both job security and decent compensation. I also want to thank all the people spreading the "C++ is complex hodgepodge of completely random things so it's impossible to learn" idea :)
"Don't get me wrong. I am quite happy with this making my skills very rare and expensive, giving me both job security and decent compensation."
Eh? C++ skills are not any more expensive than decent Java/C#/Python/(put whatever mainstream language here) skills. Who told you that? This is supply-demand. There is lower supply of C++ coders, so some may think the prices should be higher, but in fact it is compensated by lower demand. C++ is also still being taught at schools, so those skills are not that rare as, say, Haskell, Erlang or Go.
As for the price of my skills - my paycheck told me that. And sure, it could be low demand (there are vacancies in my field, that had been open for years, but I don't really know what kind of demand there is for Python, could be decades for all I care). I am just happy for my compensation and if others have more expensive skills - I am happy for them too :)
IMHO the approach is a great example of premature optimisation: generate 2^N versions then try to fold some similar ones to save space. Much better approach is: generate 1 version and specialize small parts when really needed. Technically, assuming sufficiently smart compiler, both would end up with exactly the same code, performing exactly the same, but the latter solution is preferable for other non-performance wise reasons (like better type checking, ability to distribute binary libraries, faster compile times, etc.).
BTW: I still often find Scala containers to be often much faster than C++ STL. But this is probably for much different reasons than template/generics implementation.
In the first case - I fail to see how it follows. In the second case - our levels of expertise is too far apart to discuss anything. I am not saying you are not qualified, I might be just too out of the loop and missing some dramatic advances in Java, Scala etc.
Even though the library is covered by the same standard as the language it's not a part of the language and, in practice, is not even available all the time.
A better nitpick would be dynamic_cast throwing under some conditions. But then, how does it make iostreams the part of the language? There is just a typedef in the std namespace that aliases the type for some optional, compiler specific, data.
Another way to put it is, learning the postulates for langx doesn't provide enough value for many to be worth doing.
And I agree, this is one of the assertions my opponent has made. I did not go to refute it because, if not outright false, it's disingenuous.
>Another way to put it is, learning the postulates for langx doesn't provide enough value for many to be worth doing.
This is also true, same goes for everything else. I make money using C - I learn C's postulates. If somebody paid me to use topology - I'd learned topology postulates. What would be stupid if I were paid to use topology but said: "it's hard to learn and it makes no sense for me because I don't like mathematics and don't know any other math discipline so why would I all of sudden learn topology's postulate when I can copy formulas from the net and ask questions on manifoldoverflow.com?"
Point being that there is no system that is complete. There are large bodies of mathematics these days that have specialized their notations and axioms to their domains. It is incredibly difficult to maintain a masterful knowledge of all of their idiosyncrasies and notations.
You're the one forcing the choices.
He's just saying that one should consider another metric to gauge developers, that is knowledge of good software practices over deep language arcana wisdom. And he's saying that whatever the knowledge depth (or lack of) of a candidate on the C or C++ language, he'll choose the one with better software practices. (Of course, the two metrics are slightly correlated, but being an expert on C trivia doesn't necessarily makes you a good developer)
Life is too short to become an expert on C's dark corners, and I'm not even talking about C++'s.
"I would much rather [a person do X] than [a person do Y]"
I parse this kind of statement as a person doing X is more important than a person doing Y. I stated that I thought both were achievable. How have you parsed it?
Also trivia implies some memorized factoids. The entire point of the article is having an understanding of the underlying implementation of the compiler, maybe the implementation of the standard library, or the OS or the hardware ADDS to your knowledge about C making you a much more complete/well-rounded developer. Calling it C trivia is baffling to me.
>Life is too short to become an expert on C's dark corners, and I'm not even talking about C++'s.
Agreed !
That's some of the point of the article, but some of it is also about memorized factoids. There is a lot of focus on what different versions of the C standard say about different things, which is just factoids. There is a lot of focus on how different declaration syntax behaves (statics, linker visible, etc.), which is just factoids.
Also the difference in syntax and translation unit/linker behavior is not a random fact. Its the culmination of the thought process that goes into designing the language and also knowing a bit of the theory behind classical compilers/linkers.
Sadly you are required to move for such jobs, which is not always an option.
Compare this with C. There is a degree of serious technical focus that languages like these demand for big projects. Embedded systems, Operating systems, Databases, compilers etc. A big part of software world uses these languages on a daily basis. But there is no where the kind of hipster crowd, these languages have the web frameworks or languages have.
Even if you spend some time searching for resources on the net, there are few that teach deep C skills. There might be one or two books out there which haven't been updated in 2 decades.
By and large, these are unfashionable fields to work in. The barrier to entry is really high, Success comes only with seriousness and application of well focused effort, mistakes are expensive and they warrant serious RTFM'ing the hard way- to write to some real serious code.
Sure some languages take less time to learn than others. But every single language makes tradeoffs. For e.g performance vs productivity/usability OR static typing vs dynamic typing is a common theme when discussing C vs {Java, Python, blah}.
Also, your comments falsely compare the difficulty of the TASK with the difficulty of the IMPLEMENTATION/LANGUAGE. They are not the same! Take any advanced applied math algorithm and implement it in Python. Its not 'easy' by any means.
The domains you mentioned have a higher barrier to entry because the tasks themselves are difficult. The languages and technology used are simply variables in a much larger optmization problem that also includes monetary/human/computing/time resources as variables.
Personally I don't care about C, due to its unsafe by default design and my background in safer languages for systems programming.
I got my first computer, a Timex 2068, at the age of 10. Spent two years mainly playing games and eventually started looking into programming.
When I learned C, at the age of 16, I already knew BASIC (Spectrum, GW, Quick, Turbo), Assembly (Z80, x86, 68000), Turbo Pascal.
So C was for me a bit "meh" language, given that Turbo Pascal gave me much more, granted except for portability. So I quickly jumped to C++, which allowed me some strong typing comfort and higher abstractions in a forced C land.
Regardless of the language, when I was 18, I had already coded compilers, graphics applications, GUI prototypes, MS-DOS drivers.
Nowadays kids seem just to do JavaScript/Python/Ruby frameworks for generating web pages.
That's progress, IMHO.
Instead of distributing application binaries carefully targeted to each type of user machine, we only have to run the application on one server machine (or a cluster of them) in a central location and can serve the application over the network. We now have code-generating frameworks that help automate away much of the tedious, repetitive parts of implementing an application.
It's not a good thing that software development is hard. It does mean that you're a very smart and hard-working person if you're able to do it effectively. But it's not a good thing for society that you need to be so smart and hard-working in order to do it.
It is, however, a good thing that it's not as hard as it used to be. The fact that a software developer has to do and know less today than they did 20-30 years ago, just to build an application and distribute it to a large number of users, is a sign of progress.
it has pointers, pointer arithmetic, manual memory allocation, function pointers, casts, unions, you name it.
I bet you could write could write 30 odd rules for cpp and then compile your "pascal" with cc, e.g.:
#define begin {
#define end }
#define := =
#define ^ *
#define @ &
and so on
- It has real strings
- Arrays don't decay to pointers
- It has reference arguments, which avoid pointer usage for records and out arguments
- Requires explicit casts for type conversions, instead of implicit type casts
- Provides region allocators, which allow to release a full block of memory in one go
- Has real enumerations that don't decay into ints
- Allows writing generic code over arrays without requiring the developer to manage a separate length variable
- Has proper memory allocation constructs instead of requiring the developer to know the size of memory to allocate
Finally, there are other languages in the Pascal family, that provide safety, while allowing for C dirty tricks done explicitly via a system/unsafe package which provides much fine control over the use of said features.
pascal style strings can be done with the preprocessor, as can generic iterators, allocation for idiots, and "out" function arguments.
memory pools are a library feature.
and the arrays are the same too, a pointer in pascal can be treated as an infinitely sized array, but the reverse isn't true, just like in C.
and, believe it or not, there are other languages in the C family, that provide safety in the same way there are for pascal.
tl;dr: pascal and C are more alike than different.
That still does not make them safe, implicit conversion to int will still happen.
> pascal style strings can be done with the preprocessor, as can generic iterators, allocation for idiots, and "out" function arguments.
No type safety.
> memory pools are a library feature.
That aren't part of any C compiler. So you can't count them being available.
> and the arrays are the same too, a pointer in pascal can be treated as an infinitely sized array, but the reverse isn't true, just like in C.
You lost me there.
> and, believe it or not, there are other languages in the C family, that provide safety in the same way there are for pascal.
Yes, that is why I use Modern C++, C#, Java, D and will only use C at gunpoint.
> tl;dr: pascal and C are more alike than different.
True, except in Pascal when bad things happen is because the developer explicitly wanted them to happen.
In C everything goes and to have Pascal's safety back, all modern C compilers provide static analyzers. Thus proving a point about the languages design's.
True. Talent going to waste is sadly nothing new. Has happened, will continue to happen. :(
But puzzle questions like that are silly anyways, so whatever.
I don't think you know how great mathematicians think. http://en.wikipedia.org/wiki/57_(number)#In_mathematics "Although 57 is not prime, it is jokingly known as the "Grothendieck prime" after a story in which Grothendieck supposedly gave it as an example of a particular prime number."
I'm confused. What is your basis for that statement? If you you know how they think why don't you just express it? :-S I am sorry, I have no clue what your point is, or even if you have one.
That makes him a counterexample of the claim that all great mathematicians can apply heir knowledge.
If he were a computer scientist, he would have written several brilliant papers about the pros and cons of various properties of programming languages, but he would not be able to write a compiler for any language or even to write a program.
"So, that's what you think your code will do... what does the standard say?"
or, "Why did you use int there when the API spec defines the parameter as taking size_t?"
or, "You think that code in your inner loop will compile to one machine instruction. Have you actually looked at the compiler's output?"
and the looks/responses that I got. There seem to be very few programmers who understand the language & their compilers deeply enough, or even care to learn. These programmers are far more productive because they're not constantly guessing or assuming how their code might work.If I see
static int i;
in code, I'm going to change it to static int i = 0;
even if I know that static integers are initialized to 0 in the standard, because I know that eventually someone will have to maintain the code who doesn't know that. I don't particularly care what the standard says there: the code shouldn't assume detailed understanding of the standard, which isn't realistically a valid assumption.Don't get cute and always prefer to explicitly state in code what you get for free implicitly. The maintenance guy that follows behind you will be the one to espouse how clever you actually are. That's how I mentor junior engineers as a general rule.
A person who just knows it's "bad code"--but not why--is almost certainly going to leave other bad code from lack of understanding. To pull an example from the slides, the virtual destructor: making that class virtual when it shouldn't is a waste of CPU cycles _and_ bad documentation for future developers.
Sometimes I think that having deep and expert understanding of a language may cause you to create code that other team members cannot understand... not on purpose, but due to your assumption that these are common knowledge (whether they should be or not isn't the issue.)
C and C++ can make someone loose days to track down issues, in this day and age, where teams are distributed with lots of offshoring and various skill levels across development sites.
My last C++ project was in 2006, since then I have only used C++ outside work. At work our focus has been in JVM and .NET languages.
I don't miss playing the C++ fireman expert role that has to fix a stability problem created by someone in the other side of the planet.
Nonobligatory image to support statement: http://codinghorror.typepad.com/.a/6a0120a85dcdae970b0128776...
I've since come to believe that reading the specifications and having the attention necessary to delve into these kinds of details and ask the right questions is important for mastery. It seems to me that learning 1 - 2 languages to this level of detail is worthwhile. I've been thinking of cutting back the number of languages I, "know," down to just those for which I am familiar with the specifications and how they're compiled, assembled, etc. Everything else is superficial.
Sometimes all you need is just a cursory knowledge to get something done and the ends justify those means. However if you really love your craft then mastery should be the goal, no? It seems to be the difference between, "getting something working," and, "pushing the boundaries of what is possible."
[1] http://www.amazon.ca/Expert-Programming-Peter-van-Linden/dp/...
I think a better balance may be to master things that are widely relevant at depth, while giving less focus to things that are very specific. Put another way: depth for "essence", breadth for "accident". Of the items on slide 181 of the presentation, only calling conventions, memory model, and probably optimization seem like "essence" to me, ie. every program you ever write in any language will benefit from deep knowledge of those topics. The rest of the items are somewhat arbitrary and very specific to the particular history of the C language itself.
I can study music theory for years and it won't help me play the violin well. It will make composition easier and assist me as I adapt myself to the instrument. However one still has to master the instrument and I see the two being complementary and distinct skills.
The metaphor holds together when you start thinking about the breadth of languages available. For example, my primary instrument is guitar. I've been playing for years and have a good grasp of it. However I recently took up the violin so that I could learn to fiddle and play some bluegrass and Irish folk tunes. I still can't play a single tune well on the violin but my knowledge of string instruments and music theory is certainly making my adoption faster than someone who is starting from the very beginning.
However one must be careful with breadth. I could try to learn to play the mandolin after I've "finished," with the violin. And perhaps a banjo after that. However mastery requires a deep, intimate knowledge: you have to spend time with an instrument, learn its quirks, and practice with it every day. Even the best guitar players practice every day. You don't have enough time in this world to master every stringed instrument.
Harder still is the transition to a completely different class of instrument. Mastering drums and guitar is not impossible but rather difficult as knowledge of one doesn't offer much in learning about the other. Learning C is one thing but learning Common Lisp is learning an entirely different model of computation. Having a grasp of the fundamental theories will help some here but nothing you know about C compilers will help you to understand CL (and vice versa). So be even more suspicious of the depth of a programmer's knowledge if they list more than a few languages on their resume if those languages span entire classes of computational models.
And I don't think there's enough time in this world for even a very competent programmer to say they are a master of more than one or two languages.
Update Added a paragraph on computational models and expanded the instrument metaphor.
I'd say asking for a cup of coffee in Paris is "Hello World"
Most of C programmers would know how to tell "a native Parisienne where to get off"
But that level of C is akin to write "Mes emmerdes" (youtube it, there are probably more complex songs, but I can't remember now)
I think it's feasible with an "old-school" (compact) language like C or Go. But take something like C++ or Scala, for example, and the task quickly becomes impossible.
There is also the fact that very often non-specified behavior, or implementation dependent or everything else that is not cool does not lead to a warning, so the learning is absolutely not reinforced by the compiler. Whereas a warning/error leads to questions that leads to google and some learning; you can be stepping far in the Pampa of undefined behavior for years when someone comes with a superior attitude in your company detects it and calls you a moron in a powerpoint.
And this also leads to very hard to write code sometimes, if you want to do some serious IEEE754 in C/C++ you will basically be pitting the spec of the language against the spec of numerical computation in a ring.
I agree, and I'd add that C/C++ is a bit of a leaky abstraction layer over assembly. It aims to be portable, but true portability means that the C/C++ developer really shouldn't have to know/care what assembly instructions the compiler/linker is emitting, because all of that stuff depends on which chip architecture you're targeting.
Basically, coding in C/C++ requires an intimate knowledge of how the compiler works, and sometimes even how the target chip works. C/C++ software projects of notable size oftentimes cannot be simply recompiled/linked to a different target architecture without changing things like compiler flags, the makefile, and even the application code itself. That makes it a non-portable, leaky abstraction, and dramatically increases the amount of knowledge that's required of a C/C++ developer.
I'd love to see a language as close to assembly as C is, but with a cleaner disconnect with it. As an example of what I mean:
Bitfields would be really handy when writing a device driver. A basic example of the difference they would make is "if (reg.field == VAL)" vs. either "if (reg & MASK == VAL)" or "if (GET_FIELD(reg) == MASK)". But you can't use them for that purpose because their layout is implementation defined.
There would be a noteworthy performance benefit to pay for portable bitfields, but I think it would be worth it. Right now I have to write ugly code if I want any attempt at portability (always). I'd much rather have to write ugly code where I need speed (sometimes).
I'd love to hear if anyone else has any ideas on this matter (BTW it seems like D solves many of the problems I've been thinking about, but it's memory managed).
In the Pascal family of languages, developers tend to think first about having something working and then optimize if not fast enough.
In the C and C++ communities, developers tend to micro-optimize every line of written code, even before knowing if it makes sense to do so.
I've been coding C++ for a few years now and it still baffles me most of the time. If I take a break from it for a bit, I have to re-learn how pointers function every time.
In 'foo(b++, a++ && (a+b));' the sequence point introduced by the '&&' only makes the 'a' see the effect 'a++' but the 'b' might not see the 'b++' (function args are not sequenced in C nor C++).
Put another way, the compiler can arbitrarily interleave the subexpressions within separate arguments:
foo( a() && b(), c() && d() );
The compiler is free to invoke these functions in the order a(), c(), d(), b() if it pleases. foo(std::unique_ptr<T1>{new T1 },
std::unique_ptr<T2>{new T2 });
This one's from GotW #102.But I'm very aware of "smart code" and we shouldn't be writing code that relies on the details (especially ones that might change between compilers)
Hermione (I'm sure that's her) rates her C++ knowledge at 4-5, and Stroustrup himself at 7!?
Bullshit. Either they are poorly calibrated, or they are displaying false modesty. Sure, they probably still have plenty to learn about C++, but come on, Hermione is already at the top 97% in terms of language lawyering.
Wanting to be stronger is good. Not realizing you're already quite strong is not so good.
http://hyperboleandahalf.blogspot.no/2010/02/boyfriend-doesn...
Among professional programmers, at what percentile are you, in terms of C++ knowledge?
Of all there is to know about C++, how much (in%) do you know?
Your answer will depend a great deal on which question you believe this is about.
After some time, 7, then 5
But then I got back to 6.
The more you know, the less you know!
http://stackoverflow.com/questions/16115713/how-pony-orm-doe...
This simplifies step one of the Pony ORM author's answer.
Even C++ programmers, at least the ones that had good fortune to have time away from C++, agree that C++ gotchas are unforgivable.
But in general, C++ and JS communities have the biggest cases of Stockholm Syndrome I've seen.
Disclaimer: I'm a recovering C++ programmer.
C++ has a different niche. On top of that, everyone acknowledges the warts of C++, but they're largely necessary for either backward compatibility or performance reasons.
The best exercise to start with is this one on creating the smallest ELF executable possible:
http://www.muppetlabs.com/~breadbox/software/tiny/teensy.htm...
And after that I'd do the same thing on Windows with a PE file.
Then try to break things. Try to write a C program that has a buffer overflow bug in it and exploit that to change the flow of program execution. Then exploit the same bug to open an xterm or calc.exe.
Getting really familiar with a good debugger will help you a lot with these things.
After you have a decent understanding of program loading and execution, you can dive into the more language lawyery things they're talking about so you can answer the questions about the nineteen different meanings of "static" like the hacker girl. With C, that's a worthwhile goal. With C++, good fucking luck.
http://www.amazon.com/Expert-Programming-Peter-van-Linden/dp...
The presentation's "Deep C" pun is a reference to this book, and if you make it all the way to slide 444 you'll see it mentioned as further reading. It's a wonderful book for understanding C (not C++) and what's really going on.
Though the language is full of horrors, I still quite enjoy C++ (esp. with C++11 and am looking forward to some C++14 features...)
Keep in mind though, that knowing about evaluation order, and stack frames, and sequencing, and linker optimizations is all great and all, but I definitely consider it icing on the cake for a working software engineer for most positions. If you're a senior guy, sure, you should know this stuff. But the first thing to do is learn to actually write programs well. Authoring your own projects and contributing to open source is a great way to do it.
I learned all my "deep C" on the rabbit. The architecture is based on Z80 and the compiler is a really shitty non-standard C compiler with some custom syntax.
You will very quickly learn to throw away many assumptions you might have about how things work. Learn how to write mutithreaded code without an operating system or a thread library (hint: co-routines). Learn how memory layout, very quickly get a feel for the time/space tradeoffs of different algorithms, etc.
The slide deck says, the only way is the experience. You need to read and learn endlessly over the years.
Like you I hope there was one book, that does this. Makes a nice idea for a community project. Compiling such massive wisdom into a book can't be done by a person alone.
Currently its a bit like alchemy, looks very little chemistry and much more magic. And to learn there are hardly any resources beyond your regular C books, which more or less keep talking of the same things.
1) learn to program in assembly 2) inspect the assembly generated from your compiler 3) trace a call all the way from user land down into the kernel, and back up 4) don't forget networks. Whether that means a single wire instrumented with a 'scope, or a network switch instrumented with wireshark, knowing how data gets around is important.
Until you can do that on your platform, you don't really understand it (which is fine for many jobs out there, I'm just addressing the question of how to learn it).
I suspect the shortest and best path is to write your own (toy/emulator) assembler and compiler, and to do some embedded work with no OS. That'll at least expose you to all the various issues; you might not implement register allocation well (or at all), but you'll have had to think about it and the implications (example: passing function parameters on the stack vs in registers). That's one of the advantages of a University education to my way of thinking. You will be put through all these paces if it isn't just a Java accreditation program.
There are good books documenting, say, the Windows or Linux internals, so you can gain a good understanding of things like virtual memory, device drivers and so on.
It is not as hard as it might seem. This is all basic information, one fact built on top of another. I have no idea how the ARM pipeline works, but that doesn't worry me. I know what pipelines are, some of the tradeoffs they have, and if I need to get down and dirty with a cute little ARM processor I'll know what I need to learn. You can feel sort of overwhelmed if you try to think how pressing on these keys get turned into truetype fonts on a graphical screen, and then sent across the internet to anyone bored enough to read my ramblings, while at the same time my computer updates the clock, checks my mail, and does a dozen other things, but piece by piece it is all discrete, well contained, and completely understandable.
Reading Stroustrup is pretty much mandatory to understand C++. I personally wouldn't bother with the standards (other than a brief go-over); so much of the code in those slides are just things you should never, ever program. Which is to say I disagree with the slides in large measure; it's almost suspicious to have some of that knowledge! Either your coworkers are writing truly atrocious code, or you are spending time learning corners of the language which is time that almost certainly could be better spent elsewhere, such as learning about memory mapped files or something. As others have said, knowing where the boundaries of where undefined behavior lies is important, but knowing how to avoid straying even close to that territory is far more important (always initialize variables. Even static ones! You are going to cost some sorry bastard half a day of debugging when they erroneously think the problem in the code is an uninitiated variable and keep chasing that until they finally figure it out)
Regular use of a language can build certain kinds of knowledge about the internals, but it won't be as well-rounded a study as actually working with them directly.
While I can follow most of the C problems, C++ has never interested me. As as a Common Lisp programmer, all of this seems quite demented. By reading the slides I have learned about some interesting optimization concepts, and wondered how CL compilers do it. But honestly for 99% of my Job I couldn't imagine using something like C.
Edit:grammar.
And it's pretty simple to avoid all these warts, just... you know don't write them. On the other hand I don't have any experience with old C++ codebases, which is where you're likely to find this hellish stuff.
Ergo, no one is actually a real expert in a specific language.
Stroustrup is an expert at nailing legs to a dog and calling it an octopus.
No, they are experts on the implementations they created, not on the other ones that were created.
A::A() : px(new ClassX), py(new ClassY) { }
The majority of this is undefined behavior or implementation details of your compiler. If you rely on that, I don't want your code anywhere near my machine.
On the other hand, I can and have written large, complex games that run well and appear to be just fine from the user's point of view. I can seemingly write reams of code without ever thinking of such issues, except when they trip up the compiler or cause detectable bugs. I find I very, very seldom need to think of the lawyerly issues at work, or in personal projects, and to my knowledge, nobody has suffered materially for it. (Granted, I am writing games and not NASA software.)
Whenever these types of issues come up, I always think: "Wow, I know less than I thought. This seems serious?" And yet I never seem to see real-world consequences. This seems strange to me.
"What is the point of having a virtual destructor on a class like this? There are no virtual functions so it does not make sense to inherit from it. I know that there are programmers who do inherit from non-virtual classes, but I suspect they have misunderstood a key concept of object orientation. I suggest you remove the virtual specifier from the destructor, it indicates that the class is designed to be used as a base class - while it obviously is not."
She's right that the class described on the slide probably shouldn't have a virtual destructor.
A base class should have a virtual destructor if and only if objects of its derived class are to be deleted through base class pointers.
The following four statements are wrong:
1. A class with other virtual functions should have a virtual destructor.
2. A class without other virtual functions should not have a virtual destructor.
3. A class designed to be a base class should have a virtual destructor.
4. You shouldn't inherit from classes that don't have virtual functions.
It's a narrow-minded to say that programmers who inherit from "non-virtual" classes have "misunderstood a key concept of object orientation." Which key concept is that, by the way? Object orientation isn't the be all and end all of C++. There are reasons to use inheritance that have nothing to do with run-time polymorphism. Maybe you just want to reduce redundancy and organize your data types in terms of each other.
Another gripe is that on slide 369, she says:
"When I see bald pointers in C++ it is usually a bad sign."
Naked pointers should usually be avoided for memory management. That's true. But they make great iterators, and they're useful, along with references, for passing objects to functions.
On the other hand, seeing the keywords "new" and "delete" in code is usually a bad sign. Resources (not just memory) should be managed by resource management classes. If you try to do it manually, especially in the presence of exceptions and concurrency, it's very easy to cause an inadvertent resource leak.
In C, it can be tricky to ensure that every malloc has a matching free and every fopen has a matching fclose etc. This style of resource management is error-prone, but it comes with the C territory.
Java throws its hands up in the air in disgust and basically accepts that you will generate a lot of garbage. Then it slaps a mandatory garbage collector on top to clean everything up, which I guess is an okay way to deal with memory, but it sucks for other resources that aren't managed by the garbage collector.
C++, on the other hand, makes it easy to write garbage-free code. But it's only easy if you avoid using new or similar functions that allocate resources such as file descriptors, locks, network sockets etc.
If you see people using "new", there's a good chance they're misguided C or Java programmers.
C programmers often think they should use new/delete as often as they used malloc/free in C. Java programmers sometimes try to use new every time they instantiate an object. Java did steal the new keyword from C++, and it's perfectly normal for Java code to be riddled with new's because there's no other way to do it in Java. But these styles lead to disaster in C++.
Because we 1) for some reason or another we _must_ use C or C++, 2) we're coding for a single CPU platform, and 3) we need to get the friggin' job done in this century.
Sure. What about code written by others? What about when you're debugging and something just isn't working the way you expect because you've unknowingly stepped outside that subset? (Most of the important concepts here aren't syntactic constructions you can avoid trivially by typing Y instead of X, and may or may not be detectable statically). Also, maybe you're missing opportunities to better insure your code is correct at compile time - I'm not sure whether it falls within your subset, but for instance I recently put together a macro that asserts statically that two expressions have the same type (and without any runtime cost).
"I've also seen some C 'experts' fail to come to a proper solution even if they know all the features."
Obviously. Knowledge of the spec and your compiler aren't the only thing that matter, by a long shot. I'm just saying they most emphatically are helpful.
It is made up of syntax and behaves in a certain way. If you read the documented features you'll know all of it. There may be some undocumented things which may not be of utmost importance. However, this is just the base of programming. And when you start dealing with 5-6 languages you can forget things if you don't have a good memory but the good thing is it's all documented and known features. Whereas while programming, one can often be involved in solving unknown problems, which can't be figured out just by googling.
There really is a lot more to programming than just algorithms, UI design or syntax.
Especially today some people think "great programmers" are the ones who knows all the "fashionable" tools and frameworks, wants to abstract everything (god help him if he install vim in his test environment without a fab file that configures chef) and maybe worships Uncle Bob
Knowing how to use Redis is cool, do you know what's even cooler? Being able to write it (or at least knowing how it works more or less)
Anyway the programming is such a multi-dimensional evaluation problem that you're just as wrong as you are right.
Think about it - there are so many things programmer SHOULD KNOW it isn't possible to know them all. Let's name them some of them - good architecture style, good API writing skills, good code writing style, good documentation style, company's internal workflow with all its nuances, compiler standards, compiler optimization, non-standard compiler behavior, version control system API, version control innards, database system, file system, low level operating system implementation, kernel innards, non-standard OS behavior, hardware knowledge, driver knowledge, non-standard hardware behavior, hardware compatibility, software-hardware compatibility, etc. (This doesn't cover assembler knowledge, binary represntation of important formats).
You simply have to choose to not learn everything about certain system to be good at other things.
Each programmer is great in their own right.
Some of them, however, are terrible at programming.
That's fine if you just need a nice picture to hang on your wall, but probably won't work if you want a genre defining piece that that redefines how artists think about the color blue.
- 'string literal quotes' - "inexact quotes which ${evaluate things}" - `ticked off sarcastic quotes (quoting someone back at themselves)` - << EOF quotes about shell gurus EOF - ''' string literals with 'subquotes' which aren't escaped ''' - > quoted comments by oldschool BBS junkies
:-)
Deleted comment
Deleted comment
Yes, Dynamic C had built-in coroutines :) I think it's hard for me to use threads now because of how convenient a single-threaded main loop with coroutines is. A lot of concurrency "problems" vanish when you realize you as the programmer are controlling the context-switching and that race conditions are implicitly not possible then.
I went so far as to write my code so that it compiled in GCC and Dynamic-C to speed up development time (not needing to wait for the 2 minutes it takes to compile and reflash), which meant abandoning all those things and re-implementing them in a cross-platform way buried in a mountain of #ifdefs
Although I can't remember all the details now (5 years since I touched that compiler) but I still have a lot of deep seated hatred towards it for reasons I can't remember clearly :)
For this reason, a lot of compilers have options to not strictly enforce the aliasing rules
[edit] Also C and C++ are both permitted to reorder structs, it's just that they don't because that's the easiest way to follow the standard.
A structure type describes a sequentially allocated
nonempty set of member objects (and, in certain
circumstances, an incomplete array), each of which
has an optionally specified name and possibly
distinct type.
and §6.5.8 says: If the objects pointed to are members of the same
aggregate object, pointers to structure members
declared later compare greater than pointers to
members declared earlier in the structure,…§9.2.12 says:
Nonstatic data members of a (non-union) class declared without an
intervening access-specifier are allocated so that later members
have higher addresses within a class object. The order of allocation
of nonstatic data members separated by an access-specifier is unspecified (11.1)
This means, that under a strict reading, even the following struct can have its members reordered: struct foo {
int a;
public:
int b;
}> Nonstatic data members of a (non-union) class with the same access control (Clause 11) are allocated so that later members have higher addresses within a class object. The order of allocation of non-static data members with different access control is unspecified (11).
Just type casting a struct to a char-pointer and dumping it straight out on the network or to persistent storage is code you see every day.
> Within a structure object, the non-bit-field members and the units in which bit-fields reside have addresses that increase in the order which they are declared.
Is that the same standard that says it is illegal to write to one entry and read from another in a union* ?
* Note : special situation of identical fields in a structure
If the member used to access the contents of a union object
is not the same as the member last used to store a value in
the object, the appropriate part of the object representation
of the value is reinterpreted as an object representation in
the new type as described in 6.2.6 (a process sometimes called
"type punning"). This might be a trap representation.Is that feasible? Doing so would require that the order of the member addresses not correspond to their layout in memory.
"If you compile in debug mode the runtime might try to be helpful and memset your stack memory to 0"
This is a retarded explanation (pages are set to 0 when they are recycled by the OS so you don't end up having data from dead processes mapped in your memory, with all the security implications). Also, actually randomizing memory in a debug context would actually be more helpful to trigger those initialization bugs...
People that think they know everything are a lot more dangerous than people actually aware of their limitation and safely working within them.
IOW, I've worked in environments where, yes, the value in memory is based on what the compiler does in debug mode, not what the OS might be doing. I'd say the girl's answer is, while not exhaustive, certainly correct. Without more context, we don't know why the value was 0.
What really annoyed me is the suggestion that zeroing is more helpful than randomization in a debug context. Randomizing everything not defined by the standard is a good way to trigger lots of bugs based on false assumptions...
The second is revealing incorrect assumptions / fuzzing / stress testing; zeroing is substantially less helpful than randomization in this context.
Ideally, both options would be available. When someone says "give me a debug build", I think it's probably correct that they are caring more about the former, but really it should be clear and controllable.
Predictability in this context means simplifying the problem. If uninitialized memory is nulled, then I know what I'm looking for/at more quickly. Likewise if it's set to any other particular, known value. I know it will (or won't) be failing null checks, and I know that it will (or won't) be segfaulting if dereferenced. This helps with characterizing and fixing bugs, (related to but distinct from detecting bugs in the first place). Relying on this to make the program work is bad, but that's not the same thing.
i = i + 1;
i += 1;
i++;
++i;
a[i]
*(a + i)
a->foo
(*a).fooAs for assembly, there's been several lists of "free" computer books post that have included assembly tutorials.
It uses MIPS assembly, which you can run on a simulator. I don't think many machine architecture classes teach x86 because it's more complex. The MIPS knowledge from that course has translated easily to x86 in my tiny experience of peering into my compiler's assembly output.
the most productive code i have ever written generally involves me working around the constraints of the language to implement a paradigm which is missing at compile-time, or juggling macros and templates so that i can reduce boilerplate code down to a template with a macro to fill the gaps the template is too featureless to give me (vice versa, the template is there because macros aren't complete enough either).
its good to understand this deep language stuff though because you can understand why C/C++ are limited. for instance the C sequence points limit the compiler in its ability to perform optimisation, as do struct layout rules and many of the other weird and wonderful specifics...
what saddens me most though is that nobody has offered anything to improve C and C++ in these areas which matter most to me... its not even hard. just let the compiler order structs because most programmers don't understand struct layout rules.
its not a good thing that these things are so explicitly specified for the language - its gimping the compilers, which is limiting me. also it results in pointless interview questions about sequence points.. :P
Deep understanding of the language used will make you do a better job.
This reminds me of the story that I've run into once. The junior developer wanted to raise PHP's memory_limit parameter because his code crashed almost every time while writing big file content to the output. He didn't know what output buffering is and that he can turn it off and print the file directly to the output. :D
Many of these fall under the heading of opportunities to apply the Pythonic design criterion "refuse the temptation to guess":
1) in C, according to one of the comments, it was claimed (i didn't check) that if you declare your own printf with the wrong signature, it will still be linked to the printf in the std library, but will crash at runtime, e.g. "void printf( int x, int y); main() {int a=42, b=99; printf( a, b);}" will apparently crash.
-- A new programming language might want to throw a compile-time error in such a case (as C++ apparently does, according to the slides).
2) In C, depending on compiler options, you can read from an uninitialized variable without a warning
-- A new programming language might want to not auto-initialize any variables, and to throw a compile-time error if they are used before initialization.
3) In C, code like "int a = 41; a = a++" apparently compiles but leaves 'a' in an undefined state because "you can only update a variable once between sequence points" or it becomes undefined, but on many compilers works anyway. A sequence point is "a point in the program's execution sequence where all previous side effects SHALL have taken place and all subsequent side-effects SHALL NOT have taken place".
-- A new programming language might want to throw a compile-time error in such a case
4) In C, the evaluation order of expressions is unspecified. so code like "a = b() + c()" can call b() and c() in any order. If they have side effects then this might matter, yet no compiler error is given. However, the evaluation order of a() && b() IS specified.
-- A new programming language might want to throw a compile-time error when side-effectful code is called in context in which the order of evaluation is unspecified.
Other miscellaneous gotchas:
5) In C, static vars (but not other vars) are initialized to 0 by default.
-- A new programming language might want to either auto-initialize all variables, or to not auto-initialize any variables,
6) The presentation says "C has very few sequence points. This helps to maximize optimization opportunities for the compiler.". This is a tradeoff between optimization vs. principal of least surprise.
-- A new programming language which wanted to make things as simple as possible would maximize 'sequence points', putting them in between practically every computation step. But some new programming languages would choose to minimize sequence points in order to allow the compiler to optimize as much as possible.
7) The presentation says that the standard says that source code must end with a newline.
-- Imo that's a bit pedantic and the ideal programming language would not care if code ended in a newline.
8) In one context (inside a function), the 'static' keyword is used to make a variable persist across calls to that function. But in another context (outside of any function), the same 'static' keyword is used as an access modifier to define visibility to other compilation units!
-- Using the same keyword for two different (albeit related) purposes is confusing. A new programming language might either drop one of those features entirely, or have a distinct keyword for it.
That or they're assholes.
Either way, walk.