How to C in 2016
matt.sh
matt.sh
The best we can do is write simple, understandable code with as few indirections and as little undocumented magic as possible.
Applies very well to all programming.
Of course this is a disadvantage if you don't have sufficient experience in the problem domain. Sometimes I resort to writing prototypes in a script language in that case, translating the best design to the low-level language.
While that certainly can't be a bad thing, you just have to keep in mind that it forces you to think about a specific set of details – the ones that matter for the machine ("how many bytes will the binary representation of this username require") – and it may fool you into forgetting to think about higher-level concerns ("Does å compare equally to å?").
Avoiding unnecessary abstractions is important but at the same time those abstractions were invented for a reason. Basically the Go vs generics debate, except even more rudimentary. It's fine if your code doesn't need those abstractions, it sucks badly if it does (eg. gobject)
Just look at any C collection library for things like maps compared to C++.
Algol, PL/I, Mesa, Cedar, Modula-2 and many other languages of similar age or older than C, do offer both higher and lower level mechanisms.
The tradeoff is that the task at hand tends to have more implementation details in it than in higher level languages, but as you said, this isn't always bad.
The same code would be much simpler with C++ templates, but the "C being simple" really translates into limited, then you have that some have who never learned how to use macros and this unfounded and misunderstood fear of goto that gives you a nice long lines of error checking, each that that are duplicated if statements character by character where a goto and a single error handler would have worked.
Perhaps this is an exception, but low-level in my case did not correspond to simple code.
Although developers skipping return value checks is true in most languages.
#1 The submitted article is very good. #2 I'd add to it, avoid pointer arithmetic when possible. Use array notation instead. Typically the difference in execution speed is zero to nil. #3 I'd really like a pragma that forces an exception for unchecked return values. Something like
int DoSomeThing(/* whatever*/) #pragma exit "unchecked non_zero"
Meaning if DoSomething() doesn't return zero the program bails.Yes, which is why I disagree with Go's approach for error codes as return values. I'm not saying that exceptions are the right solution everywhere, but if you return a error code, you should make it pretty hard to ignore it. In languages with Algebraic Data Types it's very easy to use Maybe or a Error(code) | Success(result) type. You've then got to explicitly deconstruct that type to get the result, so it's a bit harder to just forget handling the error.
> The first rule of C is don't write C if you can avoid it.
None of which removes the need to write simple, understandable code without undocumented magic. But we should not be throwing up our hands and treating invalid accesses, memory safety failures, buffer overflows, SQL injection, XSS, or a whole host of common failures as inevitable. It really is possible to do a lot better than most programs and a lot better than C.
Provable correctness "...has not been tried and found wanting; it has been found difficult and not tried."
We're writing programs that are tens of millions of lines of code (OSes, databases, air traffic control systems). My impression of provable correctness is that the largest program that has been proven is in the area of 100,000 lines (feel free to correct if I am in error). We need to be able to prove programs two orders of magnitude bigger than that; provable correctness has not been tried for such programs because nobody wants to wait a generation or two before they get the results.
They key is then to simply avoid emergent interactions between modules/subsystems. That's pretty doable.
What doesn't happen is that people never reason about the cost of defects. Even if they do some reasoning about this, they don't apply that to other components of systems.
You code in a functional style and/or with state machines composed. They composed stuff hierarchically and loopfree where possible to facilitate analysis. Each component is small enough to analyze for all success or failure states. The results of each become "verification conditions" for things that build on it. With such methods, one can build things that are as large as you want so long as each component (primitive or composite) can be analyzed.
Not clear at what point this breaks down. I used it for drivers, protocols, algorithms, systems, networks, and so one. I brute forced it with more manual analysis because Im not mathematical or trained enough for proofs. However, what holds it back on large things is usually just lack of effort far as I can tell.
How many modern programmers hold this mistaken belief? It's false -- the Turing Halting problem (https://en.wikipedia.org/wiki/Halting_problem) demonstrates that one cannot establish that a non-trivial program is "provably correct."
It seems the OP got carried away with his rhetoric, and the way he put it is simply wrong.
The halting problem says that there exist programs that cannot be proven correct; it says that for any property you want to prove computationally, no program can give you an answer for every last program. It doesn't say anything about "non-trivial" programs. In particular, if you write programs with awareness of the halting problem in mind, you can write extremely complicated programs that are provably correct.
One straightforward way to do this is with a language that always halts. The proof assistant Coq is based on such a language, and it's powerful enough to write a C compiler: http://compcert.inria.fr/
It is certainly correct that you can't stuff an arbitrarily vexatious program through a proof assistant and get anything useful on the other side. But nobody is interested in proving things about arbitrarily vexatious programs; they're interested in proving things about programs that cooperate with the proof process. If your program doesn't want to cooperate, you don't need an exact answer; you might as well just say "No, I don't care, I'm not trusting you."
The way you prove halting is to find some integer measure for each loop (like "iterations remaining" or "error being reduced") which decreases on each iteration but does not go negative. If you can't find such a measure, your program is probably broken. (See page 12 of [1], the MEASURE statement, for a system I built 30 years ago which did this.)
The measure has to be an integer, not a real number, to avoid Zeno's paradox type problems. (1, 1/2, 1/4, 1/8, 1/16 ...)
On some floating point problems, with algorithms that are supposed to converge, proving termination can be tough. You can always add an iteration counter and limit the number of iterations. Now you can show termination.
Microsoft's Static Driver Verifier is able to prove correctness, in the sense of "not doing anything that will crash the kernel", about 96% of the time. About 4% of the time, it gives up after running for a while. If your kernel driver is anywhere near close to undecidability, there's something seriously wrong with it.
In practice, this is not one of the more difficult areas in program verification.
[1] http://www.animats.com/papers/verifier/verifiermanual.pdf
No, but its source does. The source for the Turing Halting problem is Gödel's incompleteness theorems, which do specify that the entities to which it applies are non-trivial. The Turing Halting problem, and Gödel's incompleteness theorems, are deeply connected.
> One straightforward way to do this is with a language that always halts.
That doesn't avoid the halting problem, it only changes the reason for a halt. And it cannot guarantee that that program will (or won't) halt, if you follow -- only its interpreter.
> In particular, if you write programs with awareness of the halting problem in mind, you can write extremely complicated programs that are provably correct.
You need to review the Turing Halting problem. No non-trivial computer program can be proven correct -- that's the meaning of the problem. It cannot be waved away, any more than Gödel's incompleteness theorems can be waved away, and for the same reason.
Gödel doesn't say anything about "non-trivial" any more than Turing does. His first theorem is a "there exists at least one statement" thing, just like the halting problem. Like the halting problem, the proof is a proof by construction, but that statement is not a statement that anyone would a priori are to prove, just like the program "If I halt, run forever" is not a program that anyone would particularly want to run. So while yes, they're deeply connected, from the point of view of proving actual systems correct, it's not clear they matter. And neither theorem says anything about "non-trivial" statements/programs (just about non-trivial formalizations and languages).
Gödel's second theorem is super relevant, though: it says that any consistent system cannot prove its own consistency. But that's fine. Coq itself is implemented in a Turing-complete language (ML), not in Coq. It would be nice if we could verify Coq in Coq itself, but since Gödel said we can't, we move on with life, and we just verify other things besides proof assistants. (And it seems to be worthwhile to prove things correct in a proof assistant, even without a computer-checked proof of that proof assistant's own correctness.)
This is not waving away the problem, it's acknowledging it and working within the limits of the problem to do as much as possible. And while that's not everything, it turns out a lot of stuff of practical importance -- perhaps everything of practical importance besides proving the system complete in its own language -- is possible.
Source: https://en.wikipedia.org/wiki/Halting_problem
Quote: "Rice's theorem generalizes the theorem that the halting problem is unsolvable. It states that for any non-trivial property, there is no general decision procedure that, for all programs, decides whether the partial function implemented by the input program has that property. (A partial function is a function which may not always produce a result, and so is used to model programs, which can either produce results or fail to halt.) For example, the property "halt for the input 0" is undecidable. Here, "non-trivial" means that the set of partial functions that satisfy the property is neither the empty set nor the set of all partial functions."
Source: https://en.wikipedia.org/wiki/G%C3%B6del%27s_incompleteness_...
Quote: "Gödel's incompleteness theorems are two theorems of mathematical logic that establish inherent limitations of all but the most trivial axiomatic systems capable of doing arithmetic."
> It would be nice if we could verify Coq in Coq itself ...
I was planning to bring this issue up, but you've done it for me. The simplest explanation of these two related theorems involves the issue of self-reference. Imagine there is a library, and a very conscientious librarian wants to have a provably correct account of the library's contents. The librarian therefore compiles an index of all the works in the library and declares victory. Then Gödel shows up and says, "On which shelf shall we put the index?"
In the same way, and for the same reason, non-trivial logical systems cannot check themselves, computer languages cannot validate their products (or themselves), and compilers cannot prove their own results correct.
> And it seems to be worthwhile to prove things correct in a proof assistant, even without a computer-checked proof of that proof assistant's own correctness.
Yes, unless someone is misled into thinking this means the checked code has been proven correct.
> This is not waving away the problem ...
When someone declares that a piece of code has been proven correct, in point of fact, yes, it is.
> all but the most trivial axiomatic systems
This is about non-trivial properties on arbitrary programs, and non-trivial axiomatic systems proving arbitrary statements. It is not about non-trivial programs or non-trivial statements.
Rice's theorem states that, for all non-trivial properties P, there exists a program S where you cannot compute P(S). (The Wikipedia phrasing puts the negative in a different spot, but this is de Morgan's law for quantifiers: ∀P ¬∀S P(S) is computable = ∀P ∃S ¬ P(S) is computable.)
Gödel's first theorem is that, for all non-trivial formalizations P, there exists a statement S where you cannot prove P in S.
You are claiming, for all non-trivial properties P, for all non-trivial programs S, you cannot compute P(S); for all non-trivial formalizations P, for all non-trivial statements S, you cannot prove P in S.
This is not what Rice's or Gödel's theorems are, and not what the texts you quoted say.
> Then Gödel shows up and says, "On which shelf shall we put the index?"
You keep it outside the library. The index tells you where every single book is, other than the index.
In particular, Gödel claims that, if the index has any chance of being correct, it must be outside the library. The fact that Coq cannot be proven in Coq does not decrease our confidence that it is correct; in fact, Gödel's second theorem means that if it could be proven in itself, we'd be confident it was wrong!
> You keep it outside the library. The index tells you where every single book is, other than the index.
This approach would have saved Russell and Whitehead's "Principia Mathematica" project, would have placed mathematics on a firm foundation, but it suffers from the same defect of self-reference -- it places the index beyond validation. Russell realized this, and recognized that Gödel's results invalidated his project.
> The fact that Coq cannot be proven in Coq does not decrease our confidence that it is correct;
Yes, true, but for the reason that we have no such confidence (and can have no such confidence). We cannot make that assumption, because the Turing Halting problem and its theoretical basis prevents it.
Source: http://www.sscc.edu/home/jdavidso/math/goedel.html
Quote: "Gödel hammered the final nail home by being able to demonstrate that any consistent system of axioms would contain theorems which could not be proven. Even if new axioms were created to handle these situations this would, of necessity, create new theorems which could not be proven. There's just no way to beat the system. Thus he not only demolished the basis for Russell and Whitehead's work--that all mathematical theorems can be proven with a consistent set of axioms--but he also showed that no such system could be created." (emphasis added)
One particular project. The specific project of formalizing mathematics within itself is unachievable. That doesn't mean we've stopped doing math.
And this is in fact the approach modern mathematics has taken. The fact that one particular book (the proof of mathematics' own consistency) can't be found on the shelves doesn't mean there isn't value in keeping the rest of the books well-organized.
> Yes, true, but for the reason that we have no such confidence (and can have no such confidence). We cannot make that assumption, because the Turing Halting problem and its theoretical basis prevents it.
No! Humans aren't developed in either Coq or ML. A human can look at Coq and say "Yes, this is correct" or "No, this is wrong" just fine, without running afoul of the halting problem, Gödel's incompleteness theorem, or anything else. The fact that you can have no Coq-proven confidence in Coq's correctness does not mean that you can have no confidence of Coq's correctness.
(Of course, this brings up whether a human should be confident in their own judgments, but that's a completely unavoidable problem and rapidly moving to the realm of philosophy. If you are arguing that a human should not believe anything they can convince themselves of, well, okay, but I'm not sure how you plan to live life.)
This is like claiming that, because ZF isn't provably correct in ZF (which was what Russell, Whitehead, Gödel, etc. were concerned about -- not software engineering), mathematicians have no business claiming they proved anything. Of course they've proven things just fine if they're able to convince other human mathematicians that they've proven things.
> One particular project.
A mathematical proof applies to all cases, not just one. Gödel's result addressed Russell's project, but it applies universally, as do all valid mathematical results. Many mathematicians regard Gödel's result as the most important mathematical finding of the 20th century.
> That doesn't mean we've stopped doing math.
Not the topic.
> A human can look at Coq and say "Yes, this is correct" or "No, this is wrong" just fine, without running afoul of the halting problem, Gödel's incompleteness theorem, or anything else.
This is false. If the Halting problem is insoluble, it is also insoluble to a human, and perhaps more so, given our tendency to ignore strictly logical reasoning. It places firm limits on what we can claim to have proven. Surely you aren't asserting that human thought transcends the limits of logic and mathematics?
> Of course they've proven things just fine if they're able to convince other human mathematicians that they've proven things.
Yes, with the caution that the system they're using (and any such system) has well-established limitations as to what can be proven. You may not be aware that some problems have been located that validate Gödel's result by being ... how shall I put this ... provably unprovable. The Turing Halting problem is only one of them, there are others.
Given that both logic and mathematics are creations of human thought, is this really so untenable?
What? I have no idea what you mean by "applies universally" and "all valid mathematical results."
A straightforward reading of that sentence implies that Pythagoras' Theorem is not a "valid mathematical result" because it applies only to right triangles. That can't be what you're saying, is it?
I am acknowledging that Gödel's second theorem states that any mathematical system that attempts to prove itself is doomed to failure. I am also stating that Russell's project involved a mathematical system that attempts to prove itself. Therefore, Gödel's second theorem applies to Russell's project, but it does not necessarily say anything else about other projects.
> Many mathematicians regard Gödel's result as the most important mathematical finding of the 20th century.
So what? This is argumentum ad populum. How important the mathematical finding is is irrelevant to whether it matters to the discussion at hand.
> This is false. If the Halting problem is insoluble, it is also insoluble to a human
OK, so do you believe that modern mathematics is invalid because Russell's project failed?
Obviously, the solution is to stop writing arbitrary programs. As Dijkstra said long ago, we need to find a class of intellectually manageable programs - programs that lend themselves to being analyzed and understood. Intuitively, this is exactly what programmers demand when they ask for “simplicity”. The problem is that most programmers' view of “simplicity” is distorted by their preference for operational reasoning, eschewing more effective and efficient methods for reasoning about programs at scale.
Can you clarify your last sentence? What do you mean by "preference for operational reasoning"? What's an example of the more effective methods for reasoning about programs at scale that you're contrasting this to?
By “more effective methods for reasoning about programs at scale”, what I mean is analyzing the program in its own right, without reference to specific execution traces. There are several kinds of program analyses that can be performed without running the program, but, to the best of my knowledge, all of them are ultimately some form of applied logic.
Not really. It's also possible to reason non-operationally about imperative programs, e.g., using predicate transformer semantics. It's also possible to reason operationally about non-imperative programs.
You're overlooking the issue of self-reference. A given computer language can validate (i.e. prove correct) a subset of itself, but it cannot be relied on to validate itself. A larger, more complex computer language, created to address the above problem, can check the validity of the above smaller language, but not itself, ad infinitum. It's really not that complicated, and the Halting Problem is more fundamental than this entire discussion seems able to acknowledge.
Also, it seems you misunderstand Gödel/Turing/etc.'s result. What a consistent and sufficiently powerful formal system can't do is prove its own consistency as a theorem. Obviously, it isn't sustainable to spend all our time inventing ever more powerful formal systems just to prove the preceding ones consistent. At some point you need to assume something. This is precisely the rôle of the foundations of mathematics (or perhaps I should say a foundation, because there exist many competing ones): to provide a system whose consistency need not be questioned, and on top of which the rest of mathematics can be built.
In any case, we have gone too far away from original point, which is that, if arbitrary programs are too unwieldy to be analyzed, the only way we can hope to ever understand our programs is to limit the class of programs we can write. And this is precisely what (sound) type systems, as well as other program analyses, do. Undoubtedly, some dynamically safe programs will be excluded, e.g., most type systems will reject `if 2 + 2 == 4 then "hello" else 5` as an expression of type `String`. But, in exchange, accepted programs are free by construction of any misbehaviors the type system's type safety theorem rules out. The real question is: how do we design type systems that have useful type safety theorems...?
The language comes from Gödel's incompleteness theorems, in which there exist true statements that cannot be proven true, i.e. validated.
A given non-trivial computer program can be relied on to produce results consistent with its definition (its specification), but this cannot be proven in all cases.
It's not very complicated.
> This is precisely the rôle of the foundations of mathematics (or perhaps I should say a foundation, because there exist many competing ones): to provide a system whose consistency need not be questioned, and on top of which the rest of mathematics can be built.
This is what Russell and Whitehead had in mind when they wrote their magnum opus "Principia Mathematica" in the early 20th century. Gödel proved that their program was flawed. Surely you knew this?
> Also, it seems you misunderstand Gödel/Turing/etc.'s result.
The misunderstanding is not mine.
Provability is always relative to a formal system. The formal system you (should) use to prove a program correct (with respect to its specification) isn't the program itself.
> Gödel proved that their program was flawed. Surely you knew this?
Gödel didn't prove the futility of a foundations for mathematics. Far from it. What Gödel proved futile is the search for a positive answer to Hilbert's Entscheidunsproblem: There is no decision procedure that, given an arbitrary proposition (in some sensible formal system capable of expressing all of mathematics), returns `true` if it is a theorem (i.e., it has proof) or `false` if it is not.
If you want to lecture someone else on mathematical logic, I'd advise you to actually learn it yourself. I'm far from an expert in the topic, but at least this much I know.
Your words, not mine, but Gödel did prove that a non-trivial logical system cannot be both complete and consistent. As Gödel phrased it, "Any consistent axiomatic system of mathematics will contain theorems which cannot be proven. If all the theorems of an axiomatic system can be proven then the system is inconsistent, and thus has theorems which can be proven both true and false."
Gödel's result don't address futility, but completeness. That's why his theorems have the name they do.
> The formal system you (should) use to prove a program correct (with respect to its specification) isn't the program itself.
This refers to the classic case of using a larger system to prove the consistency of a smaller one, but it suffers from the problem that the larger system inherits all the problems it resolves in the smaller one.
> If you want to lecture someone ...
This is either beneath you, or it should be.
I'm sorry, but 'catnaroek is completely justified here. Multiple people have told you that you are simply factually wrong in your interpretation of the halting problem, and you are acting arrogant about it. If you tone back the arrogance a bit ("Surely you knew this?", "The misunderstanding is not mine") and concede that, potentially, it's not the case that everyone else is wrong and you're right, you might understand what your error is.
It's a very common mistake and is frequently mis-taught, especially by teachers who are unfamiliar with fields of research that actually care about the implications of the halting problem and of Gödel's theorems (both in computer science and in mathematics) and just know it as a curiosity. So I've been very patient about this. But as I pointed out to you in another comment, you are misreading what these things say in a way that should be very obvious once you state in formal terms what it is that you're claiming, and it would be worth you trying to acknowledge this.
Source: https://en.wikipedia.org/wiki/Halting_problem
Quote: "Alan Turing proved in 1936 that a general algorithm to solve the halting problem for all possible program-input pairs cannot exist." (emphasis added)
> Multiple people have told you that you are simply factually wrong ...
So as the number of people who object increases, the chances that I am wrong also increases? This is an argumentum ad populum, a logical error, and it is not how science works. Imagine being a scientist, someone for whom authority means nothing, and evidence means everything.
In your next post, rise to the occasion and post evidence (as I have done) instead of opinion, or don't post.
You don't need to solve the Halting problem (or find a way around Rice's theorem, etc.) to write a single correct program and prove it so. A-whole-nother story is if you want a decision procedure for whether an arbitrary program in a Turing-complete language satisfies an arbitrary specification. Such a decision procedure would be nice to have, alas, provably cannot possibly exist.
To give an analogy, consider these two apparently contradictory facts:
(0) Real number [assuming a representation that supports the usual arithmetic operations] equality is only semidecidable: There is a procedure that, given two arbitrary reals, loops forever if they are equal, but halts if they aren't. You can't do better than this.
(1) You can easily tell that the real numbers 1+3 and 2+2 are equal. [You can, right?]
The “contradiction” is resolved by noting that the given reals 1+3 and 2+2 aren't arbitrary, but were chosen to be equal right from the beginning, obviating the need to use the aforementioned procedure.
The way around the Halting problem is very similar: Don't write arbitrary programs. Rather, design your programs right from the beginning so that you can prove whatever you need to prove about them. Typically, this is only possible if the program and the proof are developed together, rather than the latter after the former.
This much math you need to know if you want to lecture others on the Internet. :-p
> but were chosen to be equal right from the beginning
While this was almost certainly the case, it's sufficient for them to have been chosen so as to be decidably comparable.
Indeed. :-)
The system you outline works perfectly as long as you limit yourself to provable assertions. But Gödel's theorems have led to examples of unprovable assertions, thus moving beyond the hypothetical into a realm that might collide with future algorithms, likely to be much more complex than typical modern algorithms, including some that may mistakenly be believed to be proven correct.
If we're balancing a bank's accounts, I think we're safe. In the future, when we start modeling the human brain's activities in code, this topic will likely seem less trivial, and the task of avoiding "arbitrary programs," and undecidable assertions (or even recognizing them in all cases), won't seem so easy.
> This much math you need to know if you want to lecture others on the Internet.
My original objection was completely appropriate to its context.
Of course. And if a formal system doesn't let you prove the theorems you want, you would just switch formal systems.
> But Gödel's theorems have led to examples of unprovable assertions, thus moving beyond the hypothetical into a realm that might collide with future algorithms likely to be much more complex than typical modern algorithms, including some that may mistakenly be believed to be proven correct.
Programming is applied logic, not religion. And a proof doesn't need to be “believed”, it just has to abide by the rules of a formal system. In some formal systems, like intensional type theory, verifying this is a matter of performing a syntactic check.
> If we're balancing a bank's accounts, I think we're safe.
I wouldn't be so sure.
> In the future, when we start modeling the human brain's activities in code, this topic will likely seem less trivial,
I see this as more of a problem for specification writers than programmers. Since I don't find myself writing specifications so often, this is basically Not My Problem (tm). Even if I wrote specifications for a living, I'm not dishonest or delusional enough to charge clients for doing something I actually can't do, like specifying a program that claims to be an accurate model of the human brain.
> and the task of avoiding "arbitrary programs," and undecidable assertions (or even recognizing them in all cases), won't seem so easy.
I don't see what's so hard about sticking to program structures (e.g., structural recursion) for which convenient proof principles exist (e.g., structural induction).
Please, get yourself a little bit of mathematical education before posting your next reply. You're so stubbornly wrong it's painful to see.
You seem to be contradicting your original claim, that it is impossible to prove any non-trivial program correct. Is this because you believe the program to balance a bank's accounts is trivial? Do you also believe that a C compiler is trivial?
I guess we never defined what you meant by "non-trivial." If your position is that the vast majority of software in the world today is "trivial," sure, I'd agree with you. But that's not really a meaning of "trivial" that anyone expected.
You seem to be mixing up where the "for all" lies. The result is:
There does not exist a procedure that, for all (program, input) pairs, the procedure will say whether the program halts for that input.
{p | Ɐ(f,i): p(f,i) says whether f(i) will halt } = ∅
You are using it as if it said:For all (program, input) pairs, there does not exist a procedure that will say whether the program halts for that input.
Ɐ(f,i): {p | p(f,i) says whether f(i) will halt } = ∅
"post evidence (as I have done)"What is at issue is your incorrect assertions about what your evidence says. We read it, it contradicts you, everyone points this out, and you just ignore it and get louder.
Not at all. My original objection was to the claim that a program could be proven correct, without any qualifiers. Any program, not all programs. As stated, it's a false claim.
Your detailed comparison doesn't apply to my objection, which is to an unqualified claim.
> everyone points this out ...
Again, this is a logical error. Science is not a popularity contest, and in this specific context, straw polls resolve nothing.
> you just ignore it and get louder.
I wonder if you know how to have a discussion like this, in which there are only ideas, no personalities, and no place for ad hominem arguments?
Let's be precise.
Let P be the set of all programs.
Let C(s) ⊂ P be the set of programs correct relative to specification s.
As I read the original statement, it was saying, {p ∈ C(s) | p, s ⊢ p ∈ C(s)} is large enough to be useful
I am not sure whether you read it differently.> As stated, it's a false claim.
You have said this repeatedly, but all you have done in support of it is point at the Halting problem. The Halting problem says:
∃p ∈ P s.t. ¬(p ⊢ p ∈ C(halts)) ∧ ¬(p ⊢ p ∉ C(halts))
You have claimed this eliminates the possibility of proving correct any non-trivial program:> [T]he Turing Halting problem (https://en.wikipedia.org/wiki/Halting_problem) demonstrates that one cannot establish that a non-trivial program is "provably correct."
I see no way to read that but as a claim that:
(∃p ∈ P s.t. ¬(p ⊢ p ∈ C(halts)) ∧ ¬(p ⊢ p ∉ C(halts))) → ((p,s ⊢ p ∈ C(s)) → p is trivial)
Please supply a proof.Note that it's quite true that
(∃p ∈ P s.t. ¬(p ⊢ p ∈ C(halts)) ∧ ¬(p ⊢ p ∉ C(halts))) → ((p,s ⊢ p ∈ C(s)) → *s* is trivial)
proved usually by reducing s to "halts" for any non-trivial s.> Again, this is a logical error. Science is not a popularity contest, and in this specific context, straw polls resolve nothing.
I am not deciding it by poll. I am pointing out that there are objections to your reasoning that you have not addressed and persist in not addressing - despite them being made abundantly clear, repeatedly.
> I wonder if you know how to have a discussion like this, in which there are only ideas, no personalities, and no place for ad hominem arguments?
You're accusing me, as a person, of not understanding that a discussion like this has no place for accusations against a person? My complaint was not about you as a person, but about the form of your arguments in this discussion.
I have already done so. Did you read my comment about existential vs. universal quantifiers? Can you respond to it?
This was the original claim:
> It's false -- the Turing Halting problem (https://en.wikipedia.org/wiki/Halting_problem) demonstrates that one cannot establish that a non-trivial program is "provably correct."
This is the statement you quoted from Wikipedia:
> Quote: "Alan Turing proved in 1936 that a general algorithm to solve the halting problem for all possible program-input pairs cannot exist."
Do you agree or disagree that the quantifiers in these two statements are different? Do you agree or disagree that "the Turing Halting problem demonstrates that one cannot establish that every single non-trivial program is either 'provably correct' or 'provably incorrect'" is a different claim than the one originally made?
http://research.microsoft.com/apps/mobile/news.aspx?post=/en...
Also, Knuth in TAoCP spends significant time proving his programs correct.
Further, there is a compiler that is asserted to be proven correct: http://compcert.inria.fr/
Now it is possible you are working with a different definition of "proved correct" than I am thinking.
... and then someone finds a bug in the specification which is provably correctly implemented.
Using gets() to read the input can be provably correct. All you have to do is remove any requirements for security from the program's specification, and specify that the program will always be used in circumstances when the expected input datum will be less than a certain length.
... and then the hardware is buggy or failing. That RAM bit flips and you're screwed. The machine as a whole isn't provably correct.
Detecting overflows and throwing nice exceptions instead of continuing execution with garbage values isn't the same thing as ensuring correctness.
Java does array bounds checking in principle, but it is usually optimized out of the machine code. Isn't it then virtually just as susceptible to hardware error as C? And the JVM is written in C anyway.
https://www.cs.princeton.edu/~appel/papers/memerr.pdf [2003]
Which one?
https://en.wikipedia.org/wiki/List_of_Java_virtual_machines
Depending which one we are speaking about, they are implemented in Assembly, C, C++ or even Java.
As a programmer, if the specification is wrong, it's Not My Fault (tm). If you want me to write specifications, well, pay me to do it! I guarantee dramatically better results than the usual crap.
> The machine as a whole isn't provably correct.
As a programmer, defects in the machine are Not My Problem (tm). In any case, hardware vendors have a much better track record than software developers at delivering products that meet strict quality standards.
We're no longer discussing correctness, but assignment of blame.
As an aside, I want to clarify that I'm not advocating being a bad team player. The point to assigning blame isn't demonizing the developer of the faulty component, or reducing cooperation between developers of different components to the bare minimum. The point is just reducing the cost of fixing the problem.
If one program misuses another, usually the fix can be an alteration in either one or both. The used program can be expanded to accept the "misuse" which is actually legitimate but not in the requirements. Or the using program can be altered not to generate the misuse.
Sometimes what is fixed is chosen for non-technical reasons, like: the best place to fix it it is in such and such code, but ... it's too widely used to touch / not controlled by our group and we need a fix now / closed source and not getting fixed by the vendor {until next year|EVER} / ...
When in doubt, ask the specification. If there was no specification, why was any code written at all?
> If one program misuses another, usually the fix can be an alteration in either one or both.
Actually, there are three possibilities:
(0) The used program doesn't comply with its specification.
(1) The using program doesn't comply with its specification.
(2) The specifications of both programs are mutually inconsistent. In which case, it's a programming mistake to connect both programs.
> The used program can be expanded to accept the "misuse" which is actually legitimate but not in the requirements.
If it isn't in the specification, it isn't legitimate, period.
> Or the using program can be altered not to generate the misuse.
Sure.
You have some goal you're trying to achieve, or why was a specification written at all? If you find that the existing specification is not the optimal path to that goal, you may very well want to change the specification. In which case, it's entirely legitimate.
So the requirements/specification/program complex exists Just Because. Humans decided they want that: for their amusement, for the sake of supporting some enterprise or solving a problem, or to try to sell to other humans.
The complex contains several different representations because that's what it takes to bridge the gap between stating the requirements and making the machine carry them out.
A proof is something internal to that complex: that the specification corresponds to the requirements, and that the program implements the specification.
If we had just one artifact: a specification that executes, then there would be no concept of proof any more.
There is only the question whether the specification that was expressed is the one that was intended in the mind. That equivalence is no more susceptible to proof than, say, the correspondence between the Ceasar salad that the waiter just put on your table (and which you clearly specified) and the idea of whether you actually wanted one, or did you mistakenly say "Ceasar salad" in spite of having wanted a soup.
Plus, there is the external question of whether a specification has unintended consequences. That basically amounts to "you say you want that, but maybe you should revise what you want because of these bad things".
"You say you want plaintext passwords to be stored for easier recovery, and that can certainly be implemented (provably correctly, too) but consider the following ramifications ..."
It is possible for an executable functional specification to have unbearably bad performance for production use. In that case, the programmer has to reimplement the program in a more efficient (but perhaps less obvious) way and supply a proof that the reimplementation is functionally equivalent to the original executable specification.
Proofs aren't going away anytime soon.
> There is only the question whether the specification that was expressed is the one that was intended in the mind.
If other people can't bother communicating what they really want, that is absolutely Not My Problem (tm).
> Plus, there is the external question of whether a specification has unintended consequences.
If other people can't analyze the logical consequences of what they wish for, that is absolutely Not My Problem (tm). The most I can do is point at contradictions in the requirements.
I used to work in a facility where they'd test systems for radiation hardness. Get the system running, then shoot a proton beam travelling a significant fraction of the speed of light directly at the processor. It doesn't matter if the code was proven correct Dijkstra, it will eventually manifest bugs under these conditions.
[1] practically we resort to empirical testing to confirm that hardware conforms to its spec. But as your sibling says, that's no excuse not to do better in the places where we can.
I was mostly C developer (80% C, 20% C++) for 4 years in a big project (I was backend developer of ICQ Instant Messenger). It has more than 2M lines of C code. Almost all libraries was written from scratch in C language (since ICQ is very old project, many of essential stuff was written within AOL).
I felt comfortable to write high level code in C.
When I read Nginx source code (it's written in C), I see better code quality than 90% of proprietary commercial software which is written on nice and comfortable high level languages.
Good developer is able to write high quality code in C language with minimum bugs in large scale project. Regardless of how nice and high level and convenient language, bad developer will write c....y code.
You can also check my personal project which is written entirely in C. (I didn't contribute for 1.5 years to my github since I was changed countries of living, jobs etc. I hope I will return to it when my life become stable again).
P.S. To be clear, I think that C++11 is great language. But unlike many C++ developers, I love pure C.
The people and the culture/organization of the shop counts for much more than your language. Language counts for something, but it's not where you get the most bang for your buck, and it shouldn't be your first priority!
That said you can create message-passing systems in either - Objective-C was originally a C library and cross-compiler.
Every professional security researcher reading this just raised an eyebrow and thought to themselves, "That's a vulnerable application." In fact, your team (or a similar one at AOL) did introduce vulnerabilities[1][2]. If we assume that, as you imply, AOL really was exclusively composed of above average C developers writing high quality code, this should really just demonstrate the point for us.
I'm sure your team wrote high quality C code (or at the very least, tried very hard to write high quality C code), and as someone who likes C I also dislike the meme that C should simply not be written, no exceptions. I'd probably change it to, "Don't write C unless you need extreme performance, have great engineering resources and you'll generate unmitigated financial returns such to make any reasonable business risk assessment moot." That means virtually the same thing for anyone reading it, but it sounds a lot less sexy and memorable as far as programming aphorisms go.
I have never professionally audited a C project and not found a vulnerability. Most of my friends in this industry would likely say the same. I've never seen a serious audit of a "large" (100kloc or more) C project by a serious, reputable firm with no vulnerabilities reported. If you wrote 2 million lines of C code, if you wrote libraries from scratch, you have serious vulnerabilities. Otherwise, your engineering practices and resources exceed NASA's in formal verification and correctness.
I believe it is theoretically possible to write 100% safe C code, for some approximation of that ideal that leaves aside the ivory tower definition ("the only safe machine is the one which is off"). But I also believe this is so expensive and resource intensive (secure coding standards, audits, attempts at formal verification and provable correctness) that it's just generally not realistic in practice.
The reason I am saying all of this is not to attack you or your sense of worth as a C developer. Rather, I'd like you to consider that you simply have no idea how many vulnerabilities exist in ICQ. In fact, I found eight different security reports for ICQ, linked at the bottom. One of the more serious ones allowed arbitrary email retrieval, which I'm willing to bet was introduced because your team liked to develop in-house libraries and decided to do that for POP3.
C is a language which is both very powerful and which requires constant vigilance while coding. It is a language which makes pushing a buffer overflow to production on a Friday at 3 pm very easy. I bet your team was comparatively well-versed in C vulnerabilities for the time, but the very fact that you rolled libraries from scratch tells me you already have a high probability of errors showing up. Rolling a library/framework in-house is basically the first thing security researchers look for in an audit, because the third party ones are (usually) more secure due to their exposure. When you further consider the probability of third party software you did use becoming vulnerable in the future, changing compiler optimizations introducing new vulnerabilities and the sheer size of the SEI CERT Secure Coding Standard itself, you are left with a very dangerous risk profile.
We cannot expect C to be a safe language for general use if you need to have the rigor of one of the best research organizations in the world coupled with the top 0.01% of all C programmers. It just isn't feasible.
[1]: http://www.coresecurity.com/content/bevy-of-new-icq-vulnerab...
[2]: https://www.cvedetails.com/vulnerability-list/vendor_id-123/...
> Every professional security researcher reading this just raised an eyebrow and thought to themselves, "That's a vulnerable application."
That may be true, but substitute "C" with "Python" and wouldn't you have the same reaction? Maybe you would expect it to be slightly less vulnerable, but the key risk that I read in that statement (I am not a security professional) is the "2M" and the "written from scratch". The risk from "C" is secondary.
I think the reason why every seasoned security researcher I've met also happens to be a heavy drinker or is damaged in some way is because they're in a war of attrition. You want perfect security but it will never happen. These are ultimately mutating, register-based machines. We will probably never be certain that with a sufficient level abstraction a program can be written which will never execute an invalid instruction or be manipulated to reveal hidden information.
Where the theory hits the road is where the action happens.
Which is how we end up with this wide spectrum of acceptable tolerances to security. Holistic verification of systems is extremely costly but necessary where human lives matter. However if someone finds a weird side-channel attack in a bitmap parsing library I think we can be more forgiving.
The whole idea that C programs are insecure by default and can never be secure is where theory wants to ignore the harsh realities. We can write languages with tighter constraints on the verification of the programs they create which will lower the risk of most security exploits by huge margins... but we have to trade something away for the benefit. The immediate costs being run-time performance or qualitative things like maintainability.
What I ultimately think will make these poor security researchers feel better is liability. Having a real system and standard in place for professional practices will at least let us soak up the damage will force us to consider security and scrutinize our code.
There are all kinds of people doing security research. In my experience with some of those people that I've met it doesn't seem like the challenges they have to contend with are not unrelated to the stress caused by the work that they do.
Sort of like how someone who works with giant metal stamping presses is likely to be missing a couple of fingers if they've worked long enough with them (and in an environment where safety regulations are too relaxed).
These are also some of the wonderful people in my life and I enjoy them very much.
But you have to take care of yourself!
They are not proven to be safe, they are proven to adhere to whatever the proof system managed to prove. This shifts the possible vulnerabilities away from the actual code and onto the proof system and the assumptions that the proof system makes.
For example, take a C program that was proven to be absolutely free of buffer overflows. And therefore labelled by the proof system as 'secure'. But, unbeknownst to the proof system, the C program also interprets user-supplied input as a format string! So it's still vulnerable to format string exploits.
Correct proof systems probably add a huge margin of security compared with the current security standards, but it's not absolute.
I think C++ can only get so far, without deprecating some of C. Or at least C style code, should come with a big warning from the compiler:
"You are using an unsafe language feature. Are you absolutely sure there is no bug here, and there are no way to use safe language features instead?"
In libraries there might be cases where C-style code like manual pointer arithmetic, c arrays, manual memory management, c-style casts, or *void pointers, uninitialized variables, etc are necessary. But in user code most often there are safer replacements.
No, it doesn't succeed at this at all. Large C++ codebases routinely suffer from the same problems that C does.
For proof, search the CVE lists for browser vulnerabilities.
I don't think that's true. There's no reason why memory safety has to cost either performance or maintainability.
It feels like there's some sort of fundamental dichotomy because we didn't know how to do it in 1980, and we're still using languages from 1980, but our knowledge has advanced since then.
This is the wrong mindset and has been the guiding principle behind over a generation of poor programming discipline.
You might be able to rationalize that you don't always need extreme performance, but it's not about performance, it's about efficiency. Efficiency always matters. We live on a small planet with finite resources. Today, the majority of computing is happening with handheld portable devices with tiny batteries. Every wasted byte and CPU cycle costs power, runtime, and usability.
Even if you're coding for a desktop PC or a server application - every bit of waste costs electricity, generates heat, increases cooling load. Every bit of waste costs the user real money, even if they don't see the time that you wasted for them.
Performance is a side-effect. The only thing that matters is efficiency, and it matters all the time.
Most of today's programming problems come down to laziness and fear of "premature optimization" (which is obviously bunk). It is possible to write perfect C code that runs correctly the first time and every time. You do that by planning ahead. Formal specs are part of this.
The same consideration applies to writing, in general. I witnessed this first hand, when I worked at the computer labs at the University of Michigan back in the 1980s. When all anyone had was an electric typewriter, they planned their papers out well in advance - detailed outlines, with logically designed introductions and conclusions. This was a requirement, because once you got into actually typing, correcting a mistake could be expensive or impossible without starting over from scratch.
When we only had Wordstar available, and users came in with the inevitable questions and requests for help, you would still see well-organized papers written with a clear logical flow.
In 1985 with the advent of the Macintosh and MacWrite, all of that changed: People discovered that they could edit at will, copy and paste and move text around on the fly. And so they did - throwing planning out the window. And the quality of papers we saw at the help desk plummeted. Surveys from those years confirmed it too - paradoxically, the writing scores in the better-funded departments dropped rapidly, commensurate with their rate of introduction of Macs.
Planning ahead, making sure you know what you need to write, before you start writing any actual lines of code, prevents most problems from ever happening. It's a pretty simple discipline, but one that's lost on most people today.
If "planning ahead" were enough to prevent security vulnerabilities in practice, someone would have done it by now. Instead, we've had C for upwards of 30 years and everyone, "bad programmers", "good programmers" and "10xers" alike, keeps introducing the same vulnerabilities.
It's a nice theory, but we've been trying for over 30 years to make it work, and it keeps failing again and again. I think it's time to admit that C as a secure language (modulo the extremely expensive formal coding practices used in avionics and whatnot) has failed.
Security is not a property of the tool, it is a property of the tool's wielder.
How many memory safety bugs do you see in Java applications compared to C applications?
Writing in Java is an effective way to reduce remote code execution vulnerabilities over writing in C. Statistics have shown this over and over again.
Remote code execution vulnerabilities in C are less a feature of the language syntax and mostly due to the abysmally poor C library, which is regrettably part of "C The Language" - but no one forces you to use it - there are better libraries. Also, it's an artifact of implementation - just as Java's supposed immunity is an artifact of implementation, not an inherent feature of the language syntax.
Case in point http://www.trendmicro.com/vinfo/us/threat-encyclopedia/vulne...
As for RCEs in C code - these would be easily avoided if the implementation used two separate stacks - one for function parameters, and one for return addresses. If user code can only overwrite passive data, stack smashing attacks would be totally ineffective. Again - there's nothing in the C language specification that dictates or defines this particular vulnerability. It is solely an internal implementation decision.
No. Increasingly, RCEs are due to vtable confusion resulting from use after free.
Meanwhile, new programming languages abound, presumably to make programmers' lives easier. But these developments often completely lose sight of the users. Python is easier to write, but a program written in python uses 1000x the CPU and memory resources of a program written in C. It does the job, but slowly, and consuming so much that the computer is unable to do anything else productive. Meanwhile, it's questionable whether programmers' lives have been improved either, because they spend too much time chasing the next new tool instead of growing expertise in just one.
As for marginal efficiency gains - that is only a symptom of the latter developer problem. E.g., LMDB is pure C, less than 8KLOCs, and is orders of magnitude faster than every other embedded database engine out there. It is also 100% crash-proof and offers a plethora of features that other DB engines lack. All while being able to execute entirely within a CPU's L1 cache. The difference is more than "marginal" - and that difference comes from years of practice with C. You will never gain that kind of experience by being a dilettante and messing with the flavor of the week.
Just out of interest: How many C projects have you audited? And have you ever looked for a relationship between the number of vulnerabilities and the percentage code coverage from automated tests?
But how is that a useful or replicable result, that will work in the real world for other projects? You're basically saying that everyone just needs to be a top 0.001% C programmer. This is the "abstinence-only sex education" of programming advice. It's super simple to not get STDs or have out-of-wedlock pregnancies - just don't have sex with anyone but your spouse, and ensure they do the same. It's 100% guaranteed to work, and if you don't follow that advice you're stupid and deserve whatever bad things happen to you.
For the record, I don't think your stupid if you have a child accidentally.
However, once you do accidentally have a child, the results are often not pretty, and I will feel bad watching as you deal with the consequences of your poorly thought out decision.
Regarding programming, I can't see the relationship you make at all.
If you write a good program, well good job. If you write a bad program, you are probably going to write a better one next time. (Ever look back at your old code?)
It takes practice to get good, and the more we work at it, the better we'll get. Just keep trying.
It's as if I say: "Step 1 of writing good code: Don't make mistakes. If you make mistakes, you deserve what's coming to you!" Some mistakes are expected, and essentially inevitable; they're the default, not the exception. Blaming the programmer might seem like the right thing to do, but there's more practical benefit to improving the tools and teaching programmers how to handle errors, so that we can decrease the negative impact that bugs will have.
I have two words for you: MAINTENANCE PROGRAMER.
Now that I think of it, there's another parallel with unwanted pregnancies. There are men that stick around, and there are boys who run and let others clean up after their mess.
What is this, 1953?
There are a great number of couples choosing to have children in stable, committed - yet unmarried - relationships.
The comment struck me as incredibly anachronistic, bordering on offensively so.
please don't.
Way to hijack a useless subthread and do something useful with it. Great meta-point. Honestly. This feature alone makes the comments on Reddit bearable for consumption.
Should be 1967.
</s>
I imagine that this poster and his team took a similar approach with C: avoiding the dangerous features, or just finding ways to factor them out.
A more constructive/charitable phrasing of the parent's point might be: you should only be allowed to code in C if you're already a top 0.001% programmer. Then, obviously, all C code will be good—not that there'll be much of it.
That's actually a more interesting point than it seems. Programmers generally tend to reject the idea of guilds and "professional" licensing—mostly because, AFAIK, most programs aren't critical and most crashes don't matter. But what if, instead of licensing all programming, we just licensed only low-level programming? What if, as a rule, everyone who was going to be hired to write code in a language with pointers was union-mandated to be certified as knowing how to safely fiddle with pointers? That could probably split on interesting dimensions, like a bunch of regular coders working in e.g. a Rust codebase, with a certified Low Level Programmer needing to "sign off on" any commit that contained unsafe{} code, where the actual failure of the unsafe{} code would get the Low Level Programmer disbarred.
I'm a very experienced C programmer and one day my boss came to me and said that the sales guys had already sold a non-existing client side module to a house hold name appliance manufacturer. The deal was inked and it had to be ready in only 3 months. Even worse, it had to run in the Unix kernel of the appliance and therefore be rock solid so as to not take the whole appliance down. It also had to be ultra high performance because the appliance was ultra high performance and very expensive. Now the really bad news: I had a team made up of 3 more experienced C developers (including myself) and 3 very un-experienced C developers. We also estimated that in order to code all the functionality it would take at least 4 months. So we added on another 4 less experienced C developers (the office didn't have a lot of C developers). The project was completed in time, a success, and almost no bugs were found, and yet many developers without much C experience worked on the project. How?
(a) No dynamic memory allocation was used at run-time and therefore we never had to worry about memory leaks.
(b) Very, very few pointers were used. Instead mainly arrays. And not just C developers understand array syntax, e.g. myarray[i].member = 1 :-) Therefore we never had to worry about invalid pointers.
(c) Source changes could only be committed together with automated tests resulting in 100% code coverage and after peer review. This meant that most bugs were discovered immediately after being created but before being checked in to the source repository. We achieved 100% code coverage with an approx. 1:1 ratio of production C source code to test C source code.
(d) All code was written using pair programming.
(e) Automated performance tests were ran on each code commit to immediately spot any new code causing a performance problem.
(f) All code was written from scratch to C89 standards for embedding in the kernel. About a dozen interface functions were identified which allowed us to develop the code in isolation from the appliance, and not have to learn the appliance etc.
(g) There was a debug version of the code littered with assert()s and very verbose logging. Therefore, we never needed to use a traditional debugger. The verbose logging allowed us to debug the multi-core code. Regular developers were not allowed to use mutexes etc themselves in source code. Instead, generic higher level constructs were used to achieve multi-core. My impression is that debugging via sophisticated log files is faster than using a debugger.
(h) We automated the process of the Makefile so that developers could create new source files and/or libraries on-the-fly without having to understand make voodoo. C header files were also auto generated to increase developer productivity.
(i) Naming conventions for folders, files, and C source code were enforced programmatically and by reviews. In this way it was easier for developers to name things and comprehend the code of others.
In essence, we created a kind of 'dumbed down' version of C which was approaching being as easy to code in as a high level scripting language. Developers found themselves empowered to write a lot of code very quickly because they could rely on the automated testing to ensure that they hadn't inadvertently broken something, even other parts of the code that they had little idea about. This only worked well because there was a clear architecture and code skeleton. The rest was like 'painting by numbers' for the majority of developers who had little experience with C.
The same team went on to develop more C software using the same technique and with great success.
Unfortunely your case is the exception that confirms the rule.
I never saw a company using C like that.
On my case the teams used to be composed from circa 30 developers, scattered around multiple consulting companies with high attrition.
Another tidbit: The developer pairs were responsible for writing both the production code and associated automated tests for the production code. We had no 'QA' / test developers. All code would be reviewed by a third developer prior to check-in. However, at one stage we tried developing all the test code in a high level scripting language with the idea that it would be faster to write the tests and need less lines of test source code. However, because we did several projects like this, we noticed that there was no advantage to writing tests in a scripting language. The ratio of production C source lines to test source lines was about the same whether the test source code was written in C or a scripting language. Further, there was an advantage to writing the tests in C because they ran much faster. We had some tens of thousands of tests and all of them could compile and run in under two minutes total, and that includes compiling three version of the sources and running the tests on each one; production, debug, and code coverage builds. Because the entire test cycle was so fast then developers could do 'merciless refactoring'.
Which is to say, your story matches my hypothetical pretty well. You effectively created the same structure as Rust has, where there are safe and unsafe sublanguages, and the "master" developers thoroughly audited any code implemented in the unsafe sublanguage.
Which makes my point: there exist a small set of people qualified to write in "full C", and a much-larger set of people who aren't; and the only way—if you're a person who isn't qualified—to write C that doesn't fall down, is with the guidance and leadership of a person who is.
Though I think maybe your thesis statement agrees with that, so maybe I'm not arguing with you. Not all the developers on a given project need to be from the qualified set, no. But at least some of the developers need to be from the qualified set, and the other developers need to consider the qualified ones' guidance—especially on what C features to use or avoid—to be law for the project. As long as they stay within those guidelines, they're really working within a DSL the "master" developers constructed, not in C. And nobody ever said you can't make an idiot-proof DSL on top of C; just that, if you need everything C offers, your the only good option is to remove the idiots. :)
I also forgot to mention about performance. Whether you code * foo or the easier to comprehend and safer(?) foo[i] then the compiler still does an awesome job optimizing. However, it's much easier to assert() if variable i is in range rather than *foo. And it's also easier to read a verbose log variable i (usually a 'human-readable' integer) than to read a verbose log pointer (a 'non-human-readable' long hex number). At the time we wrote a lot of network daemons for the cloud and performance tested against other freely available code to ensure that we weren't just re-inventing the wheel. NGINX seemed to be the next fastest, but our 'dumbed down' version of C ran about twice as fast as NGINX according to various benchmarks at the time. Looking back, I think that's because we performance tested our code from the first lines of production code, and there was no chance for any type of even puppy fat to creep in. Think: 'Look after the cents and the dollars look after themselves' :-) Plus NGINX also has to handle the generic case, whereas we only needed to handle a specific subset of HTTP / HTTPS.
I wanted to interpret it as you shouldn't write code if you can avoid it.
I'm certain that's not what he meant, but it sounds better.
Good developer is able to write high quality code in C...
the sufficently smart developer is as much elusive as the sufficently smart compiler- Was your code unit-tested?
- Did you use the 'strn' functions? (strncpy, instead of strcpy for example)
- Did your code run under Valgrind? (was there valgrind at that time? I'm not sure)
- Was it tested for memory leaks?
- Was there input fuzzying tests?
I agree that there are ways of writing great C code (like nginx, the linux and BSD kernels, etc)
But from a certain point on you're just wasting developer time when you would have a better solution in Java/Python/Ruby, etc with much less developer time and much less chance of bugs and security issues.
There was not one unit test. And there probably still isn't.
Why wouldn't you unit test in C? I do and I have found lot's of bugs as a result and maintaining code is much easier.
1. write a new one file script which parses function declarations from a source code file
2. have the script find all function declarations which start with "utest_", accept no parameters, and return "int" or "bool".
3. have the script write out a new C file which calls all of matching unit test functions and asserts their return value is equal to 1
4. have your Makefile build and run a unit test binary using code generated by your script for each source code file containing greater than 0 unit tests.
Writing a script to parse function declarations from your C source code is useful because you can extend the script later on to generate documentation, code statistics, or conduct project-specific static analysis.
https://randomascii.wordpress.com/2013/04/03/stop-using-strn...
But if you're using C++, use C++ strings, you don't need strcpy
There's strlcpy mentioned there, which seems a better C solution
I interpreted that differently than you. I didn't think take it to use another language. But to keep c code as concise as possible and to not reinvent the wheel when a good library is available.
I have seen the quite a few of those and I imagine you wouldn't enjoy to audite their code.
A good programmer will make COBOL look clean and understandable. But on average, I could way C code requires a lot more discipline, knowledge, and experience to produce quality code.
* You've never deployed one.
* You've never had to maintain one.
* You wrote the last line of code and handed it over to the maintenance team, who then never gave you feedback.
* You have no customers.
* You have done a never-before-seen formal analysis of a large scale project down to the individual source line level, and made no mistakes with the formal specification.
The author is correct: use C only if you must. There are still an enormous number of reasons why you must use C, but it should be a constraint imposed on you, and not a language you actively want to start a project in.
It's the same as with, say, goto or #ifdef's - one has to consider the alternatives and make an informed choice. Most of the time, the alternatives are preferable, but sometimes there either aren't any alternatives, or they are even worse.
Which is not to say that C can't be a whole lot of fun. I love it dearly. But I, too, tend to avoid it when I can. It's just too easy to shoot yourself in the foot.
(of course, one employer appropriately used C as a portable assembler, and respected it as such, for the purposes of developing a cross compiler and tool emulations for an older minicomputer environment; the next was doing (batch) "business applications" and foolishly spent too little on hardware and too much wasted effort on application development)
Pascal (Modula) did much the same, SAFELY, but the money wasn't put into compiler optimization :-(
Interesting that the GNU compiler collection now includes a very efficient Ada compiler.
GNAT has existed for ... at least ten years, I think. The DOD apparently paid for the development so there would be at least one open source/free software implementation. (According to Wikipedia, development started about twenty years ago, and it was merged into GCC in 2001.)
Parent seems to believe in nginx. But as a guy who actually tried to do a few of those things with nginx I can confirm that nginx's code base is pretty awful and working around all the usual pitfalls of C in nginx is a huge PITA and enormous amount of work.
(same username on github)
I like tweaking younger engineers with C eccentricities, but any commercial code tends towards being as vanilla as possible with lots of error checking and docs to be as clear as possible.
1. Almost devoid of any documentation. 2. No ADTs - all struct fields are directly accessed, even though there are sometimes very complicated invariants that must be preserved (and these invariants are not documented anywhere, of course). 3. Ultra-short variable names that really describe nothing but their types (it's like Hungarian notation with only the notation part). 4. These mystery variables are all defined in the top of the function of course, making them hard to trace. Does anyone even compile Nginx with C89? 5. Configuration parsing creates a lengthy boilerplate that's really hard to track. 6. Asynchronous code is very hard to track, with a myriad of phase-specific meaning for return codes, callbacks and other stuff that makes reading Boost.ASIO code look like a walk in the park.
I think it really shows the sad state of C programming. Die-hard C programmers always complain C++ is hard to read because of its template and OOP abstractions, and it sometimes really is a mess to read, but C is even harder since large programs always re-implement their own half-assed version of OOP and template-macros inside.
You meant "one of the worst"?
Man, this is just lame. C is a really nice, useful language, especially if you learn the internals. It requires discipline and care to use effectively, but it can be done right. The problem is that a lot of C programmers only know C and can't apply higher level concepts or abstractions to it because they've never learned anything else. This makes their code harder to maintain and easier to break.
I think Swift would be a better example here. Or objective C with ARC.
If you mean it's easier to write correct programs in C than correct programs in Rust, I'm afraid you're badly mistaken.
Actually, it's the very process of “fighting the compiler” (as you call it, though I prefer to view it as collaboration between a human who knows what he wants and a program whose logic codifies what is possible) that makes it easier to write correct programs in Rust/Haskell/ML/$FAVORITE_TYPEFUL_LANGUAGE.
The notion that a type system is "just" logic and so should be trivially followed by anyone "smart enough" or "careful enough" is laughable - there are many different type systems (sometimes with different details depending on the options or pragmas you feed to your tooling), and it can take time to get a sense of where the boundaries are.
The notion that you're not likely to find a type system helpful if you have trouble with it initially is harmful.
Please don't do this.
I would've never said "just". Formal logic is big f...reaking deal.
> so should be trivially followed by anyone "smart enough" or "careful enough" is laughable
I never said anything about being "smart" or "careful". I only said that hacker types, for whom the ability to subvert anything anytime is a fundamental tenet, tend to have a hard time sticking to the rules of a formal game, whether it is a type system or not. Just look at the reasons Lispers give for liking Lisp.
"Just" was in no way meant to diminish significance; it was meant to draw attention to your improper limiting of scope. Someone "fighting the compiler" is not fighting "executable embodiments of formal logic", but "executable embodiments of formal logic and a pile of engineering decisions". Often, enough of those decisions are made well that the tool can be phenomenally useful; some are made less well, and some simply involve tradeoffs. It doesn't have the... inevitability that your wording carried. Programmers coming from a context where enough of those decisions have been made differently are likely to get tripped up for a while in the transition, without harboring any objections to logic.
> I never said anything about being "smart" or "careful". I only said that hacker types, for whom the ability to subvert anything anytime is a fundamental tenet, tend to have a hard time sticking to the rules of a formal game, whether it is a type system or not.
True, but you left it up to us to figure out what attributes of "hacker types" were relevant. As there are many different uses of the word "hacker", it was quite unclear that you meant to refer specifically to inability to stick to rules.
While people of that description may well persist in "fighting the compiler" for longer, that is not most of what I've observed when I've observed people struggling with a type system, and I reiterate my assertion that those new to a particular tool and used to doing things another way also frequently struggle for a period.
I liken programming w/Rust to programming with C or C++ with warnings and a static analyzer barring object files from being created. If I were new to C/C++, I would probably find it baffling how anyone ever gets it to work.
Simpler in what sense? Not this one: https://news.ycombinator.com/item?id=10671800
It's probably easier to write CS101 prime-number or fibonacci programs in C.
Any nontrivial bit of software? You're going to have to learn about undefined behavior and deal with it. Rust forces you to learn this before getting started.
Plus Rust has a lot of higher-level (still zero cost) abstractions like Python which make a lot of patterns really easy to program. (And Python is _the_ teaching language these days)
See also: https://www.reddit.com/r/rust/comments/3mtwev/using_rust_wit...
In practice, I have never seen any programs where this is the best approach. Coming up with designs that don't need unsafe code is very useful, because the safe subset of Rust is vastly easier to analyze and understand, both for humans and computers.
Well, the C compiler sometimes rejects correct programs too.
This comes primarily from the fact the C standard itself is hard to interpret right. Standard ML has less issues in this department: since the Definition is clearer and less ambiguous, implementations disagree less on what counts as correct code and what its meaning is. So having only one major language implementation isn't the only way to prevent compatibility issues.
[1]: http://www.open-std.org/jtc1/sc22/wg14/www/docs/n1637.pdf
C is riddled with opportunities to do _just that_ with even the slightest mistake.
Yes, I can make a logic error in any of the above languages that performs some operation a user shouldn't actually be allowed to perform. But I can't accidentally let the user read or write arbitrary memory locations.
With the exception of unsafe blocks or similar mechanisms; being limited to these explicit scopes, most programmer's natural laziness leads to avoiding unsafe behavior whenever possible. It also greatly limits the scope of potential bugs, as well as testing and code review required.
Many of the languages you mention contain an 'eval' method, and most use environment variables in less than transparent ways. I wouldn't be so sure of this claim.
http://blog.codeclimate.com/blog/2013/01/10/rails-remote-cod...
- Using fixed width integers in for loops seems like a fabulous way to reduce the portability of code
- the statements in "C allows static initialization of stack-allocated arrays" are _not_ equivalent, one is a bitwise zeroing while the other is arithmetic initialization. On some machines changing these statements blindly will cause a different bit pattern to end up in memory (because there are no requirements on the machine's representations for e.g. integers or NULL or ..). There are sound reasons why the bitwise approach could be preferred, for example, because a project has debug wrappers for memset that clearly demarcate uninitialized data
- the statements in "C99 allows variable length array initializsers" aren't even slightly equivalent. His suggestion uses automatic storage (and subsequently a stack overflow triggered by user input -- aka. a security bug)
- "There is no performance penalty for getting zero'd memory" this is bullshit. calloc() might optimize for the case where it is allocating from a page the OS has just supplied but I doubt any implementation ever bothered to do this, since it relies on the host OS to always zero new pages
- "If a function accepts arbitrary input data and a length to process, don't restrict the type of the parameter." the former version is in every way more self-documenting and consistent with the apparent function of the procedure than his use of void. It also runs counter to a rule from slightly more conservative times: avoid void at all costs, since it automatically silences all casting warnings.
Modern C provides a bunch of new things that make typing safer but none of those techniques are mentioned here. For example word-sized structs combined with struct literals can eliminate whole classes of historical bugs.
On the fixed width integers thing, the size of 'int', 'long' and 'long long' are designed to vary according to the machine in use. On some fancy Intel box perhaps there is no cost to using a 64bit type all the time, but on a microcontroller you've just caused the compiler to inject a software arithmetic implementation into your binary (and your code is running 100x slower too). He doesn't even mention types like intfast_t designed for this case, despite explicitly indicating "don't write C if you have to", which in 2016 pretty commonly means you're targeting such a device
Can you explain this? I'm not seeing it.
You should generally know the range of numbers you're seeking to store in an integer variable, if you don't then you're asking for trouble.
You can still range check the natural integer types, just use INT_MAX and suchlike
Either you need the range, in which case you're stuck with a large type, or you don't and can use a smaller one.
I'm not sure where 'oversized' comes into it.
I ran into a bug where I was running some Arduino code (16-bit) on an mBed (32-bit). The use of `int` rather than `uint16_t` caused a timeout to become an infinite loop.
I'm pretty sure it would be impossible to have portability issues due to using fixed-width types - I mean their entire point is to ensure portability.
Their point is against portability as they may not exist at all.
The good part being that you're warned at compile time, not when your 2^16 elements loop becomes a 2^32 elements loop.
Or if you do not make assumptions of the sizes of int, something that only bad programmers do.
No you'll have the exact problem I cited in that case.
> Or if you do not make assumptions of the sizes of int, something that only bad programmers do.
Requesting a specific data size is the exact opposite of making assumptions.
I can see how the type changing to a smaller one might cause problems, but I don't see how IshKebab's example could happen without exploiting implementation specific overflow behavior.
INT_MAX (or INT_LEASTN_MAX for annathebannana's suggestion) doesn't require exploiting overflow behaviour, but going from 2^15 iterations to 2^31 or 2^61 iterations may be problematic.
And using INT_MAX or similar for what sounds like a timing loop is a whole other can of bad practice. Then the problem isn't that you used the wrong type, it's that you used the wrong value.
If the loop does significant work and was calibrated for an expectation of 65k iterations, stepping to 2 billion (let alone a few quintillion) is for all intents and purpose endless.
> And using INT_MAX or similar for what sounds like a timing loop is a whole other can of bad practice. Then the problem isn't that you used the wrong type, it's that you used the wrong value.
No objection there, doing that is making invalid assumptions, my point is that moving to exact-size integral does fix it.
And if it's not, it's not, the point being that endless in this case would be meaningless outside it's literal meaning unless we know more about the specific case.
> No objection there, doing that is making invalid assumptions, my point is that moving to exact-size integral does fix it.
No, using an exact value fixes it. Any unsigned integer type is just fine for any integer value from 0 to 65535. If you change the type to a larger integer type without changing the supposed iteration count, the code would not have this problem, and if you changed the value to something higher than 65535 without adjusting the size of the type, you would have a different problem. Thus, the problem described here does not pertain to the type of the variable used.
Also, it's less typing. :)
in 99% of the cases, except if you are working with network protocols, [u]intN_t is a bad choice
It will, that's guaranteed by the standard. If it's not a standards compliant compiler/runtime/platform, you've got bigger problems.
>> If you are working with arrays, always use size_t if not, use [u]int_leastN_t
Sure, on array indices, size_t is appropriate.
I can't see an advantage in using 'least' as standards dictate that 8/16/32/64 bit types are available.
>> in 99% of the cases, except if you are working with network protocols, [u]intN_t is a bad choice
No, it's the best choice in most cases because it makes the code a little more explicit and easier to understand, and it makes developers think about the range of the data you're using.
So, if a system is 1's complement, you won't get the intN_t types at all. If the system doesn't support a particular bit width, you won't get the types for that width.
As someone who actually programmed a 1's complement system with end-around carry, I'm going to call you out on this.
What system still exists that can run C99 and is one's complement?
I have no idea if C99 runs on any 1's complement systems. I don't know what end-around carry is. I don't care what it is. I don't care about any of this 1's complement nonsense. But the standard says it exists, and must be taken into account, and, therefore, by the rules of the game, I am obliged to assume these things.
Regarding calloc (which indeed might be slower than malloc contrary to TFA) - a variable that should have been initialized to something but instead keeps zero bits written by calloc is not necessarily a great situation. If you use malloc, at least Valgrind will show you where you use the uninitialized variable. With calloc it won't, though it might still be a bug - a consistently behaving bug, but a bug nonetheless. Someone preferring consistently behaving bugs to inconsistently behaving bugs in production code might use a function my_malloc (not my_calloc...) which in a release build calls calloc, but in debug builds it calls malloc (and perhaps detects whether it runs under Valgrind and if not, initializes the buffer not with zeros, but with deterministic pseudo-random garbage.)
In starvation, malloc+memset "usually" results in a malloc() failure, instead of the process being silently killed by the OOM killer.
Calloc has another benefit: if you calculate size by n * sizeof(kuku) then it will detect integer overflow (this can be very important for security sensitive stuff that deals with untrusted input), if you did the computation before that and passed calloc( size, 1) then this multiplication and check is another few cycles wasted.
I wonder if they could have done a standard library allocation function like calloc that does the multiplication with overflow detection, does not zero out memory and that returns NULL on zero arguments - but that would be too much to ask.
[1] http://www.openbsd.org/cgi-bin/man.cgi/OpenBSD-current/man9/...
[2] http://www.openbsd.org/cgi-bin/man.cgi/OpenBSD-current/man3/...
uint32_t array[10] = {0};
Does not initialise every element to 0 in the way it would seem to. To see the difference contrast the difference you get when you a) remove the initialiser and b) replace the initialiser with {1}.C++ lets you do "uint32_t array[10]={}", which is something C should allow as well, really. But it doesn't.
Yes it does, since atleast C99. As for the others:
a) Removing the initializer will leave the values undefined.
b) Using an initializer of { 1 } will initialize the first element to 1 and the rest to 0.
{} initialises all elements to 0.
{0} initialises the first element to 0 and the rest to 0.
The latter form just introduces confusion.
The problem I've had (and you probably are referring to) is embedded systems that may not properly zero bss. I've also had problems in embedded systems trying to place "array[10] = {0};" into a specific non-bss section (that was really annoying).
In my experience, TFA is good practices generally, but will have problems in corner cases that you run into in deeply embedded systems.
I think the real issue here is the inconsistency between C standards on details like this. If you can always assume that your code is built with c99 or later, then use all of its features, but in many cases that's not a realistic assumption.
“If there are fewer initializers in a list than there are members of an aggregate, the remainder of the aggregate shall be initialized implicitly the same as objects that have static storage duration.”
You've written a list of coding practices which assure job security for folks like me, who have to undo these gross portability problems, bugs, and security vulnerabilities these result in.
for (int i = 0; i < 100000; ++i) // most optimal int size used
Is bullshit, always.These are wonderful. Concise nuggets of knowledge that you can keep in your head, little warnings of what not to do.
Side note, it's really fun reading this article about modern C given that I'm working on an experimental programming language meant to replace C [1]. Just today I got the Guess Number Game example working, without a dependency on libc [2] (caveat: Linux x86_64 only so far). So, I'm having fun trying to think about how C could be better, reading implementations of various libc functions, and writing my own "standard library".
For example, in my language, there are no `int`, `unsigned char`, `short` types, there are only `i32`, `u8`, `i16`, etc (I shamelessly stole this idea from Rust). This mirrors the article's suggestion to only use int32_t, uint8_t, int16_t, etc. But, since it's a new programming language, we don't have the other types sitting there as a red herring.
Those are very old ideas.
While these are not new ideas - Why not use s32 rather than i32 and so on. Makes more sense if you use something like u32 too.
I think it's important when making another language to really question everything you do in it. Why use this operator, or why use this mnemonic for built in types?
i denotes an integral, as opposed to the also-signed float types?
> You should always use calloc. There is no performance penalty for getting zero'd memory.
But isn't there? This may be true the first time the program requests a new page from the OS, but when your program starts reusing memory there's going to be a performance hit.
And by zeroing pages, you prevent lazy memory allocation policies which do not allocate the pages until they're read or written for the first time. This can have a significant impact on memory usage, data locality etc.
So I'm quite skeptical that there is no difference between calloc and malloc. Is there more evidences of this somewhere?
But the OS/MMU doesn't distinguish between a regular write, and a write of zero. Thus, if you manually zero every page you get (And thus write zeros to the read-only zero-page), it'll page-fault back to the OS and the OS will have to allocate a new page so that the write succeeds - Even though if you didn't do the zeroing of memory you would have gotten the same effect of having a bunch of zeros in memory, but without having to allocate any new pages for your process.
Since malloc/calloc are generally used for smaller allocations, the chances you can actually avoid allocating some pages you ask for is pretty slim since a bunch of objects get placed into the same page (And thus writing to any of them will trigger a new page being allocated). There's also no guarantee there isn't headers for malloc to make use of, or similar surrounding your piece of memory, which makes the point moot - Just using malloc triggers writes to at least the first page. So while calloc/malloc are kinda compatible with lazy-allocation, you really shouldn't rely on it being a thing, and it probably won't matter.
It's worth understanding, but the chances it actually comes into play aren't huge. If your program does lots of small malloc's and free's, then it basically won't matter because you won't be asking the kernel for more memory, just reusing what you already have.
If you care about taking advantage of lazy-allocation for one reason or another, the bottom line is probably that you shouldn't be using malloc and calloc for that then. Just use mmap directly and you'll have a better time - more control, you have a nice page-aligned address to start with, and you can be sure the memory is untouched. malloc and calloc are good for general allocations, but using mmap takes out the guesswork when you have something big and specific you need to allocate.
E.g. imagine /etc/passwd is read and later on the page it occupied is put back in the free pool. Another process comes along and asks for memory. It gets that page and can now read /etc/passwd .
A few things though:
- Last time I ran this, in the memory dump, every second character was the letter "f". Like, in the dump, instead of saying "Words" it would say "Wfofrfdfsf". It changes each time, probably to do with my char onechar variable. (Actually now that I think about it, it should probably only be onechar[1] instead of onechar[2]. It was called onechar because I was trying to printf 1 char at a time)
- Before I changed "p" to &memcpy, it was "p = &p", and it would only read my environ before segfaulting.
- This program, on linux, dies almost immediately. It only spits out a few bytes.
So, no, you aren't printing out bits of files. You're printing out functions that your program has the ability to call.
He's starting somewhere in /usr/lib/system/libsystem_platform.dylib (where memcpy lives) and printing the rest of the executable code his program has mapped in.
It doesn't surprise me that some of this contains public keys and certificates. They are probably used by some OSX networking library.
On linux you can dump a processes address space mapping from file "/proc/<pid>/maps". Or to see an example just run: cat /proc/self/maps That cat command will dump the address space mapping for that instance of the cat command. If you run it multiple times it will show different address ranges because of a security feature (ASLR, address space layout randomization).
Also you shouldn't use fprintf or printf to print 1 char. printf will try to look for format characters in the "string" you are passing it. It would be better to use fputc(*p, stdout), then you don't even need to call memcpy() or the "onechar" array.
If you set "p" to a block of allocated memory, then your program should segfault shortly after reading past the end of the memory. How many bytes you can read past the end of the allocated block of memory depends on how the memory allocator manages pools of memory and it's own bookkeeping information is stored.
I can't get it to print 3g of stuff on OSX. It prints about 300m , which makes a lot more sense.
If you're sitting in a loop calling malloc(), filling it with data, processing it, and calling free(), changing malloc() to calloc() is just wasting cycles.
I guess one can rely on such behavior when writing a platform-specific libc. But in portable code? Yuck.
2. One could imagine this being configurable for embedded or HPC systems which don't care about security, though I'm not aware of people actually doing that.
3. The memory could equally well be overwritten with 0xdeadbeef and you have no control over that.
4. I once heard that Windows may sometimes fault-in nonzero pages if it runs out of pre-zeroed ones, but I couldn't locate this in MS docs so I'm not quite sure.
Thanks.
Heap allocation can easily contain data that have been previously freed in the same process and hence calloc has to perform some extra work to wipe them.
IIRC FreeBSD will zero freed pages in the background when it has idle time, and calloc can draw from that pool of zeroed pages "for free".
Of course it does.
> That only works when the pages are actually freed. Since actually asking the kernel for more memory is expensive, most malloc/calloc implementations will hold-on to some pages since it assumes you're going to allocate more. Since it reuses this memory without ever giving it back to the kernel, the kernel won't zero it for us.
It's what the standard system malloc does, possibly in cooperation with the kernel but not necessarily so (why would it need the kernel involved after all?)
Just as OpenBSD's malloc framework can do way more than trivially request memory from the kernel: http://www.openbsd.org/cgi-bin/man.cgi/OpenBSD-current/man5/.... In fact, since OpenBSD 5.6 it'll junk small allocations on freeing (fill them with 0xdf) by default. And as usual that uncovered broken code…
> When we started junking memory by default after free, the postgres ruby gem stopped working.
To be clear though, the parent commenter (and I) are talking about the pages that malloc keeps around in the processes memory to be reused (And thus are never released back to the OS). Are you saying FreeBSD/OpenBSD/others have a system to tell the kernel when a user process has pages it plans to zero, and then a system for the kernel to notify the process when/if it does? That would be pretty interesting to see, but I've never heard of that being a thing.
I was talking about https://www.freebsd.org/doc/en/articles/vm-design/prefault-o... but after actually re-checking the docs rather than going from bitrotted memory pre-zeroing only happens for the initial zero-fill fault, so the original calloc is "free" (if there are zeroed pages in the queue) but freeing then re-calloc'ing will need to zero the chunk (unless the allocator decides to go see if the kernel has pre-zeroed space instead I guess). My bad.
> To be clear though, the parent commenter (and I) are talking about the pages that malloc keeps around in the processes memory to be reused (And thus are never released back to the OS).
Yeah so was I, but I was misremembering stuff.
> Are you saying FreeBSD/OpenBSD/others have a system to tell the kernel when a user process has pages it plans to zero, and then a system for the kernel to notify the process when/if it does?
Turns out no.
I'd very much recommend that advice; use it all the time, remove it only when it proves to be a problem, and then, as always when optimizing, bear in mind that you've chosen to juggle lit sticks of dynamite and prepare accordingly.
[1]: I hypothesize that this is because what "really" comes back from a malloc is basically too impossibly complicated to model by a human. Even calling it "random" isn't correct; it's not random, but what it is is very hard to characterize.
I think "uninitialized" is the word you are looking for :)
Uninitialized isn't "a mess that you really ought to clean up before using", it's radioactive waste.
If you want to tell me that you really do think of "uninitialized" that way every time you use it, I have no grounds to disagree with you personally... but based on the code in the wild, it does not seem that most C programmers do have a correct understanding of the issue.
(People who develop the proper level of concern about handling radioactive waste tend to find ways to stop programming in languages that don't default to handing back nuclear waste on every memory allocation. I suspect some form of evaporative cooling may be a problem here.)
But how does replacing that radioactive waste (neat analogy) with zeros improve correctness here?
It doesn't improve correctness it limits liability of incorrect code.
Coming to C with a mental model of memory from another language is a recipe for security problems.
The next person may have equally likely (or maybe more likely?) simply forgotten to initialize some particular field and now you've got a problem if the field is supposed to read 0 most of the time and maybe sometimes 1.
Yes, calloc() is a few ms slower than malloc(), but does it really matter in 99% of the use cases vs being on the safe side of clean data?
Why would you even be using C if it didn't?
- Exposing OS APIs to other languages
- Integration with an SDK that only provides C APIs
- Improving an algorithm until it runs at the desired speed for the customer (no need to shave every last ms)
- Targeting embedded processors
Although, I personally would rather use either C++ or a safer systems programming language.
Please list said safer systems programming languages.
- Spark
- Rust
- D
- Swift (it will eventually cover all use cases with 3.0)
- FreePascal
- Delphi
- ATS
If I would be in an hipster mood and revive old faded ones:
- Dylan
- Modula-2
- Modula-2+
- Modula-3
- Oberon
- Oberon-2
- Active Oberon
- Component Pascal
- Algol
- PL/I
- CPL
- Mesa
- Cedar
- Turbo Pascal
- Apple Pascal
- Quick Pascal
There are a few other others that I certainly forgot.
> Swift (it will eventually cover all use cases with 3.0)
There's nothing in 3.0 that will make Swift any better at safe systems programming than it is today (last I checked, anyway), and Chris Lattner is on the record as saying that anything akin to Rust-style static analysis is explicitly out-of-scope for 3.0."Swift is a successor to both the C and Objective-C languages"
https://developer.apple.com/swift/
And I read somewhere that there are still some fine tuning in regards to covering the C use cases, which should be part of Swift 3.0.
I was really hoping that the list would be longer than that. :(
> Ada
> Spark
> Rust
Ada (Spark is effectively Ada) and Rust
> D
Nope. Garbage collected doesn't work for systems programming.
> Swift
Not sure I buy this. When I see someone chewing on an OS in Swift I might believe it.
> FreePascal
> Delphi
Dunno about FreePascal, but my anecdotal experiences with Delphi programs have left me wishing that their authors had written them in C. It is, of course, possible that the programming teams were bad, but it sure seems like Delphi isn't any better than C in that respect.
I wonder how Xerox PARC managed to write Interlisp-D, Mesa/Cedar.
I wonder how Symbolics managed to write Lisp Machines OS.
I wonder how ETHZ managed to write Oberon, Oberon System 3,EthOS, AoS 2.
I wonder how Olivetti and Compaq managed to produce SPIN OS.
I wonder how Microsoft was able to produce Singularity and Midori OSes.
Actually I don't need to wonder, plain technical achievements that fail to convice religion against GC systems.
All of those operating systems got their ass kicked by a 1970's reject of an operating system rewritten by a college student who didn't know what he was doing.
And, in fact, Symbolics got its ass kicked so badly in terms of performance that they had to port everything to a non-garbage collected OS to even begin to compete.
If you want to try to refute me, choose Erlang. It actually still exists and is in use. It also garbage collects per process and runs down on bare metal levels so its ecosystem could be called an OS.
However, Erlang succeeds in spite of its GC, not because of it. Per process GC combined with "shoot and restart" mentality means that when the GC goes pathological the whole process gets shot and the memory gets released all at once and reallocated.
Of course, if I'm being truly cynical, I would ask why we should do GC at all in Erlang and not just shoot the process regularly.
After all, RC is usually the first chapter of any CS book about GC algorithms.
I dream of the day when non-believers will be forced to use Swift, .NET Native, C++/CX, Go on their beloved OSes.
efficiency really matters and costs real money.
Of course if you're using dynamic allocation in an embedded system then you have other issues anyway...
If a developer is doing micro-optimizations without numbers of what it brings in terms of business case, then no.
Sometimes.
10 million times $.02 is only $200,000.
And the probability of you shipping 10 million units of anything is vanishing small. 10,000 is more typical--at which point the savings is $200. Not having to think saves you more money.
vanishingly small... I'm currently designing a product that will ship around 30 million units and have done similar in the past.
So, no... the point stands.
If you believe a millisecond is a short amount of time you're probably not writing video games. At 60 fps, you only have 16.67 milliseconds to render each frame, so wasting a couple of milliseconds would have a huge impact on what you'd be able to do for each frame.
Luckily, calloc doesn't actually take milliseconds unless you're making a very large allocation.
Care should be taken when it really matters, not in every single line of code.
Calling calloc/free on a 1 MB buffer seems to take about 50 microseconds on my computer (with no optimizations), and I consider that an unusually large allocation.
Do you actually touch it and write to it? If you don't the allocation may be provisional and not actually happen until you try to write memory, in which case depending on the system you'll either get a COW fault (IIRC linux points the allocation to a zeroed COW page, first time you need to write it'll copy the page and let you write in) or a zero-fill fault (e.g. freebsd doesn't allocate the BSS by default, write access will cause a zero-fill fault and in response the VM will actually allocate the page — which it can do from a queue of pre-zeroed pages)
calloc itself calls memset in user space. That's what separates it from malloc. Is it so hard to believe a modern computer can memset 20 GB/s?
If I run "dd if=/dev/zero of=/dev/null bs=1M count=10000", I get 21.0 GB/s. Reading from /dev/zero is basically performing a memset to dd's buffer in kernel space. Writing to /dev/null is a no-op, so it's a very similar test.
Have you considered reading my comment past the first phrase? If you don't actually touch the page, the system doesn't need to allocate it, let alone zero it. If you don't touch the page, assuming you're benching on linux you're measuring the cost of mapping a preexisting zeroed COW page, not the cost of actually allocating that memory.
This is why I said "in user space".
If you fail to do that, pre-zeroing might just give you consistent incorrect behavior. It can also inhibit a compiler's ability to detect that you're accessing an uninitialized object. The compiler has no way of knowing whether that zero-initialization was setting it to a meaningful zero value, or just "clearing" it for safety.
Other than that, you're right.
The first rule of C is to ALWAYS use C
The second rule of C is to NEVER use [u]intN_t, use [u]int_leastN_t instead and (unsigned) char when you want to deal with bytes
The 3rd rule of C is to never use anything else when size_t should be used
The 4th rule of C is that everyone has his own style, so declaring variables on top is totally acceptable (and more clean IMO)
The 5th rule of C is "enjoy your stack overflow" when using VLAs instead of malloc (not to mention that VLAs are optional in C11)
>which also means it's capable of holding the largest memory offset in your program.
wrong
>C99 gives us the power of <stdbool.h> which defines true to 1 and false to 0
I do use it but it is useless
>readability-braces-around-statements — force all if/while/for statement bodies to be enclosed in braces
don't use a shitty editor
>You should always use calloc
good thing calloc has: checks for overflows when multiplying the typesize and the arraylen
bad thing: makes programmers think that it will make your pointers null
>you can wrap it with #define mycalloc(N) calloc(1, N).
but the two arguments is the only good thing calloc has
>growthOptional
I can't see any reason why the second one is considered better
printf("Local number: %" PRIdPTR "\n\n", someIntPtr);
which is quite unreadable to say the least. To address this issue I even wrote a small library (https://github.com/cppformat/cppformat) that allows you to write something like this instead: print("Local number: {}\n\n", someIntPtr);
It is written in C++ rather than C, because the latter doesn't have facilities (at least overloading is necessary) to implement such library.- Option to insert commas or underscores every three digits (for decimal) or four digits (hex/binary/octal).
- Option to print floating in engineering format (make exponent a multiple of 3).
- Option to print floating but with a specific given exponent.
uint32_t numbers[64];
memset(numbers, 0, sizeof(numbers) * 64);
That is certainly true because that is a buffer overrun. memset(numbers, 0, sizeof numbers);[1] https://github.com/btrask/stronglink/blob/master/SUBSTANCE.m...
In C, with CHAR_BIT defined as 8, unsigned char is the only type that satisfies the requirements of uint8_t.
( It is also possible that it is defined as char, if implementation defines char to have the same range, representation, and behavior as unsigned char. In this case, char effectively becomes unsigned char, but is still a distinct type. The types are compatible (can alias).)
Now why is unsigned char the only possibility. Type uint8_t is defined to have two'2 complement, have exactly 8 bits, no padding, and be unsigned. So we need an unsigned integer type. The unsigned type that follows unsigned char in rank is unsigned short char. But this type is defined to have ranges at least from 0 to 65535. Since CHAR_BIT is 8, unsigned short int cannot be used, because it has to have more than 8 bits. As there is no type in rank before unsigned char, it remains the only possibility.
See also this suggestion that the gcc people, true to form, considered making uint8_t something other than unsigned char: https://news.ycombinator.com/item?id=10868953
Or it could use a typedef .e.g typedef __special_u8 uint8_t; where the particular compiler internally knows about __special_u8 and treats it differently from unsigned char.
6.2.5.15 - http://port70.net/~nsz/c/c11/n1570.html#6.2.5p15
6.3.2.3.7 - http://port70.net/~nsz/c/c11/n1570.html#6.3.2.3p7
B.19 - actually I think 7.20.1 is a better demonstration that these are integer types (and not, since it's never mentioned, character types) - http://port70.net/~nsz/c/c11/n1570.html#7.20.1
7.20.1.1.3 - http://port70.net/~nsz/c/c11/n1570.html#7.20.1.1p3
http://port70.net/~nsz/c/c11/n1570.html#6.2.5p4 leaves open the option for extended signed integer types, and the following paragraph provides for an unsigned integer type for each signed type, extended signed types included.
So it would be perfectly possible, if not actually very reasonable - not that this is stopping anybody at the moment - for an implementation to provide __int8, a signed 8-bit integer type that fulfils all the requirements of uint8_t, but is not a character type, and use that as uint8_t. Then, conceivably, it could fail to support the use of uint8_t pointers to alias other objects, something supported by the more usual situation of using unsigned char for uint8_t.
I'm not sure that anybody would do this, but they could. I'm rather surprised gcc doesn't do it, come to think of it, just to teach its users a lesson. Maybe I should file a bug report?
I looked briefly previously and somehow got the sense that extended types could only be larger than standard types, but this time I found that 6.3.1.1 explicitly mentions same width extended types.
Personally, I've been happy with GCC's decisions on such issues.
As for the actual C content, frankly this was where I went to when I was wondering how to write better (modern) C, so I don't really know any better. The writing is almost certainly opinionated and as above, could be a lot more comprehensive.
(oops edit window left open for a while, see sibling!)
make -j doesn't help linking speed, unless multiple binaries are begin built.
The linker invocation is a single make job, which is the unit of make -j parallelism.
It was very challenging! I was learning how to do threading in C for what was one of the most complex assignments in my degree. However, I don't regret a moment of it. Even with a mark that was lower than my average programming assignment marks, I learned more in that assignment than I did in anything else while at uni.
Being forced to RTFM because it wasn't abstracted away for you was exhilarating once I understood what was going on. And man, it was fast! Much faster than others who had programmed in Java.
I wish I had more reasons to write C, because I love it.
>The only acceptable use of char in 2016 is if a pre-existing API requires char (e.g. strncat, printf'ing "%s", ...) or if you're initializing a read-only string (e.g. const char hello = "hello";) because the C type of string literals ("hello") is char .
Heh, basically you need to use 'char ' for strings. I would far prefer if strings were unsigned so that I could use 'uint8_t ', but if you try it you'll get tons of conversion warnings and your code will look weird.
Void pointers do not allow pointer arithmetic (GCC allows it as an extension).
The correct type for pointer math is uintptr_t defined in <stddef.h>.
I've worked on systems (Cray vector machines) where converting a pointer to an integer type, performing arithmetic on the integer, and converting back to a pointer could yield meaningless results. (Hardware addresses were 64-bit addresses of 64-bit words; byte pointers were implemented by storing a 3-bit offset in the high-order bits of the pointer.)
Pointer arithmetic works. It causes undefined behavior if it goes outside the address range of the enclosing object, but in those cases converting to uintptr_t and performing arithmetic will give you garbage results anyway.
The casts to uintptr_t are not needed in this case, as the result of subtracting two pointers already yields ptrdiff_t.
Also, the mere subtraction of two pointers that do not point to the same "array object" yields undefined behaviour as per the standard.
In practice it works only if sizeof(ptrdiff_t) == sizeof(uintptr_t) AND the system uses two's complement for signed integers. Neither property is guaranteed by the standard.
I think the correct way to find the byte offset is this:
ptrdiff_t diff = (char*)ptrOld - (char*)ptrNew;~1 beat~
Yikes!
~leaves self-shaped dust cloud behind as hurriedly runs toward higher [level] ground~
But I agree, a short guide for C++ with the most important stuff in it would be nice.
I was dreading getting back to C++ with the Rule-of-3 and now the Rule-of-5 -- what a pain! Fortunately I ran into the Rule-of-Zero as described here: http://accu.org/index.php/journals/1896. The CppCoreGuidelines have a more restrictive Rule-of-Zero but allude to the link's form . . . the only way I'll get back to C++ is if I can avoid move-constructor and move-assignment for the majority of my work ... and the link provides the answer!
We've got these facilities, they're not exactly bleeding edge (the clue's in the name C99), let's use them!
I like using anonymous structs with variable initializers to provide parameters to functions with default values -- nice and clean!
C99 is not your grandparents' C. It has kept up with the times.
A more thorough resource I cannot recommend highly enough: 21st Century C[0]
I would not even dare to use threads without the support of these tools. In fact, this seems a side where C has a big advantage over other languages.
So, you're back to writing C89, in 2016.
For bonus points it shouldd prefix the clang-format command with "exec" but that's not a correctness issue.
Now, if you want to talk about signed overflow, bitwise operations on signed types, wacky undefined behavior, or memory management, that stuff is pretty hard to deal with in C. But 99% of the time when people are shaking their fists at "shitty ol' C ruining the Internet for everyone", they mean buffer overflows, and it just shouldn't be an issue these days.
Consistency means avoiding special cases. Variable length arrays are too prone to special cases - IMO, just don't use them. malloc vs calloc is too prone to special cases, that rule is worthless. The rule "never use memset" is full of special cases, also worthless. (And in fact, for your larger automatic initializations, gcc just generates code to call memset anyway.) I would say "always use memset when you need zeroing - then you'll never be surprised." The fewer special cases you have to keep track of, the better.
"Never cast arguments to printf - oh but for %p, always cast your pointers to (void *)" again, worthless advice with an unacknowledged special case.
The way to write code that always works right is to always use consistent methodology and avoid any mechanism that is rife with special cases.
The advice "always return 'true' or 'false' instead of numeric result codes" is worthless. Functions should always return explicit numeric result codes with explicitly spelled out meanings. enum is nice for this purpose because the code name is preserved in debug builds, saves you time when working inside a debugger. But most of all, it's a waste of time for a function to return a generic false/fail value which then requires you to take some other action to discover the specific failure reason. POSIX errno and Windows GetLastError() are both examples of stupid API design. (Of course, even this rule has an exception - if the function cannot fail, then just make it return void.)
I actually like declaring my variables at the top. To me, this seems more readable than searching for declaration statements throughout the code.
Why should I not do this?
If there's a section of code in your function that may or may not be executed, depending on a conditional, then allocating the memory for variables used within that code block will be a waste if you end up not executing it. Now, you could argue that that block should be turned into a separate function, but now you're doing another function call (unless it gets optimized out) which wastes performance, and you're making the code more cluttered if the code block is relatively small and not really worth turning into a separate function.
Allocation of N variables on the stack is O(1) since it just moves the stack pointer by the total size of all variables.
It makes the code more readable.
{
int x = 3;
do_something();
int y = 4;
}
you write this: {
int x = 3;
do_something();
{
int y = 4;
}
}
As a Lisp programmer, I prefer clear binding constructs which put the variables in one place. The (define ...) in Scheme that you can put anywhere makes me cringe; you have to walk the code to expand that into proper lambdas before you can analyze it. Other Lisp programmers don't agree with me. The CLISP implementation of Common Lisp, whose core internals are written in C and which historically predates C90, let alone C99, uses a "declarations anywhere" style on top of ANSI C. This is achieved with a text-to-text preprocessor called "varbrace" which basically adds the above braces in all the right places.You get better locality with explicit binding blocks, because their scope can end sooner:
{
foo *ptr = whatever();
/* this is a little capsule of code */
/* scope of ptr ends */
}
{
foo *ptr = whatever_else(); /* different ptr */
/* this is another little capsule of code */
}
These tighter scopes are much more tidy than the somewhat scatter-brained "from here to the end of the function" approach.I can look at that capsule and know that that ptr is meaningful just within those seven lines of code. After the closing brace and before the opening one, any occurrence of the name "ptr" is irrelevant; any such occurrence refers to something else.
https://hg.mozilla.org/mozilla-central/file/tip/mfbt/SizePri...
It's not exactly 100% complete in the literal sense, but everything covering the language spec and usage is there. I highly suggest working through it, as the book itself is incredible for going from beginner to expert in the short span it does.
Also, VLAs are optional in C11, and may not be supported in otherwise-near-C99 compilers for small processors. (Of course, alloca() is not standard at all.)
There's no way to recover from any stack allocation.
#define TOO_BIG 1000000
This: {
double big_array[TOO_BIG]; // not a VLA
}
is no safer than this: {
size_t size = TOO_BIG; // size is not a constant
double big_vla[size];
}
If you use VLAs (variable length arrays) in a manner that permits arbitrary unchecked sizes, then yes, they're dangerous. If you carefully check the size so that a VLA is no bigger than a corresponding fixed-size array would have been, they don't introduce any new risk.(I'm not sure why VLAs were made optional in C11, at least for hosted implementations. Support for them was mandatory in C99, and I'm skeptical that they impose any significant burden on compilers.)
Note that the overarching theme of C99 was making C a better language for numerical code (ie competing with Fortran), thus the introduction of VLAs (as well as _Complex, tgmath.h, restrict, and possibly other things I forgot ).
There is also the fact that you permanently lose one register in your function (ebp) if you use alloca(). This is important on 32-bit x86 but is rarely discussed.
There is essentially no implementation difference between VLAs and alloca().
The "machine dependent" reason is out of date though.
Respect earlier developers and try to minimise _unnecessary_ changes that break version control history. Did earlier developers use names you don't like? Formatting you don't like? InconsistentCamelCasing or other_naming_paradigms you don't like? Let it be -- at least until you do a refactoring where _you_ claim responsibility.
It may not be a good point to uncritically and retroactively apply the good points from the original article in an old code base.
2 bit example:
00 => -1 01 => 0 10 => 1 11 => 2
or: 2 bit example:
00 => -2 01 => -1 10 => 0 11 => 1
This even wrapsaround naturally.
test.c:1:2: error: #import is a GCC extension
This ought to carved into the walls of R&D departments everywhere.
I realize this is a bit of a style question, but I'm curious.
Unlike gcc, clang does not have good documentation for its warning options. The best sources I've found are the clang code itself (https://github.com/llvm-mirror/clang/blob/master/include/cla...) and http://fuckingclangwarnings.com/.
- Add security measures provided by your compiler to your builds: -fstack-protector-strong on GCC. -fsanitize=safestack on Clang if you have a version that's new enough.
- Build the program with AddressSanitizer and run it with typical data. Run it through a fuzzer afterwards.
for (infer x = 0; x != 10; ++x) ...
If the compiler can infer the type, use the fastest integer. If it can not infer, use the largest integer.
I'm sure this would create difficulties with structs, but at least for local integers it would be nice.
This is IMO very close to the inference you describe, while still providing some stipulations about what you need.
int8_fast_t and its siblings are defined in <stdint.h>, described in this article (though it would be nice if the article mentioned them).
Unless you're doing memory-mapped IO, or other very-specific things, you probably only need the _least_t or _fast_t types (int16_least_t, e.g.).
When I hear buffer overflows I think.
Programmer trying to be smarter than he is. Depreciated unsafe string functions. Unnecessary pointer arithmetic. No unit tests Failing to use static code analysis tools
My only concern about buffer overflows in C is that memory error continue to happen in software that I depend upon in my daily life.
Chrome, Firefox, Linux, OpenSSL, all these things suffer memory errors that compromise my security. Anyone doing security work in C in 2016 is in my opinion committing malpractice and putting user's at risk because their ego's can't take not fiddling bits by hand.
> The solution here is to always use an automated code formatter.
Spot on. Python got that part right.
Python, not so much.
But on the writing side, it's harder (impossible?) to have tools that format an entire file, because without braces, various things could be at different levels of indentation. You have to indent as you go, and changing the level of indentation can be a little painful.
In curly brace languages, you can sometimes work very fast by crapping things onto the page and then formatting the entire file.
I don't know which I absolutely prefer, but there is a tradeoff to whitespace based blocks. That said, I think C & Java suffer for the combination of optional braces and lack of a standard format tool.
Instead use C99 ... which is from 1999. (Yes, I know most people didn't learn it in the 90s)
In practice a lot of c99 got into gcc after 3.0 (2001) and it still took a bit of time for any distribution to include it as stable. If I remember correctly, lots of distributions held on to 2.9x for as long as possible. I'd say realistic c99 support has ~10 years now, rather than 17.
I see this a lot shouldn't O3 produce better optimizations? Is that people have large codebases? OR is it that O3 does some CPU/memory tradeoffs? OR something else?
ALSO it generally produces a larger binary which means less of it fits in the cache. I know a while back it was often suggested to try O3 and OS because often OS ended up being faster.
I usually compile with -O2 and annotate certain functions with O3 - if it looks like vectorization will give the function a boost.
In some ways -Wextra is too much: it's annoying that -Wall with -Wextra complain about unused parameters.
I am not a fan of the use of macros here. You should never use macros.
Instead maybe: static inline void *mycalloc(size_t sz) { return calloc(1, sz); }
lmaoed at that one
> Standard c99
If you have to write C, why not go with ANSI C/C89 for compatibility/portability reasons?If publishers and programmers with popular blogs can't get a consensus on what constitutes correct C, how could I ever hope to produce it?
Case in point https://news.ycombinator.com/item?id=10864601
int is a type that's at least 16 bits large. I can't see the problem with using it. If I specify that I want a type that's 32 bits large, my code won't run on 16 bit systems.
Why should I assume? Because the author doesn't believe in writing code that works on "20 year old systems"?
Why not?
I've read that if one uses int64_t with a compiler that targets a 32 bit system, doing various operations on the int64_t variable will be much slower because more machine code instructions need to be generated.
The point stands: Why not use int if one is looking to target any architecture that C will run on?
In most places, if you're using an `int` for non-address-y stuff you're just wasting space on wider archs, or you're overflowing on less wide ones. The only exception off the top of my head is "I have an array of bytes, but want to bulk process more efficiently than u8 would allow". Things like `size_t` (or explicit SIMD types!) are still probably a better choice there, if only to simplify the rule to "just don't use int".
This will often generate more code. So don't do this except when it saves a modest amount of memory.
Yep, but presumably you picked int64_t for a reason. If you want a 32-bit variable, use int32_t.
It would be an extremely rare platform on which int16_t is noticeably slower than int. The very minor platform-specific performance improvement you might get from using int rather than int16_t is very rarely worth the risk of introducing an overflow on the few platforms where int is 16 bits[1], and not noticing it because int is longer than that on the platform you develop on.
[1] remember that a signed integer overflow is undefined behaviour which should be regarded as equivalent to a security flaw - the compiler is permitted to send all data your program has access to to the mafia if your program contains a signed integer overflow anywhere, and some compilers will
If they used 'int_fast16_t' from <stdint.h>, you have at least an idea that they considered that it might only be 16 bits, and that it can be a larger type on systems where operations on 64-bit types are faster.
That said, there's really no way to correctly pick the "perfect" integer size for every platform - There's just too many variables. IMO I much prefer to just use 'int' when I know the values are going to be small - smaller then a 16-bit value - and I don't care about the actual width. But obviously it's a hot topic for some - As long as you follow the standard in regards to it's size then it doesn't really matter what integer value you choose to use. I prefer saving the fixed-width integer types for situations when I know I need such a size.
If the loop variable is a int_fast16_t the loop increment step compiles to
addq $1, -16(%rbp)
If it is int16_t it compiles to movzwl -4(%rbp), %eax
addl $1, %eax
movw %ax, -4(%rbp)
The 64 bit width version is shorter, and so conceivably faster. Performance seems the same though, but that's probably because this program isn't a proper benchmark.Without profiling there's no way to tell on modern processors.
Comparatively, while one is three instructions and the other is one, that single instruction is still doing everything the other three are doing. The only reason they aren't both one instruction seems to be that `addw` must not being a thing (Why, I don't know). Since the I/O done by both will be nearly identical, and the addition's should be very comparable in speed, it's not crazy to say they should perform basically the same. If you compiled with -O2 I bet you'd see almost identical code, since it would probably remove the memory access.
Yes, the int is technically less portable. No, I don't see any reason to use that longer name.
`int` is at least 16 bits, `long` is at least 32 bits, `long long` is at least 64 bits, and `size_t` is big enough to hold the size of the largest object you can allocate - and in most cases that is all you need to know.
The exact width types can be useful if you're marshalling data to and from an exact-width format, like a network protocol or file format.
You should know the range of values that you are expecting to be holding in an integer, and choose the type accordingly.
The only advantage to choosing an exactly-N-bit type over an at-least-N-bit type is possibly saving some space in very large allocations. In most cases this won't be important - and when it is, the mininmum-width types from stdint.h (eg int_least32_t) are more apropos than the exact-width types.
The disadvantage is the impedance mis-match with all the existing libraries of code using the basic types.
IMHO, "int" is perfectly acceptable if you just want to use the platform's preferred word size (and are fine with the possibility of it being only 16bit). Also, "char" is perfectly acceptable too if you're dealing with ASCII. This is valid for both the signed and unsigned variants.
The general rule should be using the language types as long as you're not assuming something about their size, only then use "stdint.h" because it makes your assumption more explicit.
Now, nobody uses "short" or "long" to mean "something smaller or shorter than the platform's int", so these are usually something to avoid.
Same goes for unsigned versions and long long.
Granted, divisions are rare, but that's still nearly an order of magnitude difference in that corner case.
32-bit and 64-bit addition, multiplication, comparison, logical operations etc. are in almost all cases just as fast.
You'll still need more storage on the stack for 64-bit variables. So more cache misses (and perhaps TLB misses and page faults) when corresponding boundary is crossed.
Have a reference to the norm easily accessible and don't use options of the compiler which impact you don't fully know.
Most developer are poorly understanding multithreading and it can be tricky to make a portable library that has this property. Don't hesitate to stipulate in your documentation that you did not cared about it, people can then use languages with GIL (ruby, python ...) to safely overcome this issue.
Modern language are social, take advantage of sociability (nanomsg, swift, extension, #include "whatever_interpreter.h" ....).
Programming in C with dynamic structures is guaranteed to be like a blind man walking in a mine field: you totally have the right either to make C extension to modern language OR include stuff like python.h and use python data structure managed by the garbage collector from C.
Old style pre allocated arrays may not be elegant, but that's how critical system avoid a lot of problems.
DO NOT USE MAGIC NUMBERS...
USE GOTO for resource cleaning on error and make the label of the goto indentend at EXACTLY the same level the goto was written, especially when you have a resources that are coupled (ex: a file and a socket). Coupling should be avoid, but sometimes it cannot. Remember Djikstra is not that smart, he was unable to be a fully pledged physicist that can deal with the real world.
Avoiding the use of global variable is nice, but some are required, don't fall for the academic bias of hiding them to make your code look like it has none.
Copy pasting function is not always stupid (especially if you use only 1 clearly understandable function out of a huge library).
Be critical: some POSIX abstraction like threads seems to be broken by complexity: KISS.
Modern C compiler are less and less deterministic, learn about C undefined behaviour and avoid clang or GNU specific optimisations and CHECK your code gives the same results with at least 2 compilers.
Don't follow advices blindly: there are some C capos out there that dream an ideal world. Code with your past errors in mind and rely on your nightmarish past experiences to guide you to avoid traps. Feeling miserable and incompetent is some time good.
Resources are limited : ALWAYS check the results of resource exhaustion in C instead of relying on systems and always have error messages preallocated and preferably don't i18n them. Don't trust the OS to do stuff for you (closing files, SIGSEV/SEGFAULT ...).
KISS KISS KISS complexity is your enemy in C much more than in other languages. Modern C/C++ want you to build complex architecture on libraries that are often bloated (glibc).
Egyptians are nice looking.
Once you finished coding you have done only the easy part in C: - dependency management is hard in C; - have a simple deterministic build chain; - you should write your man pages too; - you should aim at writing portable code; - static code analysis, maybe fuzzing, getting rid of all warnings is the last important part.
Coding in C is just about 15% of the time required to make great software. You will probably end up debugging it much more than you write it: write code for maintainability as your first priority and do not under any circumstances think that you can handle multiple priorities (memory use, speed, ...).
Above all : just write maintainable code, please. This rule is the golden rule of C and it has not changed.
I would add one thing.
Know your target. Batch programs, daemons, microcontroller programs, cross-platform programs are all different and you're going to end up with different styles.
Well that is the difference between writing code and bringing a solution that fits a problem.
I fully agree.
But IT has become shove sellers to other shove seller hoping for a new gold rush to happen.
Computers have still not proven to be an actual improvement to the economy. Positive impact cannot be measured. I guess I know why.
It looks like the lost bet of Diesel to empower the poorest to have access to the economy and fight equally with capitalists able to buy steam machine thanks to the banks: IP laws makes it impossible for a new economy to rise.
Those who do not learn from the errors of the past are doomed to reproduce them.
Ha ha ha. What kind of programmer would argue this? Do they ignore Windows 64? The type long is 32 bits there.
so defining a string happens how if not:
char *my_string = "Hello";
i meant a variable
char * my_string
It must be a const char * my_string You can still assign a different string to that, as my_string is not const.
For my_string to be const, you'd have to use char * const my_string
Now the location pointed to by my_string is writable but I cannot change the location.
To make it entirely unmodifiable, you'd use const char * const my_string.
On the other hand, it might make more sense to use an array here.
const char my_string[] = "hello";
No thanks; sticking to C90.
> LTO
Violates ISO C, which says that semantic analysis is done in translation phase 7, and only linking takes place translation phase 8.
Link-time optimization is a language extension which alters the concept of a translation unit. In some ways, multiple translation units become one unit (such as: they inline each other's code, or code in one is altered based on the behavior of code in the other). But in other ways they aren't one translation unit (e.g. their file scope static names are still mutually invisible.)
One reason I don't use C99 is that many of its features aren't C++ compatible. Most of C90 is. With just a little effort, a C90 program compiles as C++ and behaves the same way, and no significant features of C90 have to be wholly excluded from use.
Portability to C++ increases the exposure of the program to more compilers, with better safety and better diagnostics.
In some cases, I can use constructs that are specially targetted to C or C++ via macros. Instead of a C cast, I can have convert(type, val) macro which uses the C++ static_cast when compiled as C++, a coerce(type, val) which uses reinterpret_cast, or strip_qual(type, val) which uses const_cast. I get the diagnostics under C++, but the code still targets C. I'm also prevented from converting from void * without a cast, and there are other benefits like detecting duplicate external object names.
This way of programming in C and C++ was dubbed "Clean C" by the Harbison and Steele reference manual.
C90 is a cleaner, smaller language than C99 without extraneous "garbage features" like variable-length arrays that allocate arbitrary amounts of storage on the stack without any way to indicate failure to do so.
The useful library extensions in C99 like snprintf or newer math functions can be used from C90; you can detect them in a configure script like any library.
The overall main reason is that I simply don't have faith in the ISO C process. The last good standard that was put out was C90. After that, I don't agree with the direction it has been taking. Not all of it is bad, but we should go back to C90 and start a better fork, with better people.
git checkout -b c90-fork c90
git cherry-pick ... what is good from c99-branchBut here is the thing: all the initialization comes from the initializer, not from the struct type, and the only default value you get is the good old 0 (0 integer, 0.0 floating point, null pointer). So, if you want any other default, designated initializers are basically useless.
The designers of a data type want to provide initialization mechanisms with defaults which they control from the site of the definition of the type. We can do that in C90 with macros. Macros ensure that everything is initialized. If we add a member that must be initialized with a non-default value everywhere, we add a parameter to the macro, and the compiler will catch all the places where the macro is called with the old number of parameters.
struct foo {
char *name;
int bogosity;
};
#define foo_anon_init(bogosity) { "anon", bogosity }
#define foo_init(name, bogosity) { name, bogosity }
You provide all the variants in one place which do all the forms of defaulting which you need, and which stick the parameters into the correct place, and then just use them everywhere else.Even if you have designed initializers, it's still good idea to wrap initialization that way; they don't replace its benefits.
Suppose I add a frobosity member, and there is no obvious way to default it. All the places where the structure is instantiated have to be analyzed case by case to choose the appropriate value. Designated initializers (alone) won't help; we add the member to the struct and nothing happens. It defaults to zero everywhere it is not explicitly mentioned with a designated or positional initializer. With the macros we just add it as a parameter:
struct foo {
char *name;
int bogosity, frobosity;
};
#define foo_anon_init(bogosity, frobosity) \
{ "anon", bogosity, frobosity }
#define foo_init(name, bogosity, frobosity) \
{ name, bogosity, frobosity }
Now "make -k" and fix all the places where the macros are called with too few arguments, choosing a value for the frobosity.With designated initializers we can of course do:
#define foo_anon_init(bogosity, frobosity) \
{ .name = "anon", .bogosity = bogosity, \
.frobosity = frobosity }
But that brings very little to the table when we've gathered all this stuff in one place (which alone reduces mistakes) and we are initializing every member to a nonzero value.Languages that have keyword initializers are great when the class supplies defaults for all the ones that the constructor doesn't specify.
(make-instance 'foo :frobosity 42) ;; 50 other slots defaulted for us
If you want to change a default, you don't want to go to all the places where the type is instantiated and update the initializer!!!Eek! Not what I was expecting. In the video game world there is a not so small contingent of "write everything in C unless you have a really damn good reason not to".
Great post though. Now go forth into the world and write C, my friends!