> I doubt anyone is willing to pay 2x performance for their C code to be more friendly. If they were, why are they writing in C?
Beats me why anyone is writing in C full stop.
> I doubt anyone is willing to pay 2x performance for their C code to be more friendly. If they were, why are they writing in C?
Beats me why anyone is writing in C full stop.
I fully agree. But this is not without tradeoffs: eliminating undefined behavior pretty much means (a) the performance and runtime overhead of a GC (most languages); (b) forbidding malloc (verified "mission-critical" variants of Ada, C, etc.); (c) requiring programmers to learn a lot of new concepts (Rust†). I do think that there is rapidly becoming little reason to use C except for throwaway programs, and as an industry we need to be more open to (c) if we are ever going to move beyond making the same memory management mistakes we've been continuously making since the 1960s. But I also understand the reasons why programmers continue to choose C and C++.
† It's interesting to me that the biggest reason C++ programmers give for bouncing off Rust is fundamentally "I want my undefined behavior back", though very few actually word it like that.
Disallowing undefined behavior simply means the result is defined behavior. Might be safe, unsafe, any number of things. Just clearly that to developer and tools.
In your proposal, how do you propose to solve the load-load forwarding issue described in the article?
This is a technical problem that needs a specific technical solution, not philosophy.
My position is the same as Tony Hoare, optimizations that exploit undefined behavior or contribute to possible malfunctioning programs are not to be done at all.
Then how come programs written in languages which don't have undefined behaviour, like Java, are not 2x slower than those written in C?
I really don't consider this a criticism of C per se or its designers. It is unreasonable to expect that a language specified in the 1970s would be tuned for either modern processors or have the benefit of an additional 40 years of collective experience. But it is, nevertheless, true.
(It's part of why my position is very much that this entire approach just won't work. C is pretty much what it is; it can not be changed. It is literally easier to use another language than to try to specify a "friendly C" in 2016. And I don't say that because I don't know how large a task it is to start writing code in another language; I say that because people are really underestimating how much work Friendly C will take. I suggest that everyone observe that all who have now seriously tried to sketch what this would look like have immediately given up and totally surrendered. This means something.)
Back when the language was being standardize, there was way more hardware diversity than there is now and C implementations outside UNIX, even K&R ones, had lots of semantic differences.
As no compiler vendor involved in the first ANSI C process wanted to give up on their semantics, all of them were swept into "undefined behavior".
https://en.wikipedia.org/wiki/List_of_Java_virtual_machines
It is like I would measure C performance by picking a specific compiler, even though I can choose any of these ones:
http://www.azulsystems.com/products/vega/processor
Well, would be if I was a Java salesman. ;)
A proper comparison would be to something like Component Pascal, Modula-2/3, Ada, Fortran for numeric algorithms, and so on. Their results were high performance with higher correctness and predictable behavior. So, parent's point stands and so does mine that only C's bad design and culture lead to relying on undefined behavior for performance. They could just as easily optimize from well-defined semantics into assembly like the other languages mostly did.
The answer varies based on which specific optimization you're talking about, but for the optimization described in the article the answer is simply that Java is type safe. Notice that the load-load forwarding optimization relies on int pointers and float pointers being guaranteed not to alias. Strict aliasing is one way to guarantee this. But another way is simply being type safe! Type safety means that int pointers and float pointers (in Java, int[] and float[]) can't alias by construction--otherwise the types would be wrong. So a Java compiler can just do this optimization without worrying about it. And it's free of undefined behavior.
I like this example because it shows how the "close to the metal" nature of C and C++ can actually hinder optimization opportunities, as counterintuitive as that may initially sound.
You avoid relying on a compiler specific behavior, that will break the moment you switch compiler, compiler version or CPU target.
Instead you get a very clear set of Assembly instructions, that will never change for the specific CPU family.
The argument of writing Assembly being non-portable, doesn't play a role, because such optimizations are not even guaranteed to stay stable for the same compiler vendor.
In any case, only after a profiler has proven they are really a case for spending development effort on.
My experience with the C and C++ developer culture back to the mid-90's, is that many suffer by anticipation with needless premature micro-optimizations.
Not feasible. Quoting DannyBee yet again: "Most applications have flat profiles. They have flat profiles because people have spent a lot of time optimizing them."
When your application has a flat profile, you're spending most of your time in most of your code. So by saying that only code written in assembler should be optimized, you're asking most large C/C++ users to rewrite their codebases in assembler or suffer an unacceptable performance hit. Not going to happen.
I am aware of it.
They rather rely on something that isn't guaranteed to even survive a minor upgrade of the same compiler, and always act surprised when they discover what compiler does and the standard says don't go always hand in hand.
So no reason to complain when those long weekends come in to track down optimizer induced bugs.
EDIT: There is always Assembly, after properly using a profiler, instead of relying in undefined behavior outcomes that are compiler specific and can even change between releases.
Plus, most of the software runs on x86 (now ARM) anyway. So, that's two ISA's to account for outside of operating systems, embedded, and ISV's for unusual platforms. Even then, one might use a subset of C with no undefined behavior and then make platform-specific optimizers. So, two possibilities that make a lot of sense vs depending on known unknowns plus transforms without safety arguments.
My overall point is that you can do an optimization from a set of statements with specific meaning at least as easy as you can from one with unknown meaning. So, a similarly low-level language with well-defined behavior could be optimized at least as much as C. However, since compiler faces uncertainty, the argument doesn't work the other way around. That's why C and its undefined behavior are bad by design.
And hence all the discussions on workarounds like the OP.
1. Interoperability code with some library or the OS.
2. C as the closest thing to a portable assembler there is (e.g. to implement something like the Ruby interpreter).
What I need for this isn't perfect safety; it's the ability to reason with some confidence about the code I'm writing (I may want to be able to rely on code review by merely mortal programmers or create my own tools to enforce this). If I wanted a safe, expressive language, I wouldn't use C in the first place.
But right now even something as simple as a malloc implementation is riddled with traps even for reasonably skilled programmers. The Linux kernel uses -fno-strict-overflow and the FreeBSD kernel uses -fwrapv because the price to be paid by enabling the undefined behavior of -fstrict-overflow was too high.
If by "interoperability" you mean "binding a more modern language to a C library", no argument there. Binding to C is a legitimate reason to write C.
> 2. C as the closest thing to a portable assembler there is (e.g. to implement something like the Ruby interpreter).
My question is: why do you want a portable assembler? Why not write Ruby in, for example, Go? It could certainly be done; look at JRuby, for instance.
The vast majority of the time, the answer to this question is "performance". Which brings us back to the point of the article. Without compiler optimizations enabled by undefined behavior, you can easily lose 2x performance or more. A 2x performance loss for MRI is unacceptable.
Malloc implementations embody this even more. In fact, it's kind of hard to think of any piece of code that's more performance critical than malloc. Many large applications (apps, games) spend 10%-20% of their time in the malloc and free functions. Allocator performance is so important that Facebook and Google have invested a huge number of man-hours into these two routines (jemalloc and tcmalloc respectively). It all comes down to performance again: under these extreme contraints, malloc simply can't afford a 2x performance loss. A Friendly C that produced slower code would have little chance of being adopted by allocator writers.
I would in principle like to do this in another language. Go isn't that, because Go's runtime makes some very specific assumptions (about things like stack layout and how it interoperates with syscalls). I may not want to be weighed down by these assumptions. For similar reasons, you may want to eschew JRuby, as it locks you into the JVM ecosystem (which is great if that's what you need, not so great if you don't want to suffer from the poor interop with non-JVM libraries).
> The vast majority of the time, the answer to this question is "performance".
The answer for me is generally that C is ecosystem-agnostic (what I'd expect from a portable assembly language; practically any other language that isn't called BCPL makes more ambitious assertions about its environment). Note that when I talk about interoperability above, I don't just mean interoperability between one language and C, but also to build bridges between two non-C languages (which is actually fairly important for some of my current work). Performance is a nice additional benefit.
> A 2x performance loss for MRI is unacceptable.
We have a few points to chew through here. First, I didn't say that undefined behavior is unacceptable. My problem is with undefined behavior that is difficult to reason about or to investigate (as an extreme case, when an infinite loop is turned into a no-op).
Second, I don't buy the 2x performance loss. What makes most bytecode interpreters slow is how C compilers optimize the dispatch loop (poorly), not the exploitation of undefined behavior or the lack thereof. For example, -fwrapv has virtually no effect on Ruby's performance (even though clang/gcc unnecessarily disable strength reduction in some cases where they don't have to). Ruby would likely benefit a lot more from having its bytecode dispatcher rewritten in assembler LuaJIT-style (I'm talking about the LuaJIT interpreter, not compiler) than it would lose from not exploiting undefined behavior on a large scale.
Finally: code that crashes really fast (or may suddenly start crashing for inexplicable reasons because the compiler was upgraded) is not going to help anyone. This is not a hypothetical concern; you'd be surprised how much C code doesn't survive an encounter with UBSAN.
If every important program on your computer (The ones written in C -- OS, web browser, text editor, language interpreters, etc.) suddenly became twice as slow for "safety", assuming you never modify those programs on your own, you'd probably be a little bit cross right? You wouldn't really care that the language is safer, you'd just notice your top of the line computer is suddenly feeling pretty slow.
Exactly. Read the comments on literally any thread about browsers on HN to understand why sacrificing that degree of performance is unacceptable.
i can list many reasons why i continue to use C, please, there's no need to suggest or imply that C has outlived its usefulness - it clearly hasn't.
This is what "undefined" means: what it say! It doesn't mean "illegal", and it doesn't mean "bad".
However, "undefined" does mean "illegal". Not only does the standard not say what will happen, the standard says that such behavior is not allowed in well-formed C programs. C compilers, then, tend to optimize based on the assumption that certain behaviors will not happen. This can causes weirdness when they encounter programs that exhibit this behavior - and yes, such programs have a bug. But often, the bug manifests in an even stranger way. Check out the work of John Regehr for more (academic publications: http://www.cs.utah.edu/~regehr/papers/; blog: http://blog.regehr.org/).
The C11 standard is however quite clear on what undefined behaviour is: "behaviour...for which this International Standard imposes no requirements". That sounds like a far cry from illegality to me! It sounds more like it is merely... yes... undefined.
This means that gcc (et al) are allowed to do what they do. It doesn't mean it's not crap.
3.4.1
1 implementation-defined behavior
unspecified behavior where each implementation documents
how the choice is made
2 EXAMPLE An example of implementation-defined behavior is
the propagation of the high-order bit when a signed integer
is shifted right.
3.4.3
1 undefined behavior
behavior, upon use of a nonportable or erroneous program
construct or of erroneous data, for which this International
Standard imposes no requirements
2 NOTE Possible undefined behavior ranges from ignoring the
situation completely with unpredictable results, to behaving
during translation or program execution in a documented manner
characteristic of the environment (with or without the issuance
of a diagnostic message), to terminating a translation or
execution (with the issuance of a diagnostic message).
3 EXAMPLE An example of undefined behavior is the behavior on
integer overflow.
And later, 4. Conformance
2 If a ‘‘shall’’ or ‘‘shall not’’ requirement that appears outside of
a constraint or runtime constraint is violated, the behavior is
undefined. Undefined behavior is otherwise indicated in this
International Standard by the words ‘‘undefined behavior’’ or by
the omission of any explicit definition of behavior. There is no
difference in emphasis among these three; they all describe
‘‘behavior that is undefined’’.
...
5 A *strictly conforming* program shall use only those features of
the language and library specified in this International Standard.3)
It shall not produce output dependent on any unspecified, undefined,
or implementation-defined behavior, and shall not exceed any minimum
implementation limit.
Substitute "strictly conforming" where I said (colloquially) "well-formed". Footnote 3 basically says that a program can still be strictly conforming if it uses of conditional features has conditional guards.All of that above tells me: "implementation-defined" and "undefined" are different concepts in the C standard, and programs with "undefined" behavior are erroneous - or, colloquially, illegal programs.
Oddly, I still agree with your final conclusion: gcc and other compilers are definitely allowed to do what they do; they are following the standard.
Embedded software? All of the platforms I work on provide C/C++ compilers and driver libraries that are C only. I'm working on an ARM Cortex project right now and there isn't even C++11 support.
C is so poorly designed [1] that we're just now getting formal semantics, reliable analysis, and CPU's that handle it safely. UNIX ecosystem is similarly broken [2] [3]. Intel CPU's have all kinds of legacy holdovers, bugs, over-privileged components, and complexity in general. So, calling the current stacks broken by design makes sense without any implications for future stacks that might be better or worse.
Want to have one that's better? Just make a sensible design, document it well, aim for simplicity where possible, keep it efficient, ensure its safe in normal operation, and give it good interfaces for extension/maintenance. Those tend to lead to better designs or implementations. The crap we use today largely evolved due to social and market forces, not good technical reasons.
[1] http://pastebin.com/UAQaWuWG
[2] https://queue.acm.org/detail.cfm?id=2349257
[3] http://linuxfonts.narod.ru/why.linux.is.not.ready.for.the.de...
So, no, C is not easy to reason about. Nor are the tools that can cheap if you want low, false positives.
No, people write C because it's the only choice in many cases. Like where you need total control over memory management, or because you want to interact with a kernel or hardware directly, or because you want to write functions in the lingua franca of compiled code.
I wouldn't be quick to use it for a large scale project which does not require 110% performance but for anything were performance is your nr. 1 goal I've found it quite refereshing.
You know you can't make such statements without putting your money where your mouth is. We want to see the code.
Here's three different 4x4 matrix multiplication routines I've written (the mmmul functions). Depends on your cpu which is fastest.
https://github.com/rikusalminen/threedee-simd/blob/master/in...
Granted, Intel's implementation was for 8x8 (floats), perhaps that makes a difference in the instruction pipelining. I'll see if it does later.
MxM_4x4 is my function, say we have AB=C then it loads B transposed into registers, then for each row it uses 4 multiplications to calculate to row(A)column(B) products, 2 HADD for summing half of the generated products, then 2 permutations to line up the vectors, and one final sum to compute a row of C.
MxM_4x4_2 is a direct port to doubles instead of floats of intel's example on the provided link. When compiled with -O3 my compiler produces the same code in terms of assembly instructions as Intel's example explictlly writes.
(Note the name of the repo, I know it's not pretty ;) )
Interestingly, this is why I enjoy working with Go.