C++ design goals in the context of Rust (2010)
pcwalton.blogspot.com
pcwalton.blogspot.com
C is not that fast. One of the major problems is that it's close to hardware. 1970s hardware that is. Ken Thompson reportedly once said: "I'm not going to do nibbles. I have an 8 bit processor".
A good example of how bad it has become is that modern processors have a rather good understanding of the 'string' concept, and offer instructions to process them. C offers a char*.
Another problem is that C has strict contracts on how parameters are to be passed through, and combined with separate compilation units, this hurts compilers when they try to optimize things.
I see greater potential for a safe higher level language to be able to align closer to modern day hardware than C. Some nice examples: Linear types can avoid garbage altogether, and coroutines can be expressed clearly and correctly using monads. On the other hand, raw performance is rarely needed, and most cycles are burned interpreting things like python and php.
Why is that so? Do more modern languages need another twenty years before they can compete with C (and eventually beat it performance-wise) ?
"Not that fast" doesn't necessarily mean that it isn't the fastest we have got.
When I started coding in the mid-80's C compilers were dog slow and no better than the alternatives (Pascal, PL/I, Modula-2, Cedar, ...).
In fact, game developers would talk about high level languages the same way they talk about current modern languages vs C/C++, and use nothing else other than Assembly.
However with UNIX spreading into the enterprise and bringing C along, it meant compiler vendors focused on optimizing for the language they were getting money for.
The main issue when people compare languages, is that they forget although the design drives the implementation, not all implementations are alike.
ps: for instance in some cases, jitted code will go faster than C, some Jitted kernels (never tested personally) allowed for a good 30% performance increase.
The x86 string functionality that I'm aware of was directly derived from C string functionality, and any decent library of course uses them. You don't directly map to those underlying opcodes because that would be silly, and would completely undermine any platform independence you might have.
The same for vectorization. You don't explicitly express vectorization in your code, but of course all decent C compilers can easily and robustly generate such code.
C isn't close to the hardware (beyond very high level notions like "contiguous memory"). But it's a simple enough language that it's very heavily optimizable.
From what I've seen auto-vectorization is anything but robust, often you need non-portable annotations to get anything decent. See e.g.: http://locklessinc.com/articles/vectorize/ where there's an interesting case where using "if" instead of a ternary operator is enough to prevent optimization.
The string instructions tend to be slower than using generic instructions, especially in the presence of SSE.
The simplicity of C means you can't make many assumptions of how it is handled. Same problem in C++. You can make a lot of optimizations to assembler if you can guarantee a value won't change, or if the relative location in memory needs to remain static because anything can rip the address and do awful pointer arithmetic with it.
I like to compare to asm.js - it is a reduced language in that it restricts the environment from doing things it can't easily optimize away, except in that case its to assembler. Reduced set versions of C exist to do the same already, but it would be nicer if you just had a fast language default with straightforward rules with ways to just declare wonderland functionality that can mess things up that means the compiler needs to preserve the machine assumptions rather than the language structures.
I once wrote a SIMD C++ template library based on valarray: http://www.pixelglow.com/macstl. Sadly it has languished over the years but I would love to work on it again. It had novel (at the time?) vectorized trigonometric functions, for example.
These days the most important factor to performance, in almost all programs, is memory accesses.
Writing cache-friendly code is made relatively easy by C. You control the placement of your data into aligned cache lines. You control the order in which the cache lines are accessed.
The cost of the computational instructions is usually swallowed by the memory misses.
C's common ABI (for x86/64) now passes up to 6 parameters via registers, and doesn't unnecessarily spill to the stack. Additionally, link-time optimizations now allow cross-compilation-unit inlining which can optimize away the ABI costs.
Of course, a C program doesn't have to use C strings, but so many libraries and APIs do (including the C standard library) that they are unavoidable for most practical purposes.
D eliminates this problem by using dynamic arrays to represent strings. A dynamic array is a pointer/length pair.
Let's consider now why C is a great language. It is commonly believed that C is a hack which was successful because Unix was written in it. I disagree. Over a long period of time computer architectures evolved, not because of some clever people figuring how to evolve architectures---as a matter of fact, clever people were pushing tagged architectures during that period of time---but because of the demands of different programmers to solve real problems. Computers that were able to deal just with numbers evolved into computers with byte-addressable memory, flat address spaces, and pointers. This was a natural evolution reflecting the growing set of problems that people were solving. C, reflecting the genius of Dennis Ritchie, provided a minimal model of the computer that had evolved over 30 years. C was not a quick hack. As computers evolved to handle all kinds of problems, C, being the minimal model of such a computer, became a very powerful language to solve all kinds of problems in different domains very effectively. This is the secret of C's portability: it is the best representation of an abstract computer that we have. Of course, the abstraction is done over the set of real computers, not some imaginary computational devices. Moreover, people could understand the machine model behind C. It is much easier for an average engineer to understand the machine model behind C than the machine model behind Ada or even Scheme. C succeeded because it was doing the right thing, not because of AT&T promoting it or Unix being written with it.
I've spent some time working on VLIW architectures where C is a terrible abstract representation of the CPU's inner workings. Get anything done efficiently was awful and required huge, cumbersome, carefully structured intrinsics. Or raw assembly.
Is this is fault of VLIW? Not really. Its mostly the reality of a world that would rather run C code than use VLIW architectures. So we don't have VLIW or other CPU architectures that don't work well with C.
"people could understand the machine model behind C" -> that is why every substantial C program is ridden with undefined behavior, overflows, memory leaks, race conditions....
"C succeeded because it was doing the right thing" -> If the right thing is giving rise to the software exploit industry and causing billions in damages..
I suggest you read Richard Gabriel's "Worse is Better" (http://www.jwz.org/doc/worse-is-better.html) and forget anything that Stepanov has to say on the matter. Either he is utterly clueless or a dangerous imbecile.
C has been in use for decades now. We only have to look at the facts, not listen to fallacies that various cretins feel the urge to proclaim.
Back in th 8 and 16 bit days of the home computer systems, I never felt the need to use C.
I could touch all the hardware using Assembly, Turbo Pascal.
Apple was using Object Pascal for system programming and a few friends were doing Amiga games with Amos.
Sure I did learn C at high school, but it only became a requirement for the university work that had to run on UNIX systems.
Calling conventions can be a bottleneck but this is not unique to C, any other compiled language allowing separate compilation units to be linked together has the same issue. Techniques like LTCG can avoid it.
But given that C is almost always at or near the top of benchmarks both for size and speed, maybe it's not that much of a bottleneck after all.
I hadn't heard of this -- is it still true now that segmented stacks are gone? If so, why does C need a separate stack?
tldr; C's stack is interleaved with Rust's stack.
It seems that one can choose to both use runtime assertions and static proofs of correctness proofs in Spark. But I don't know if that extends to statically ensuring memory safety in low-level code.
On the early days it was deemed too complex to implement, although I would say C++ became even more complex.
The companies that sold Ada compilers had customers with deep pockets, so Ada compilers were too expensive and required worksations to be used properly.
When affordable Ada compilers became available, not many cared about it.
Nowadays it has found its place where human lifes are at risk. Many avionic systems, train control systems, hospital devices are coded in Ada.
I do attend FOSDEM regularly and also get the feeling its use is increasing in Europe thanks to the security exploits in languages tainted by C compatibility.
Plus, it's not like semicolons are any big deal. If they are in a language, you add them and move on. Dead simple to add, minimal noise, instantly familiar to most C-derivative programmers. The only comminity that regularly complaints about them are hipster (for lack of a better term) javascript programmers.
But your post would still be decently readable if you just made a new paragraph for each sentence ;)
I like Python. I think Python's whitespace-based blocking is interesting. But it is not the most interesting part of the language, and anybody who obsesses over that kind of stuff is just bikeshedding.
Developers always seem to be complaining about having to type this or that, when they should be much more concerned with what they have to read!
Really, I couldn't care less about typing a few extra characters. I can type pretty fast anyway and usually spend more time thinking than I do typing. What I do care about is being able to tell exactly what some code does by glancing at it, without worrying too much about whether someone wrote = instead of ==. Rust seems to be setting itself up for loads of those kinds of errors by trying to make the syntax terse at the expense of making it readable.
I have high hopes for Rust and will reserve judgement until it is stable; from what I've seen it's still in a high state of flux. However for now, I much prefer the KISS approach taken by Go, in spite of a couple of things missing from the language (that will probably get there in the end).
Or maybe they find terser syntax to be more readable. If verboseness was always more readable, if not necessarily more "writable", than terseness, then no one would have a problem with Java since it has good IDE support, including autocomplete and generation of boilerplate code. But it turns out that it's not just a matter of being lazy typists.
I'm pretty comfortable writing complex regular expressions, but I don't know many other developers that are and I certainly don't like trying to grok anything longer than about 10-15 characters that I didn't write in the preceding 15 minutes. This is a perfect example of where terseness is fine for simple problems, but it does not scale.
I already mentioned Go because it has a pretty terse syntax, but one that is very carefully optimised for readability. One of the things the designers of Go are careful about is not adding too many operators, keywords, or usage rules to the language, which keeps everything nice and simple.
Rust OTOH seems to have a metric boatload of operators that work in different ways depending on the specific context. That's a recipe for a ton of cognitive load, which isn't helpful for writing code, but reading suffers even more.
I don't yet know enough about Rust to say whether this is as bad as operator overloading in C++. As I said, I will reserve judgement until 1.0 because it's entirely possible things will change drastically before then.
What operators does Rust have that Go does not? I believe Rust has no more operators than Go.
You're right.
That is two separate levels. When you first glance a page of code the first time in 1 second, you should tell what the structure of the program is. How many blocks (for/while/if) it has. How many functions, how big they. You haven't yet had time to read each individual character. That is one level.
Here variable and ambiguous indentation rules get in the way. If there is non-uniform, non-standard indentation then you have to start reading individual lines in detail. If there is standard indentation then it doesn't even matter about little commas and semicolons, it is a level higher than that.
Then past that it is about individual functions, classes, modules, and what have you. Then it becomes about == vs = or . vs , and so on. If there is ambiguity there it could be harder. Like in Python I added a , at the end of a some variable. So that turns it into a tuple. And it resulted in a strange exception down the line. In C the = vs == is notorious. But there are others. None of this make the task impossible, but just slightly harder. The problem is that if this is done many many time over the course of a lifetime of a piece of code. Maybe it take 10 extra seconds for reader to understand the code, that multiplied by thousands of times will add up.
Prefer "x = y + z" to "ADD y TO z GIVING x".
Prefer "[[NSThread alloc] initWithTarget:t selector:s object:o]" to "pthread_create(&p, &a, &f)".
That the last value in a block is the value of the expression is something that makes a lot of code much more readable, as it frequently obviates the need for additional temporary variables.
I was sceptical until I actually used it. Now I’m converted; though I still also like Python’s pure statement-oriented approach in various ways, I prefer Rust’s model.
https://github.com/rust-lang/rust/blob/master/src/librustc/m...
Personally I think most other things in Rust (namespaces, constant as_slice() unwrap() etc.) are currently too verbose, although I'd say that also impedes readability.
I do agree with the rest of your first paragraph though.
We've written hundreds of thousands of lines of Rust and this has never been an issue. The typechecker will catch any misuses of semicolons.
> What I do care about is being able to tell exactly what some code does by glancing at it, without worrying too much about whether someone wrote = instead of ==.
The typechecker will catch misuses of = versus == as well. So assuming that the code you're looking at passes the typechecker, you don't have to mentally distinguish between = and ==.
> However for now, I much prefer the KISS approach taken by Go, in spite of a couple of things missing from the language (that will probably get there in the end).
They're different languages. Go does not have memory safety without garbage collection, and will never have it while remaining backwards compatible. But that was an explicit design goal of Rust. That is why Rust has the lifetime and unique pointer support, which allows Rust to support safe low-level programming in a way that wasn't possible before.
> This is like optimising for the least readable code possible. I imagine it's going to be a nightmare in practise.
I have pushed a reasonable amount of code to the rust repo, my own libraries, and some to servo, and it has never been an issue, in fact it has been the opposite (proof: https://www.ohloh.net/accounts/bjz and https://github.com/bjz/).
> What I do care about is being able to tell exactly what some code does by glancing at it, without worrying too much about whether someone wrote = instead of ==.
Rust solves this by having assignment expressions always returning `()` from assignment expressions and not having implicit conversions. The issues with `==` vs `=` completely vanish.