Things Rust shipped without
graydon2.dreamwidth.org
graydon2.dreamwidth.org
I haven't done this for a while, but once upon a graduate program I wrote a compiler from a made-up-language (MUP) to C. MUP had some strange control structures, and if C did not have "goto", it would have been a lot more difficult to implement those structures. Since then, I have always thought languages should have a "goto" statement that human-written code is not allowed to use. :)
(I don't recall whether the borrow checker actually depends on the control flow being reducible, but some of the possible future improvements we've talked about with "non-lexical lifetimes" definitely do. There are also nasty interactions between RAII and irreducible control flow…)
The benefits are enough that data structure and algorithm designs in the JVM compiler world often take advantage of assuming reducible control flow even though Java bytecode can express irreducible CFGs. Instead, such programs are left to the interpreter.
Most problems on reducible graphs are not easier than for general graphs, because you can always compute a loop forest (for generalized loops, not natural loops with a single entry point) and just consider a derived acyclic graph.
> you can always compute a loop forest
It's easier to not have to.
https://pp.info.uni-karlsruhe.de/uploads/publikationen/braun...
In the case of irreducible control-flow it may produce redundant cyclic phis, but those can be eliminated with a pretty simple post-pass.
https://github.com/rust-lang/rfcs/pull/396
It does have the effect of picking a single dominating entry point for a loop with multiple entry points, but if you want a single-entry region that's necessarily true. Multiple-entry regions are probably possible, but they could be counterintuitive and produce even stranger error messages than the current region system.
The borrow checker is based on dataflow analyses that should work on arbitrary CFGs, assuming an appropriately generalized notion of region.
True, but only in node-splitting approaches. If you use a label threading variable (as emscripten's relooper does), there is a guaranteed reasonable limit on code size increase (at the cost of performance).
"Just recently, however, Hoare has shown that there is, in fact, a rather simple way to give an axiomatic definition of go to statements; indeed, he wishes quite frankly that it hadn't been quite so simple....
"Informally, α(L) represents the desired state of affairs at label L; this definition says essentially that a program is correct if α(L) holds at L and before all "go to L" statements, and that control never "falls through" a go to statement to the following text. Stating the assertions α(L) is analogous to formulating loop invariants. Thus, it is not difficult to deal formally with tortuous program structure if it turns out to be necessary; all we need to know is the "meaning" of each label."
In C it would still have use in the implementation of orderly error handling -- the pattern where you hand-implement exception handling in C by putting an on_error: label at the end of the function that is goto'd on error. The addition of some orderly construct for this in C would eliminate that case, leaving no real role for goto there either.
In 20 years of using C/C++, I've found one use for "goto" that's hard to substitute: simulating coroutines that yield to an outer context. (Similar to C# yield return).
A "break label" gets you out of a loop. I needed to use "goto" to jump back into the middle of a loop to resume where the "coroutine" previously left off. The keywords "break/setjmp/longjmp" wouldn't have been substitutes for this particular use case.
Here's the base code without gotos:
typedef enum { ADD, MUL, ..., END } opcode;
void run() {
opcode ins;
while (1) {
ins = fetch_next_inst();
switch (ins) {
case ADD:
perform_addition();
break;
case MUL:
perform_multiplication();
break;
...
case END:
wrap_up();
return;
}
}
}
You have 3 jumps on each loop. From the break to the end of the loop, then from the end to the top, and one from the switch to the right case. The first one might be optimized away, but let's remove it explicitly. typedef enum { ADD, MUL, ..., END } opcode;
void run() {
opcode ins;
start:
ins = fetch_next_inst();
switch (ins) {
case ADD:
perform_addition();
goto start;
case MUL:
perform_multiplication();
goto start;
...
case END:
wrap_up();
return;
}
}
Assuming a non lousy compiler, we haven't improved anything yet. But now the fun starts. We can go down to one jump for each iteration. typedef enum { ADD, MUL, ..., END } opcode;
#define NEXT() \
do { \
ins = fetch_next_inst(); \
goto *jump_table[ins]; \
} while(0)
void run() {
opcode ins;
static void *jump_table[] = { &&add_l, &&mul_l, ..., &&end_l };
NEXT();
add_l:
perform_addition();
NEXT();
mul_l:
perform_multiplication();
NEXT();
...
end_l:
wrap_up();
return;
}
Voila! a single jump every time around. Now, depending on what kind of architecture you're running on, the size of the cache, etc, this may or may not be faster.Granted, this is not the kind of code you write everyday. But sometimes speed matters, and good luck writing this without gotos.
But assuming a sufficiently smart compiler when you depend on your code being fast tends to cause issues. Especially if you expect it to run on a bunch of platforms with compilers of varying qualities. Last time I saw code last this, we certainly couldn't rely on compilers being particularly smart, and we did care about speed.
Do you have a pointer to the relevant part of OCaml's source? The whole thing is probably worth a read, but don't think I can set aside enough time for that soon, and it could take me a while to find the right part.
See the "Next" macro definition.
To be fair, I've seen code like this exactly once.
static void* address = &address;
This is helpful for things that take void* context pointers that you want to match with logical equality.
In my practice these exceptions are ubiquitous. Multiple tiny interpreted DSLs (which cannot be compiled more efficiently for the latency reasons), efficient protocols, all that stuff. System programming, in other words, and it's exactly the stuff Rust was supposed to be designed for.
> But it's definitely not the rule and the speed difference between the one and the other is so small
2-4 times difference is not "small" in system-level programming.
> Even the inner loop of an interpreter usually calls routines
These are too high level interpreters for the too dynamic languages, they're beyond any hope in terms of performance, by design. I'm talking about some simpler and better designed things (like OCaml, for example).
In your example:
void (*[10])() function_table = {
&perform_addition,
&perform_multiplication,
...
}
while(1) {
ins = fetch_next_inst();
*(function_table[ins])();
}
Edit: Seeing it written like this now, it's clear you could save yet another jump by defining a macro for what's in the while loop and putting it at the end of every function call.If you were to inline all of the perform functions in the GOTO version vs putting a macro at the end of each function, you're right that there's some function overhead, but I think it's as small as a single instruction to put the return address on the stack. Maybe that would be optimized away, maybe not.
To your point: My argument isn't that you can't do the GOTO version. With optimizations, it's essentially identical. My point is, that is a lot more hard-to-grok code to maintain for something that can be achieved in a simpler way.
Sometimes this can be optimized away, but not always.
Over time it tends to evolve into ever messier and harder to understand versions of the initial run after which a future maintainer will end up losing sleep and or hair chasing some production bug.
Consider this (slower!) much clearer alternative, and consider too that if you need to eliminate one goto for speed reasons that you're most likely doing something wrong:
typedef enum { ADD, MUL, ..., END } opcode;
int process_instruction(opcode ins) {
switch (ins) {
case ADD:
perform_addition();
break;
case MUL:
perform_multiplication();
break;
...
case END:
return FALSE;
}
return TRUE;
}
void run() {
do {
} while (process_instruction(fetch_next_instruction());
wrap_up();
}
That's a whole function call overhead (but you can eliminate that with an 'inline'), it's easier to test and much easier to follow what it does.It would be interesting to see what the actual difference is in speed when comparing those two versions, I suspect that the contents of 'perform_addition' and 'perform_multiplication' are going to be the key here, not whether or not the loop uses a short-cut or an instruction (or two) less. Oh, and you could have eliminated that 'ins' variable.
In summary, the assembly of the goto version is much shorter and there's a 10% speedup over SPEC JVM98. The optimization works out because: a) instruction bodies are often short; b) instruction dispatch is the hottest code path of any interpreter; c) it's orthogonal to all or almost all other optimizations.
Exactly. And stop-gap measures are fine if and when they're followed up by a proper solution. Unfortunately stop-gap measures tend to be a lot more permanent than originally intended in practice.
I like the way this is handled in Go with the defer keyword: https://blog.golang.org/defer-panic-and-recover
This construct gives you most of the power of C++ RAII without the overhead. Except that you can't use it to cleanup resources after exiting an anonymous block—it strictly defers to function exit.
In all other cases (which is almost all usages), it is semantically equivalent to RAII, so the language doesn't force any runtime overhead. The only difference is that the compiler is less mature than an average C++ compiler, but this is an implementation problem, not a design problem.
RAII has no overhead because it is a pattern designed within the context of a zero-overhead language. defer allows you to implement a superset of cases that RAII handles, including those with runtime overhead.
I hadn't considered the type of overhead you're talking about, which is admittedly more important for the kind of software that tends to get written in C.
But it leaves out the most useful part: in C++ releasing the resources is completely implicit:
{
std::ifstream in(fn);
// Do something with 'in'
} // Released
while with Go defer, you have to call explicitly for the method that you want to call. If you forget this, you still have a resource leak.Sure, it's an improvement over languages where you have to cover every possible scope exit, but it definitely doesn't give you 'most of the power' of C++ RAII.
static int foo_impl(int arg, int **resource1, int **resource2) {
*resource1 = (int *)malloc(sizeof(**resource1));
if (*resource1 == NULL) return EXIT_FAILURE;
/* Do something with resource1 (omitted)... */
*resource2 = (int *)malloc(sizeof(**resource2));
if (*resource2 == NULL) return EXIT_FAILURE;
/* Do something with resource1 and resource2 (e.g.,
stick them in a global hash table as key/value pairs,
or whatever, omitted)... */
return EXIT_SUCCESS;
}
int foo(int arg) {
int *resource1;
int *resource2;
int ret;
resource1 = resource2 = NULL;
ret = foo_impl(arg, &resource1, &resource2);
if (ret != EXIT_SUCCESS) {
free(resource2);
free(resource1);
}
return ret;
}
The advantages are that it makes it very explicit which resources need to be cleaned up (making it easier to review to be sure you haven't forgotten one, or forgotten to initialize one, because they're in a nice list), and it's harder to screw up the control flow. You can't get out of the function without passing by the clean-up code, and you can't accidentally fall into the clean-up code without explicitly returning an error.A way to fix it would be to return different failures for every allocation and to use a switch without a break to clean up, something like:
switch (ret) {
case EXIT_SUCCESS:
free(resource2);
case EXIT_FAILURE2:
free(resource1);
case EXIT_FAILURE1:
break;
}
EDIT: added potential fix to the code. Sorry if there is any bug in the fix, my C is a little bit ... rusty.EDIT2: fixed a bug in the fix. C is hard, let's go shopping.
EDIT3: the explanation for the bug was also wrong. It is not a memory leak for but freeing a null pointer. Fixed too.
Although the concern of a memory leak is valid. I would have expected the resources to be free'd regardless of the result.
Seems to be a common misconception [1], thanks for pointing out, that makes the GGP code correct in all counts and my switch redundant.
Two things about goto and jmp instructions. Most people have no idea how badly they were abused back in old days. For instance to jump into the middle of a subroutine. Yay just saved 9!!! words of core memory!!! And most people forget that old computer scientists were obsessed with creating grammars that you could write formal proofs for.
It's an interesting idea.
The problem I see is that users can get around it by wrapping goto in the simplest possible macro. Then the idea has backfired: we now have goto under as many names as there are goto-using programmers. :)
I think goto should be in almost every language. It's one of the most primitive instructions, why shouldn't it be available when needed? Yes, it can be misused, just like any other feature in the language, but it can also be used to great benefit. State machines, for example, can make good use of gotos.
Hell, look at any unix library your machine uses every day. The ones you don't even think about. Start with libncurses: I promise there's a few gotos where needed.
$ find ~/ncurses-5.9 -name '*.c' -print0 | xargs -0 grep goto | wc -l
70
Raise your hand if you plan to stop using ncurses because of how opposed to how "harmful" goto statements are.Oh, here's tmux, if you're interested (one of the most beautifully written C programs): https://github.com/tmux/tmux/search?utf8=%E2%9C%93&q=goto
I see what you're saying, though I don't find this a compelling argument. By this logic, why not allow direct register access in all languages?
Just because something is a primitive operation doesn't mean you want to include it, especially not if it's more difficult to enforce guarantees your language would like to make about valid programs.
What if I just avoid contributing to the ncurses codebase? I've used plenty of useful tools with absolutely horrific codebases that I'd never want to touch in a million years. Not sure if ncurses is one of them.
The whole "it gets used, ergo it must be a good idea" argument doesn't hold much traction with me - even if I think using it when in C, to enforce single exit style, to work around the lack of RAII constructs, goto is the lesser evil.
I started with GOTO in BASIC. I used it a lot. It structured my initial reasoning about control flow. Despite this, in the past few years, I've used a naked goto maybe once or twice, and in all cases later rewrote it without the goto, which in my opinion increased it's readability. (I generally always have the option of C++ over C, and choose it, rendering single exit style 'useless'.)
> Oh, here's tmux, if you're interested (one of the most beautifully written C programs): https://github.com/tmux/tmux/search?utf8=%E2%9C%93&q=goto
Most of those are single exit style gotos. Those that aren't, do cause some concern, despite being "one of the most beautifully written C programs", even to their original author from the looks of it:
if (errno == ENOMEM)
goto retry; /* possible infinite loop? */
For what it's worth, it seems unlikely to be an infinite loop, short of encountering a bug in sysctl, or another process/thread constantly adding data. I had to google the header path to find an appropriate manpage (i.e. not _sysctl, not sysctl the program) to figure this out...I've encountered worse edge cases before, however, and I'd really prefer my programs crash properly, instead of hanging when they do.
> State machines, for example, can make good use of gotos.
The performance complaints about an additional branch misprediction when using the "for(;;) switch(...)" style without gotos, is one of the few arguments that moves me, if only slightly. That seems like a case your standard optimizing compiler really should be able to handle, however. I'll assume they don't, as I'm too lazy to test if this is merely hearsay...
That said, I'll even use "goto case" in C# on occasion where I'd use case fallthrough in C++, if I'm feeling particularly lazy and don't want to turn things into proper method calls that can simply call each other just yet. I usually clean it up before I start to confuse myself.
It's not something I'd miss if it were gone, however. It's something I use only rarely, and only as a crutch to stave off cleanup. Not exactly a ringing endorsement.
EDIT: Code formatting, proper insertion of subject...
Anyways, exceptions are sort of gotos or at least they can behave that way.
It does, yes, although many Lisps got a pretty limited backend for this sort of things. My preferred metaprogramming environment must have a fallthrough mode for allowing generating low-level code where Lisp semantics is not sufficient. Rust seems like a very good target platform in this sense, so the lack of goto really hurts.
> It doesn't have goto.
Of course they do (tagbody in Common Lisp, for example). And when they don't, it's often relatively easy to add one.
I was wrong, however it is much more limited than the goto everywhere from C, as where the labels are located is clearly defined in the tagbody and not in every possible statement.
I don't plan to stop using ncurses in particular, but one of the reasons I am supportive of Rust is that I would like to stop using all C software sometime in my lifetime.
Note that a union of two pointer-containing structs is safe as long as the pointers line up, having the same type and offset in each struct.
https://github.com/rust-lang/rfcs/pull/724
If you really need them, you can implement them in a library by declaring a suitably sized byte array and having methods that return casted pointers to it. This would be illegal in C due to strict aliasing, but IIRC Rust does not have a strict aliasing rule (because almost all pointers being the equivalent of 'restrict' drastically reduces the benefit).
There are issues with size_of not being a constant expression and such (at least in stable), but those are definitely going to be fixed.
The main advantage of using higher-level languages is that you can talk about manipulations of data. Goto doesn't manipulate data, it manipulates the machine. And if you want to _really_ manipulate the machine, you're going to want more than just goto.
>Raise your hand if you plan to stop using ncurses because of how opposed to how "harmful" goto statements are.
Goto statements are not harmful to computers or to code. Ncurses isn't what is harmed by goto, nor does "using ncurses" imply any interaction with goto at all.
Muhammad Ali was a good boxer. He was Muslim. Am I to conclude that one must be Muslim to be a good boxer, now?
Yes, ncurses is a good program. Yes it uses goto. But that doesn't imply that goto is necessary to it being a good program. Or that it is even a mildy efficient way of doing well. (Especially not when there is a whole universe of alternatives.)
If you're going to defend goto (and there are reasons to do this) you should probably do so without employing such blatant logical fallacies. It's irresponsible and detracts from your point.
All its use cases are better served by other language constructs.
Then again, C is a portable macro assembler.
I don't see that as a bad thing though, it's not great if you are trying to write Rust programs/libraries -right- now as sometimes you will have to jump through some hoops to only use stable Rust APIs but it will be worth it in the long term.
I'd also question why the module system is "braindead", of course.
I don't recall why incremental compilation requires a breaking change, maybe it doesn't. I shouldn't have looped it into that statement without being certain.
The module system _is_ braindead, though. Last time I brought this up, I made a proof of concept that compiled the exact same program two ways - one by using the module system, and one by running the code through the C preprocessor and literally #include-ing other rust files. The purpose of this PoC was to point out that Rust modules are functionally identical to #include-ing C files would be in C, which is a well known antipattern.
Rust is very close to C when you consider the guts of the toolchain. C has solved several problems with respect to linking and I feel like Rust could have taken several more hints from C. This is probably a consequence of the Rust devs inherently disliking C and wanting to distance themselves from it.
But they trivially aren't.
lib.rs:
mod foo;
mod bar;
foo.rs:
fn f() {}
bar.rs:
fn f() {}
No name conflict. foo::f and bar::f happily coexist.In C:
lib.c:
#include "foo.c"
#include "bar.c"
foo.c:
void f() {}
bar.c:
void f() {}
Name conflict; fails to compile.> Rust is very close to C when you consider the guts of the toolchain. C has solved several problems with respect to linking and I feel like Rust could have taken several more hints from C. This is probably a consequence of the Rust devs inherently disliking C and wanting to distance themselves from it.
No, it's that header files are are a big problem in C (DRY violation, hostile to code inlining, slow compilation) and it was felt that a real module system would be an improvement.
mod foo {
#include "..."
}
Which is barely enough to say that Rust is hugely different than just #include-ing C files.I was talking more about linking objects incrementally and the consequences of that design, rather than singing the praises of headers (though I do rather like headers). I understand C from the compiler's perspective as well, having written my own linker and assembler from scratch myself, and I really appreciate the elegance of the design.
>slow compilation
That's objectively untrue. It's much faster to compile with something like headers.
As far as inlining and DRY are concerned, back before Rust shipped I spoke with many Rust maintainers about solutions to all of these problems, but it was dismissed because "we're trying to ship". Maybe you shouldn't sail a boat when you need to replace the hull later?
Well, sure, if you want to get fancy enough you can make a module system out of #include. (Your code snippet isn't enough because it doesn't replicate privacy or imports.) But replicating something approximating Rust's module system with "#include plus other stuff" doesn't show that Rust's module system is "just #include".
> That's objectively untrue. It's much faster to compile with something like headers.
I don't think that's true once you have the proper incremental build setup. With headers, you have to parse large source files over and over again. With a proper module system, the compiler can use a more efficient binary database format (with an index).
For example, consider math.h. If your program is using one function (say, sin) from math.h, you have to parse all the prototypes in math.h. (And if you have inlined functions or templates in the header files, you have to parse those too!) But with a module system, the compiler can serialize an index of all the signatures of functions inside libmath.so, so the compiler can do a direct, O(1) hash table lookup for "sin".
We aren't there today, of course, since we don't have incremental compilation, and C is certainly simpler, but I think doing it right from the start will pay dividends down the road.
> As far as inlining and DRY are concerned, back before Rust shipped I spoke with many Rust maintainers about solutions to all of these problems, but it was dismissed because "we're trying to ship". Maybe you shouldn't sail a boat when you need to replace the hull later?
I don't see anything backwards-incompatible about incremental compilation, and I believe where we're going will end up better than header files when all is said and done.
Isn't this pretty much a solved problem with precompiled / pre tokenized headers? PTH are language / arch / compiler agnostic.
> Modules making it into C++17 is less likely.
However, they do say > That said, from a user’s perspective, I don’t think this is any reason to
> despair: the feature is still likely to become available to users in the
> 2017 timeframe, even if it’s in the form of a TS.Are you sure about that? One of the main advantages of coroutines/greenlets IMO is writing simple and straightforward blocking code (e.g. an echo server); without them, you either need to use threads (which are slower and much more heavyweight) orcallbacks or related constructs (async/await, futures, ...).
Even if it could, many real-world servers actually do non-trivial work in their threads, so the cost of spawning a thread is dwarfed by the actual work the thread ends up doing. There are serious drawbacks to M:N: complexity, fairness, problems in interoperability with the 1:1 world (including essentially unavoidable performance problems with the FFI), etc.
There is a Rust library called "mio" which provides a lot of the plumbing for such systems: https://github.com/carllerche/mio.
The path forward is likely going to involve adding a way to build cheap state machines (call them generators or async/await) with a clean syntax and giving mio hundreds of thousands of reusable instances.
I don't understand, and the link seems unclear. Perhaps a more direct question: I get a request X, and I need to consult a backend service to answer the request. Do I write synchronous code calling that backend? Or do I have some callback mechanism?
> ... generators or async/await
Ah. This perhaps answers my question. Both of these are essentially compiler-written callbacks.
If this is going to be like C#, then I presume there will be a thread-pool where user code will execute. It seems like a non-ideal story for concurrency. Users will have to take inordinate care not to call any blocking code; otherwise they will prevent one of the threads in the pool from doing useful work.
The downsides of going M:N are worse. The cgo-like FFI performance problems, for example, are killer for Rust's use case.
To be honest it seems to me like your explanation is an attempt to downplay just how nice fibers/coroutines are rather than acknowledge their utility in many existing languages.
Userspace is not equipped to make reasonable scheduling decisions that provide any significant performance advantage, and library/language runtime control of M thread register/stack contexts on top of N kernel threads plays absolute havoc with most operating system's standard libraries.
Go works around this by explicitly not calling into libc et al -- all system calls are issued directly. One big problem with that: directly invoking syscalls is supported on Linux, but NOT supported on OS X.
End result is that Go literally must rely on undefined behavior on any system that does not support direct issuing of syscalls.
From my brief review just now, what MS appears to be proposing for C++17 isn't coroutines in the traditional M:N threading sense, but rather, an explicit mechanism (with syntactical sugar) for capturing reachable variables in a lambda (without preserving the stack), and issuing a call to that magicked-up lambda later via promises.
This is interesting if you love the idea imperative mutable promise-based concurrency, but it's not likely to win you any performance gains, and it's useless in the extreme if imperative mutable promises aren't your cup of tea.
But since Rust is intended for runtime-less lower-level programming it doesn't make any sense here.
https://mail.mozilla.org/pipermail/rust-dev/2013-April/00355...
https://github.com/rust-lang/rust/issues/217
> I'm sorry to be saying all this, and it is with a heavy heart, but we tried and did not find a way to make the tradeoffs associated with them sum up to an argument for inclusion in rust.
> -Graydon
https://github.com/rust-lang/rfcs/issues/271 is a better link today.
return as foobar(x, n-1)
return in foobar(x, n-1)
override return foobar(x, n-1)
final return foobar(x, n-1)
It's a little surprising to see entirely new construct invented for something that is just a different implementation of `return`, after all, the end result of computation with tail-return is the same as with normal return, only performance differs.A safe stdlib - aborting on malloc failure is not safe.
Abort occurs on oom, that's right and also on double panic (most frequently found if a destructor panics during unwinding).
In C++, this would look like:
#include <iostream>
class B { public: int get_id() { return id; }; int id; };
class C : public B { };
int main() { C c; std::cout << c.get_id() << "\n"; return 0;}
Maybe I'm ignorant, and there's an easy way to do this. I was screwing around with Rust a few months before 1.0 and didn't see any mention of it in the docs, though. (Everything I saw required me to provide an impl for functions defined in an interface for every class that implemented that interface.)Rust provides this. E.g., the "talk" method in the "Animal" trait below: [0]
trait Animal {
// Static method signature; `Self` refers to the implementor type
fn new(name: &'static str) -> Self;
// Instance methods, only signatures
fn name(&self) -> &'static str;
fn noise(&self) -> &'static str;
// A trait can provide default method definitions
fn talk(&self) {
// These definitions can access other methods declared in the same
// trait
println!("{} says {}", self.name(), self.noise());
}
}
[0] From Rust By Example, http://rustbyexample.com/trait.htmlA question (In C++ syntax, as I'm not a Rust programmer): How would one call Animal::talk inside Dog::talk? The obvious thing (commenting out the println and adding Animal::talk(self) ) causes infinite recursion. The other vaguely obvious thing ( Animal.talk(self); ) is a syntax error, which makes sense.
Edit: I'm referring to the code at the linked Rust By Example page.
There is no way to call the overwritten `Animal::talk`.
Animal::talk(self);The default implementation, if overridden, does not exist for the given type.
pub fn super_talk<T:Animal>(this: &T) {
println!("{} says {}", this.name(), this.noise());
}
trait Animal {
//...
fn talk(&self) {
super_talk(self)
}
} let mut arr = [1, 2];
let a = &mut arr[0];
let b = &mut arr[0];
This code could be allowed: let mut arr = [1, 2];
let a = &mut arr[0];
let b = &mut arr[1];
And it gets Hard once variable indexes are involved. And since the whole point of using arrays is to get variable indexing, the Rust developers choose to only implement reborrowing for structs and tuples. fn main() {
let mut x = 10;
let y = &mut x;
*y = 11;
println!("{}", x);
}
This will complain error: cannot borrow `x` as immutable because it is also borrowed as mutable
This is because an `&mut` borrow is exclusive: while `y` is alive, we cannot use `x`. We can fix this by making a new scope for `y`: fn main() {
let mut x = 10;
{
let y = &mut x;
*y = 11;
}
println!("{}", x);
}
This works, and will print `11`.Non-lexical lifetimes would allow the compiler to demonstrate that these two things are the same, and allow the first one to compile with the behavior of the second.
It's an interesting tradeoff, because right now, the rules are very simple and conservative. Scope is fairly easy to reason about. Non-lexical lifetimes would make certain things easier, but also a bit harder to reason about, because the rules are more complex.
I want memory safety with as few hazzle as possible and that Rust doesn't understand the safety of the first example means hazzle. I expect the compiler to try to understand even if it's ’hard’. It doesn't need to understand everything but the mentioned situtation should be doable.
I'm not sure what you mean by a 'more detailed tradeoff.' You mean a more complicated example?
> I expect the compiler to try to understand even if it's ’hard’.
It's not a matter of difficulty, exactly, it's a matter of how easy it is to understand what the compiler is doing. Figuring out non-lexical scopes means that my mental model of what the compiler is doing is more difficult than it is right now, which may or may not be the right tradeoff. I would say that most people want non-lexical lifetimes/SEME regions to be implemented, though.
I might be wrong here but it always seemed to me that Rust was built to replace C/C++ in critical infrastructure like Firefox, Nginx, Redis, etc. Basically critical network dependent infrastructure.
With respect (because I understand the difficulties that come with time constraints/resourcing/building an efficient async API), currently I'm not sure how Rust expects itself to be a viable replacement (let alone the best replacement) language for any of those types of applications.
It's important to remember that IO is truly a library concern in Rust. 1.0 means the language is stable, but there's still tons of libraries to build on top of that language. Holding back the language itself for a library that, while important, is only needed for certain applications wouldn't make a whole lot of sense.
The argument might be similar to using benchmarks for a parallel web rendering engine (a Servo-like) in Go. It would probably have similar median latencies to Servo. And Go may even scale out better if tabs were sandboxed in goroutines instead of processes. However, ofcourse unlike Servo it would have a higher variance/p90/p99 latencies because of the GC. P.S. I'm not a big fan of Go I was just using it to try and bolster my argument that Servo's current benchmarks shouldn't be used as an argument against async io prioritization. Actually, side note, I would love to hear more flaws about building a Servo-like in Go.
https://www.reddit.com/r/rust/comments/1v2ptr/is_nonblocking...
The timing of the move away from green threads didn't really offer enough time to implement a stable async IO option before 1.0
Personally I'm starting to dislike index operations since they can panic (and are the shortest way to access), and I rather use the explicit Option based APIs. Though I'm not too worried about those, since a lint disallowing them shouldn't be too hard.
Seriously, nearly all data storage formats we use today have some kind of version number in them - why are we treating code as dumb text rather than interesting data?
I think rust is neat, but this is somewhat arrogant.
Rust is amazingly young. 10-20 years from now, when rust hopefully has bajillions of users, if this is still true, then you can say it. I mean, do you really believe that Rust won't have things that turn out to be warts from 1.0 it can't remove 10-20 years from now?
The point is that you can at least try to avoid the things you view as sad design choices--you're not necessarily stuck with them.
http://www.kylheku.com/cgit/lisp-snippets/tree/tail-recursio...
This also provides some facilities for doing cross-module tail recursion among top-level functions. Here, continuation to the next function is provided by wrapping the function call in a closure, performing a non-local exit which abandons stack frames up to a dispatch loop, which then invokes the closure.
Language creators won't endear themselves to me by ranting. The problem I have with Rust is _not_ the language itself.
It makes everything so much easier to read.
fmt-tools offer the amazing possibility that one programmer writes and reads the same code with different formatting than another programmer. I'd love to be able to set my formatting in my editor so that I see it how it's best for me but on saving or sharing code the formatting is reverted to the standard.
So, I think there should be a default style for rustfmt, but also support for other styles.
Even if I repeat myself: I think there should be one default formatting that is standardized and there should be the option to emit in other formats such that everyone can read in the individually preferred format.
With fmt you don't need to establish formatting rules on a project basis, anymore. Everybody can just configure their editor to format the code how they want it to look. That is why I think rustfmt should be compilable as a library, too.
Every language carries with it a culture. The culture of Go is one which allows for gofmt to define one true style and refuse to deviate, just as it allows the language designers to refrain from adding generics.
The culture of Rust is not like that, for several reasons. For one, the Rust community loves a good bikeshed. For two, the syntax of Rust is more complicated than Go, and includes situations (match statements & where clauses come to mind) in which people are just going to want to different things.
I know the advantages of one true style - everyone's heard the arguments - and there's a sane, median default as the official style guide, which will be rustfmt's default output. That seems like a good compromise.
Its testament to the fact that there are some people that want that; that's very different from that being what people in general want.
Good
fn inc(a: u32) -> u32 { a + 1 }
fn foo(a: u32, b: u32) -> u32 { let x = a + b; a * x }
Bad fn bar(...) -> bool {
let mut success = false;
let conn = getConnection();
...
if x > y {
return false;
} else if z < q {
success = false;
}
foo.barify(x, y);
...
success
}
It looks especially bad when the function has multiple early returns, and then the final return looks different."Nicer" is a subjective thing. BTW in the trivial case above one may judge this or that to be nicer, but in a large method, 'return func(a,b,c)' is obvious, whereas looking at 'func(a,b,c)' it is super non-obvious that the value is being returned.
Then "foo.map(|x| :x+1)" is nice. If that doesn't stand out enough for some people on a line of its own, make your editor render the unary colon line in a very bold color for you.
Maybe you are writing code for the Rust standard library, where following the core projects style recommendations would be important for consistency.
Maybe you don't want to start from scratch coming up with your own style, and want a decent starting point from which you can vary as your team figures out what does/doesn't work for them.
Lots of reasons to have a language project also provide default style recommendations.
On the one hand, it is an objective fact that it is considered, by the authors, poor style.
On the other hand, that it is "considered" anything is an explicit (not merely implicit) statement that it is subjective (and "poor style" -- or good style, for that matter is inherently subjective in any case.) So, characterizing that language as making it sound "like their recommendations are objective facts" seems quite bizarre.
Hanging braces style for C/C++ is rarely allowed in the coding standards I've had to use in the past for embedded and real-time stuff in the defence industry, because it can be a source of errors.
Aligned opening and closing braces are much more common (in line with ADA's style).
By default it just warns. That's not enforcing.
How about making it the opposite, '#![allow(default_style)]'? Or at least '#![allow(non_default_style)]'?
But I get your point. I think people would be open to that naming change if you file an issue.
Maybe I could at least try, yes. I mean... It's kind of silly, I know, but that would already change the way I look at Rust.
Given the frequency of manipulating large bodies of text in contemporary programming I do not agree it is best left to a library implementation.
Without ropes as core std average rust developers will do what they did in java, c++, objC and simply use and abuse std::strings in all cases including those where it will perform poorly.
https://en.wikipedia.org/wiki/Rope_(data_structure)
[pdf] http://www.cs.rit.edu/usr/local/pub/jeh/courses/QUARTERS/FP/...
union foo {
int x;
float y;
};
union foo bad;
bad.x = 100;
printf("%f\n", bad.y); // undefined behavior (though usually works)[1]: http://dbp-consulting.com/tutorials/StrictAliasing.html
> If the member used to read the contents of a union object is not the same as the member last used to store a value in the object, the appropriate part of the object representation of the value is reinterpreted as an object representation in the new type as described in 6.2.6 (a process sometimes called ‘‘type punning’’). This might be a trap representation
float half_r=r*0.5F;
union
{
float y;
int32_t i;
};
y=r;
i=0x5f375a86-(i>>1);
y=y*(1.5F-(half_r*y*y));
return y;But where I'd more likely use that optimization at this point is on Arm or ATmega processors. ARM doesn't seem to have an approximate inverse square root, based on a quick check, and ATmega are frequently still stuck with software floating point, so I'd hardly say that the optimization is dead.
fn approx_invsqrt(r : f32) -> f32
{
let y : f32 = unsafe {
let i : i32 = std::mem::transmute(r);
std::mem::transmute(0x5f375a86 - (i>>1))
};
return y*(1.5-(0.5*r*y*y));
}
fn main()
{
println!("approx_invsqrt(2.0) = {}", approx_invsqrt(2.0));
}
Result: approx_invsqrt(2.0) = 0.70693This is usually done to "get the byte representation of a float" or things like that.
Except... it's undefined behavior.
No longer true since C99.
for C99 see http://www.open-std.org/jtc1/sc22/wg14/www/docs/n1256.pdf
- random-access strings
- auto-increment operators
(Note that Rust _does_ have a ranged syntax here, which returns bytes, and is O(1))
Pre vs post increment just leads to a lot of confusion, and doesn't give you much over `x += 1`, in my opinion. Not sure what Graydon's reasoning is here.
It just reveals the inherent complexity in handling string data.
If you want to be able to slice at grapheme boundary n, you have to iterate through the string counting up to n to find out where it is. Might as well just provide an iterator over graphemes in the first place.
You can optimize slicing on codepoint boundaries by storing strings internally in a fixed-width representation, like Python does. But now you have to convert encodings all the time just to use UTF-8 for your I/O. And that still only gets you codepoints.
It's not a great tradeoff to do it that way. It wastes memory and CPU time, and randomly accessing characters of strings is just not something you actually need to do often enough to optimize for it.
what should be the result ? what is the result in C++ ? (undefined)
1. It could be a compiler error. The existence of certain operators doesn't mean they can be combined arbitrarily.
2. Sequence points could be defined differently, allowing for such a convoluted line. This would likely hurt performance, as the compiler would have fewer operations to reorder in each sequence point.
Typically, increment and decrement are just syntactic sugar. I don't mind if a language lacks them, but I can see why people like them. A convoluted example of UB in C++ is not a good argument against having them.
Also: In C++, overloaded operators are function calls, and function calls are sequence points. So if i is an object, the behavior is defined. That said, it would still be bad code. :)
Note: things like this are part of the potential advantages of static linking.
Knowing that `++i` is `{ i += 1; i }`, we have `i += { i += 1; i };` (which should compile already).
The nested assignment can be hoisted to obtain `i += 1; i += i;`.
That means Rust would compile `i += ++i;` as `i = (i + 1) * 2;`.
If those are things you do want to have, well, we didn't have a design for AIO that we were happy yet, as even the most advanced Rust library, mio, is still a work in progress. https://github.com/rust-lang/rfcs/issues/1081
As for async/await, https://github.com/rust-lang/rfcs/issues/388
In the Java space some libraries are moving to compile-time code generation instead of relying on reflection. It is a huge win, since a lot more can be checked beforehand. Dagger 2 is a good example how it can be beneficial. It provides dependency injection at compile-time, which will in turn check whether all dependencies are satisfied. I haven't seen this being done at compile-time before, but it is definitely a step up from reflection-based DI.
I'm not sure whether macros of Rust can provide such functionality, but the developers seem conservative when it comes to adding functionality. That seems like a good thing.
Yes, in C++ there is a similar library called ROOT which generates c++ files called "dictionaries" storing class information by running an executable over the files. I don't see (or understand) the downsides in providing the functionality for performing those steps at compile time though. The developers of ROOT are currently pushing for it to be included in C++17 (or beyond).
And thanks for the pointer to dagger, this looks like an interesting alternative.
I am not very familiar with rust, is what I was describing currently possible?
> It turns out that with branch predictions and the relative speed of CPU
> vs. memory changing over the past decade, loop unrolling is pretty much
> pointless.
http://lkml.iu.edu/hypermail/linux/kernel/0008.2/0171.html int remaining = length % 4;
switch (remaining) {
case 3: h ^= (data[(length & ~3) + 2] & 0xff) << 16;
case 2: h ^= (data[(length & ~3) + 1] & 0xff) << 8;
case 1: h ^= (data[length & ~3] & 0xff);
h *= m;
} match x {
0 => { /* do some stuff */; fall_through; },
1 => true,
_ => false
}
However, I don't know yet how useful it would be. I can't remember ever really needing it, so it would probably need a few practical examples before it became a reality but its an idea. int remaining = len % 4;
if (remaining)
{
do
{
remaining--;
h ^= (data[(length & ~3) + remaining] & 0xff) << (remaining * 8);
}
while(remaining);
h *= m;
}
Fall-through's interesting, but at the same time, as architectures have changed, has become less useful. Self-modifying code at one time was near vital, but has fallen by the wayside, fall-through is doing much the same.If you're truly sure you're better than the compiler, I'd imagine an assembly language implementation would be easier to understand than nested switch/while fall through madness.
That said. There's an implementation of Duff's device in the Wikipedia entry which (IMO is much easier to understand and maintain) that doesn't use fall-through that should be just fine to write in Rust.
I'll be wanting UTF-16 support. Going the other way matters too; if I'm on Windows I may need UCS-32 support.
The Rust std library had to pick a string encoding, and it picked UTF-8 (which is really the best Unicode encoding). The String type is platform neutral and always UTF-8.
However, it does provide an OsString type, which on windows is UTF-16. Maybe there is a library - and if not, one could be written - targeting Windows only, and implementing stronger UTF-16 string processing on the OsString type.
EDIT: To be clear, Rust's trait system makes this very easy to do. You just define all the methods you want OsString to have in a trait WindowsString, and implement it for OsString, even though OsString is a std library type. One of the great things about Rust is that its trivial to use the std library as shared "pivot" which various third party libraries extend according to your use case.
https://simonsapin.github.io/wtf-8/
But there is actually prior art here - Java's contribution to perverse Unicode encodings is called "Modified UTF-8" and encodes every UTF-16 surrogate code unit separately.
http://docs.oracle.com/javase/6/docs/api/java/io/DataInput.h...
Utf-32 is only fixed length if you don't care about diacritics, variation selectors, RTL languages, and others. Unicode is not one code point or one char/wchar/uint32 per glyph.
Few string libraries actually deal with grapheme clusters as the native underlying representation (Swift being a notable exception).
I mean good for Rust, but C is a pretty low bar. Does any language created in the last 20 years make those same mistakes?
Maybe we should be embarrassed as an industry that it took us 20 years to get there, yes.
'goto' should not be daemonized, we know better today.
LLVM used to have a C backend, yes, but that hasn't been maintained for a while now.
Rust is a meta-language with very powerful macros in it.
What is a macro? Macro, essentially, is a compiler. You can have some petty macros implementing tiny syntax sugar on top of your language, that's totally fine, but that's not their purpose.
The real metaprogramming kicks in when you implement very high-level eDSLs on top of your macros. And for this sort of things, having a proper compilation pipeline is a must. Many eDSLs may end up being represented as an SSA. E.g., I'm doing this with a Packrat eDSL - it pays to represent its intermediate language as an SSA for doing some high level optimisations before spitting out the host language code.
Sorry, I do not understand what are you talking about.
Of course it is the same host compiler. But, every macro itself is a small compiler. Sometimes, when your DSL is an elaborate, complex thing, the macro itself is a complicated, big compiler.
With all the bells and whistles of a big compiler - multiple stages, multiple intermediate representations, all that stuff. And SSA is one of the most efficient intermediate representations ever, suitable for a huge number of semantic classes of DSLs. So, making it harder to generate your same host language code out of an intermediate, in-macro representation does not serve any reason, it's plainly conuterproductive.
But, every macro itself is a small compiler.
I'm not following you here. What does that mean? The macro gets expanded at compile-time into 'normal' source code, by the compiler.Any marginally non-trivial macro is a compiler, by definition.
> The macro gets expanded at compile-time into 'normal' source code, by the compiler.
Macro expands a DSL inside it into the underlying host meta-language code. In other words, it compiles this DSL into the host language.
For example, a macro which defines a parser. Inside it is a BNF (or PEG), a very high level language. Macro must compile this language into Rust, and, since it is a very high level language, there are tons of optimisation opportunities that would be totally missed by the underlying Rust and LLVM because they lack this domain-specific knowledge.
Your macro will compile this source DSL in multiple stages (well, because this is the only sane approach to compilation anyway, read about the Nanopass framework for more details). First it will operate on an AST level, do some analysis, error reporting, inlining, may annotate the detected left recursive nodes and binary expression nodes. Then you'd lower it down into the trivial Packrat building blocks - still a tree. But now you notice that there is a lot of redundant reads from the input stream, and if you flatten this tree into an SSA you can optimise them all away.
Alas, after such an optimisation you'd have to promote it back to tree in order to generate Rust, because there is no `goto`.
And this is only one trivial example. In my practice there were dozens such DSLs. For example, same story is with an optimising WAM-based embedded Prolog DSL, multiple querying DSLs, tree walking DSLs (which are essential for implementing macros efficiently).
But either way, what about targeting LLVM IR makes it unable to replace C? Portability?
(But you're right, "replacement for all of C's use cases" was strictly incorrect. I'd argue that this is a bad use of C, unlike e.g. kernels or bootloaders, but that's a matter of opinion.)
However in those 20 years we saw the rise of UNIX in the industry and with it C and C++, to the point those languages lost mindshare and are only known to those that experienced the IT world before they got widespread.
So new generations have to re-learn system programming without C and C++ way of doing things is possible.