‘~’ is being removed from Rust
github.com
github.com
A lot of people respond as if he made the following two claims:
1) Recursive data structures are rare
2) Singly linked lists are useless
When in fact he made these two claims:
1) definitions of recursive data structures are not prevalent in code
2) unshared singly-linked lists are not useful
#1 is true since you define a recursive data structure only once, no matter how often you use it.
#2 is slightly less obvious. One of the big advantages of a singly linked-list over other structures is that you can share tails with multiple lists. With unique ownership that advantage goes away. (There is still the use-case of loaning out tails, but that is a fairly weak argument).
Another advantage of a singly-linked list is ease of implementation; that is negated when your standard library provides dynamic arrays and stacks already baked in.
That said, I don't think it will be. It's only really useful in function signatures, since types are inferred everywhere else, and `~` is harder to search for, as well as not on all keyboards. Plus, in Rust, you should always use stack allocation unless you absolutely have to, and so making heap allocation a little more weighty syntax-wise fits with the goals of the language. `~` -> `box` isn't that much heavier, either.
1: One way it's a superset: previously, only `~` supported 'placement new', and now, all pointers can with `box`. `box` also allows you to specify which allocator you'd like to use.
Which keyboards do not have a '~' key? How do those people type a path relative to their home directory? Use $HOME/foo/bar all the time?
It's not difficult to type (Alt+¨ followed by a space), but you can't find it by looking at the keyboard.
Vertical bar/pipe isn't shown on most European Macbook keyboards, nor is backslash in fact, so I guess backslash has to be removed too.
I'm just saying it's not a great argument.
Most euro keyboards need at least AltGr + key, possibly 2 keys (because ~ is a dead key, so a space is needed to insert the character itself)
> How do those people type a path relative to their home directory? Use $HOME/foo/bar all the time?
It's common to `cd; cd foo/bar`
The keyboard situation doesn't seem to be as bad as for the degree sign °, which doesn't seem to be typeable on an English keyboard at all.
French Canadian: http://www3.uakron.edu/modlang/french/images/kbd4.gif
European French: http://www.fentek-ind.com/images/frenchkeytop.gif
The situation with ~ is simpler; it's used by plenty of programming systems today, and any programmer who uses it and can't type it as easily as on a default US keyboard (i.e. with one modifier+key press) is not using their tools effectively.
(To anyone who feels the urge to educate me on the Fable of the Keys, remember that typing speed is not the only worthy criterion. Comfort, RSI, and ease of learning, matter much more.)
With my typing style, the former is a complex press-release sequence with the left pinky and ring finger, whereas the latter is a single chord executable with the right thumb and ring finger.
let x = box (GC) 5;
if my memory serves. The argument to `box` defaults to HEAP, so these are the same: let x = box (HEAP) 5;
let x = box 5;
And of course, boxing an integer is silly. This is just for illustration.The @ for managed boxes has been removed from the language for a while now. It's not technically gone because there's one or two instances in the compiler, IIRC, but it should be gone Real Soon Now.
It's GC in full uppercase FWIW, it's a static marker[0] not the type name. And the syntax already works in 0.10 (although funnily enough only HEAP and GC exist, no e.g. RC)
> The argument to `box` defaults to HEAP, so these are the same:
Yep[1].
[0] http://static.rust-lang.org/doc/0.10/std/gc/index.html
[1] http://static.rust-lang.org/doc/0.10/std/owned/index.html
Unique pointers are one of Rust's newest, shiniest features, so they do get talked about a lot. Probably disproportionally so.
Honestly, it's probably that the docs are still coming together. The tutorial is notoriously poor, Mozilla has hired someone to work on it, but the fixes haven't surfaced yet.
Anyway I'm very excited about Rust, really looking forward to 1.0.
And thanks for the explanations and the tutorial !
I get the arguments, and they're logically sound, but in the end it's beginnning to read like C++03.
On the one hand, the Rust community's generalising a lot of stuff, which sounds great, but then adding a bunch of macros (I know, this ain't yo daddy's C macros) to replace first-class citizens ([1, 2, 3] => vec!(1, 2, 3) for example), and I worry how IDEs will be able to handle macros from an autocomplete/intellisense standpoint, not just checking/displaying type information when using them, but also writing new ones.
I guess it's a language primarily by and for text editor users, and I'm beyond that point these days.
As I said, I started following the language daily a year ago, and a couple of months ago I was thinking of jumping in around this point in time, but right now I feel like I want to wait another year to see if the syntax has become more readily intelligible.
I still think it's a fantastic effort, and kudos for all the hard work.
What do you mean?
EDIT: Imagine a programming language designed around an IDE. Fun thought. Not sure whether it's a useful thought, but it's fun to think about.
A trivial example which is not related to rust but I think serves as a decent example, you can assume the IDE will color/italicize/etc member variables to get rid of _var __var var__ nonsense. You can also use slightly less verbose terms even if they are less informative if you assume a tooltip with a concise and complete explanation pops up.
The unnecessarily complex design of many Java libraries and frameworks is generally what pushed developers toward using an IDE, not the language itself.
What are your thoughts regarding IDEs vs editors? It's an important topic.
I do all of my javascript coding in vim and don't miss IDEs so much, but that's because I mostly only work on my own javascript code. For large, multi-person projects I massively prefer working inside an IDE.
> uses and instead of &&, or instead of ||, list instead of [] and so on. Most people find it far easier to read and search for these than the sigil alternatives. The avoidance of more than one way to do the same thing (like including both ~T and Box<T>) is another reason why people find Python easy to read.
Python is a great language for beginners and people who get into programming through nontraditional channels without having much of a symbol-pushing background; presumably, being able to parse symbolic expressions naturally is a skill that needs to be learned, and perhaps even needs to be learned early in life to wield effectively. If you don't have the ability, classical systems languages will look like malicious ASCII soup and a language whose code you can actually read like a paragraph of text will be long-overdue respite from this misery; the same perhaps even applies to other non-verbose shorthand (as an anecdote, one of my graduate program's logic-heavy classes is currently suffering from a student who more or less constantly interrupts the lecture to get all the mathematical notation on the board read out to him in natural language). The latest startup/trendy tech boom has made a disproportionate number of people of this type gravitate towards programming, many of them coming from the wider blogosphere/"social internet"; as a result, their opinions regarding what is good or bad and legible or illegible have come to entirely dominate the airwaves. It is all to easy to forget that there is a barely visible core of often significantly more prolific "native" programmers hiding in newsgroups/IRC channels/mailing lists and more often than not having very different measures regarding what constitutes good language design. When Rust came about, its design held the promise to finally deliver something to this group that might offer a way out of the C monastery; having it reorient itself to appeal to more popular sensibilities at this point is bound to cause a lot of hard feelings.
Personally, I always found the flat monotonicity of Python code to be rather unpleasant to read and write. Having what some people like to denounce as "line noise" actually enables a very useful visual contrast between "structural" (parentheses, scope boundaries, operators) components of the code and largely user-defined "names". Looking at a piece of code like "if(vals[1]==1 && vals==[1,2,3]) { ... }" allows you to break up the code visually before having read and echoed to yourself even a single word of natural language; this can't be said of a (hypothetical) example like "if vals.at(1) equals 1 and vals equals list(1,2,3):". Adding to that the general lack of code layout flexibility and the deliberate lack of alternatives to express a given pattern, the resulting picture is that the way an experienced programmer would usually gather a slew of information from diagonally skimming over a piece of code (segmenting all of it using the visually distinct "line noise", gathering when and how data is accessed, inferring the original writer's priorities from how the code was spaced out and what constructs were chosen, spotting the names of external dependencies...) is severely crippled.
Syntax is a very small part of the overall design. Rust won't stop being Rust just by removing a few confusing sigils.
It's ~[T] that is now Vec<T>, which was a good change in emphasis (the former is still valid, though it will become Box<[T]> with this RFC).
Besides the implementation differences, I think readability is improved when `Vec<T>` is written.
As a concisionist, this change stinks. I appreciate that the "not on my keyboard" argument is a big deal, but I'd argue that just merits a change of sigil, not adding verbosity.
Designing a language to be prolix just makes code more laborious to understand.
There's two side to it: it makes creating unique pointers slightly harder, but at the same time it makes (unnecessary) overuse of unique pointers slightly harder/less likely, and that looks to be a concern of the core team.
> Designing a language to be prolix just makes code more laborious to understand.
Does it make the code more laborious to understand though? The concept becomes easier to search for and it's easier to talk about it (both because box/boxing is a term of art, and because "box" is a single syllable).
(philosophy) A language should support concision to allow those skilled in the art to use their time effectively.
~Vec<~Vec<~Vec[~T]>> (not 100% realistic example)
becomes Box< Vec< Box< Vec< Box< Vec[Box<T>]>>>>>
which really looks weird. It's like you have tire marks over your code.In other words, one would write
Vec<Vec<Vec<T>>>
That might be considered ugly, but it's not ugliness caused by the `~` change.Hardly. As thestinger argued so eloquently, legitimate use of the ~ are few and far between: typically, recursive data structure definitions, which are rare even when you go Recursive Crazy. (Functional programming would even avoid unique pointers, since they prevent sharing.)
From an information theoretic perspective, giving a special syntax for such an uncommon use doesn't make sense. If you want your language to be concise, you want the short-cuts to be used for frequent cases. Before we give a special syntax to unique pointers, we should address everything that's more often used. That would turn Rust into APL.
Since we don't want such a combinatorial explosion of special symbols, we're back to the good old Huffman encoding: short cuts for the frequent use cases, more verbosity for the rare ones.
Wait, don't you need it to instantiate the corresponding recursive data structures? As in, a typical cons-cell list would be created in a fashion like Cons(1, ~Cons(2, ~Cons(3, ~Nil))).
Exposing unshared nodes like that does no good. Just write a cons function that embed the unique pointer, that will get rid of the tilde, and make for an even better syntax than what you just showed.
Now 2 things.
First, the tiled denotes unique pointers. Your example was a singly linked list. Singly linked lists have 2 important characteristics: adding or removing the first element is O(1), and the tails of those lists can be shared. With unique pointers, you can't share. By design. If you want several references to a "node", you need to use shared pointers, whose allocation is manage by reference counting or garbage collection. So you have a data structure whose only advantage is O(1) insertion and removal… of the head. That's not very useful to begin with, considering the absurd amount of heap allocation you need to do. Other data structures fare better (vectors, ring buffers…).
Second, your example supposes the existence of a `cons()` function to begin with. If you really want to use unique pointers, you should write a function that accepts values, and wraps them in a pointer instead. That way, you can write `cons('A', cons('B', cons('C', empty)))`. There, no more pesky tiled: they have been factored in the definition of `cons()`. Less repeating yourself for the win.
> are you saying that for every self-referential constructor of an algebraic datatype, the correct thing to do is to write a boilerplate function that simply wraps the constructor and a call to whatever the boxing operation winds up being...?
Not quite. I'm saying the constructor itself should take care of the boxing operation. When devising a data structure, you generally know what it will be used for. It's memory allocation scheme should be a part of it. Hidden, if possible. For instance, unique pointers have value semantics. As such, they're an implementation detail. Leave them out of the interface. If it turns out you didn't need them after all, you can scrap them without breaking outside code.
Hmm, I guess that makes three…
From what I have read, there is a general mechanism for pattern matching: let the custom data structure implement a "pattern match" method or something, which is then implicitly called with the pattern matching syntax. The explicit box type would be no different.
My guess is, parse-trees would be just as easy to match.
Punctuation and other built-in language features give me a baseline for what I can understand and what I can always rely on. If I see ~T, even if I don't know what it means immediately, I still know that it's part of the core language and will get me a step closer to useful understanding; here is a thing I can probably use to solve some problems. Moving as much of the language into a stdlib as possible is a worthy goal that makes for a "cleaner" and more "elegant" language (whatever those mean), but it vastly reduces this effect.
The problem is that all of the core syntax and builtins are probably good to know, but not everything in the stdlib is useful. At least half of the stdlib for any given language tends to be worthless junk. Do you do Python? Did you know about the formatter module? popen2? asynchat? sunau? Probably not. But you almost certainly recognize 100% of Python's syntax and at least 90% of its builtins.
Contrast with Perl, which (besides having a ridiculous amount of built-in syntax) relegates such basics as OO support into the standard library. Subclassing is done by using a library, and there are even two distinct libraries for doing it in the stdlib! So to understand Perl code, you have to understand a good chunk of its stdlib as well, but not all of it because a lot of it is weird obscure junk, but there's nothing telling you which bit is important because it's just a thing everyone knows. (And the reliance on third-party libraries to fill in gaps in the language makes this far worse.)
One of C++'s major offputting properties is that everything is in a bloody library. (The other is that the features it does have all interact in obtuse ways.) You can learn what all of C++'s syntax does, and still be unable to make sense of real-world C++ code. I'd be pretty sad to see that happen to Rust.
> Rust has strived to remove non-orthogonal features for a long time, and the language has been getting steadily smaller and simpler.
Hereby avoiding the one mistake that made C++ so horrible: lots and lots of non-orthogonal features. It's not just a matter of being cleaner or more elegant. It's a matter of having less to learn.
As for half the standard library being worthless junk… Rust probably won't be doing that mistake. If they simplify the language, they are likely to simplify the standard library as well, which means cutting the worthless junk out. Plus, it will likely be easy to use whatever is most useful: it will be all over the place in tutorials, manuals, examples, and of course actual code.
> One of C++'s major offputting properties is that everything is in a bloody library.
On the contrary, it's an assurance that you can implement bloody efficient data structures yourself, if the STL doesn't do what you want. This is what makes C++ a generally fast language: it's not optimized for special cases.
> You can learn what all of C++'s syntax does,
Most programmers can't. Experts do, but for the rest of us… C++ is impossible to parse, and has many, many, MANY pitfalls: http://www.yosefk.com/c++fqa/
Which is precisely what Rust is trying to avoid.
Will they change * to `RawPointer<T>` too? ;)
I'm not happy about change of [T] to Vec<T> either.
I hope all these changes will make a full circle and get a first-class syntax back.
Perl-y symbol nonsense is not a first-class syntax
[T] wasn't a thing in Rust. ~[T] hasn't gone and won't go away either. See this email from acrichton: https://mail.mozilla.org/pipermail/rust-dev/2014-April/00935...
(The borrowed slices &[T] and &mut [T] are certainly here to stay.)
I would argue that this change elevates the clarity and quality of Rust syntax.
Do note that the Cons cell construction becomes `Cons(1, box Cons(2, box ...))` which doesn't seem particularly bad.
As Rust becomes more popular while we try to bring home the last few major changes, we've had this problem a few times, where the broader community is surprised and partially shocked and disappointed at a change that has been brewing for a while.
Certainly, more effort could be put into messaging and explaining how big changes are going to effect users. Figuring out how much, and what sort of, messaging is sufficient is not easy, and somebody is always going to find something objectionable about nearly everything. Additionally, Rust is still in an alpha state, in heavy development. The project is not in production mode yet.
We're always learning from mistakes and evolving the process.
Finally, Rust is in the home stretch. In some sense, we're at a stage where we need to hurry up and just get it done, at the expense of some community fallout along the way. If Rust is good, the ill will from the churn will be forgotten in time; if Rust is bad it will die. With the attention Rust is attracting, community momentum wants Rust to stop changing, but Rust must change a bit more. Rust needs to get to a point where the design can be justified and maintained for years to come, before its own popularity forces it to slow down.
To clarify, what has distraught me the most about this change is not even the change itself (although I do think that, however small or warranted on its own it may be, it may have long-term ramifications that need to be considered much more carefully), but the way the process seemed to diverge from the fairly open and inclusive discussion that seemed to regularly take place in the past. This change may well have been brewing somewhere for a while, but from the point of view of someone who is not in the "inner circle", what was perceivable of it essentially was the thread I linked above some 16 days ago (which was remarkably full of responses by apparent "insiders" that aimed to discourage people from discussing the topic at that point in time) and the thread announcing the RFC, in which even some of the most basic questions needed for a casual observer to understand the tradeoffs were unanswered when the proposal of the RFC was mainlined something like half a day later.
The end result is something that looks similar to what happened with GNOME 3 to someone who was in the opposing camp in both cases, where the "in-group" at some point decided that they had identified a development agenda of utmost importance to what has in fact been part of the central mission of the project all along, and increasingly came to send the message that they held disagreeing members of the "out-group" to be obsessive hecklers and people who should just get out already if they don't like it - although nothing says it has to end up like that, where this took the state of Linux desktop environments is easily observed nowadays.
When considering the overall lifecycle of a project, the number of remaining "difficult decisions" is a far more accurate metric of completion than code or concept. Frustratingly, it is pretty much bound to increase towards the full end of the scale of the latter when the development process is a highly public affair. People wouldn't argue this topic with such ferocity if it wasn't for their appreciation of your work so far and high expectations for this project; please don't let that go to waste by shortcutting the decision process when you have come this far.
And just out of curiosity, where did the rust designers get the inspiration for "fn"? I know plan 9 rc and clojure have "fn". Or maybe it's independent invention?
As far as I know it came from ML.
ML etymology sounds pretty plausible (I didn't know ML uses fn) since I can see there is a strong influence of ML on Rust.
The language is still in flux and there is only one way to find out whether we will miss '~'. Let's do it.
Note: I'm not laughing at you, I'm laughing with you.
It may be headed in the direction of stability, but that's quite different from stabilizing.
We can consider it to be stabilizing once the language and standard libraries have been frozen, and the only changes happening are very minor ones that are fixing critical bugs.
Until then, it's still in the research and development stage, at best.
But meh, this maybe ours is a (boring) purely definitional disagreement - I think we can consider it stabilized once the language and standard libraries have been frozen and only bug fixes and minor changes are occurring.