Such a Little Thing: The Semicolon in Rust
lucumr.pocoo.org
lucumr.pocoo.org
foo() => evaluates to the return value of foo
{ foo() } => evaluates to the return value of foo
{ foo(); } => evaluates to nil
The last case is not a special case. The last expression in the block determines the returned value. There is an expression separator in there, so there must be two expressions which are separated by it. The first one is foo() and second one is ... wait for it ... the empty expression. Which value should EmptyExpression have? Of course: nil, which would be called void in C-land.Hence, ; remained as a statement terminator, and the return keyword served to indicate returning an expression.
However, the D lambda syntax does not require a ; when the lambda body consists of only a single expression:
http://dlang.org/expression.html#Lambda
and in practice this has turned out to be well liked.
// This function lives at the top level of the module, so its types must be annotated
fn foo(bar: int) -> int {
// Here's a closure without type annotations:
let baz = |qux| { qux + 1 };
// And a closure with type annotations:
let baz = |qux: int| -> int { qux + 1 };
// Which means that this would be a compile-time error:
let baz = |qux: int| -> int { qux + 1; }; // returning nil, but expected int
// But if you expect this to return int, then there are very few cases where
// this might not be a compile-time error:
let baz = |qux| { qux + 1; };
log(error, baz(bar)); // will print "()", but did you want bar + 1?
io::println(fmt!("%?", baz(bar))); // same as previous
io::println(fmt!("%d", baz(bar))); // this one *is* a compile-time error
// macros ftw
baz(bar); // this will also be a compile-time error
}AFAIK, Rust doesn't have nil. So, the closest thing is using the option type.
Though I can think of an even more crazy variation. Declare the semicolon a binary operator and allow overloading it.
As funny as that sounds, Haskell provides something like this, as do-notation can behave differently depending on the monad it is in.
http://en.wikibooks.org/wiki/C%2B%2B_Programming/Operators/O...
My mind is blown. You can write bottom-up code in C++ :)
EDIT: no, you still can't the expressions are evaluated first, before calling "operator,"
In other words, the semicolons of Haskell are only tenuously related to semicolons in languages like C.
This is also not such a big problem when a human reads the function. We can always see at a glance the return type of the function. Either it's explicitly declared, or it's a lambda whose usage is readily visible. Semicolon or not, you can easily guess if the function is supposed to return its last expression, or not.
By the way, the compiler could do the same. Knowing that, there probably will be helpful error messages such as "did you forget the last semicolon?", or "should you remove the last semicolon?".
The semicolon is really just a small confirmation. That's why they didn't chose a heavier syntax.
foo :: Int -> [a]
It should always evaluate to a list. There is no such thing as a null pointer. The only thing that comes close to a nullable type is an option type such as Maybe: data Maybe a = Just a | Nothing
The function bar :: Int -> Maybe [a]
can either evaluate to a Just [a] or Nothing.[1] Actually, I lied a bit, since there is bottom: http://www.haskell.org/haskellwiki/Bottom But bottom does not fullfil the same role as, say null pointers.
Almost a year later, I'm just as in love with Rust's semicolon rules as Armin is. However, I bet new users will still be just as instinctively revolted as I initially was.
The downside is that you would have to put () (Rust's version of “nil”) in a bunch of functions to fulfil the requirements of the callback's signature since otherwise the type inferred from the function would be the value of the last expression
Wouldn't co/contra-variance solve this entirely? It works just fine in Scala for example scala> def runTwice(f: () => Unit) = {
| f()
| f()
| }
runTwice: (f: () => Unit)Unit
scala> runTwice{ () =>
| println("moo")
| 1
| }
moo
moo
Note how it expects a function that returns Unit, i'm passing in a function that returns an Int (1), but the compiler is perfectly happy.The semicolon is the least-annoying non-letter character to type. It's right there on your home row.
Still, even on an American keyboard () are not terribly well placed.
I've been thinking about adopting the NEO2 layout(with some modification so that I get all of åäö) but haven't gotten around to it yet. Has anyone tried that for programming?
I have all my symbol characters easily accessible from the normal A-Z keys using a third shift state. It's awesome, and works with any keyboard layout (QWERTY, or even Swedish QWERTY are all fine, as is Dvorak, Colemak etc). Typing a previously uncomfortable sequence like "for (int i = 0; i < count; ++i)" is even pleasant now.
It takes about two weeks to get used to. If you try it, make sure the third shift state can be accessed from either hand (i.e. you have to have left and right keys, just like with shift or control).
I wasn't aware of NEO2 so I developed my own symbol layout using a genetic algorithm running over my code corpus, then adjusting to taste.
"practically all computers (except custom-made, e.g., in public sector and some Apple computers) use standard US layout (commonly called Polish programmers layout, in Polish: polski programisty) with Polish letters accessed through AltGr (AltGr-Z giving “Ż” and AltGr-X giving “Ź”)." http://keyboard-layout.info/#Polish
Russia uses shift+4. Can't imagine how annoying that is.
The Dutch commonly use the US keyboard layout, which makes sense because Dutch and English have the same alphabet.
We're trying to mix logic that executes on each element in a sequence, with logic that controls how to iterate over that sequence. That is, a "return 42" statement will tell .each to stop iterating, but its contained within a block that is supposed to do things to individual elements.
The mathematical concept just doesn't sit right with me.. I guess if all you have as an iterator is 'each' then that would necessitate finding an additional way to modify iteration in some way, but still.. I don't like the fact that a return statement can break out of something outside of its own scope
edit: even the low-level alternative seems nicer
for(blah;blah;blah) {
stuff;
}It also means you can't put more expressions on the same line because doing so requires a semi-colon which then eliminates the special semicolon behavior. Having an explicit 'ret' keyword means you could accomplish that and have more expressions for a separate block on the same line if desired.
1: http://c2.com/cgi/wiki?SyntacticallySignificantWhitespaceCon...
http://lucumr.pocoo.org/2011/2/6/automatic-semicolon-inserti...
Technically, Python is a much better example of optional semicolons since you can use them and they're not required. Of course, the real reason they exist is for compound statements.
http://stackoverflow.com/questions/8236380/why-is-semicolon-...
You pretty much only need to remember to separate tokens where lack of whitespace might create a different vald token, and ensure any expressions that you want to let span more than one line ends with something that expects following tokens to make a complete expression.
Personally I can live with that easily. I can not live with significant indentation, on the other hand...
a ; b is the operator which returns the value of b. And, there is some syntactic sugar to make a ; equivalent to a ; nil which returns nil.
However, that template code was utterly hideous.
fn find_even<T: Num>(vec: &[T]) -> Option<T> {Only sorta - consider Haskell's indent rule: code which is part of an expression should be indented further than the start of that expression. Generally, this is something you should be doing to make your code readable regardless of the statement termination.
The presence/absence of the ; is both explicit and subtle enough to not annoy.
map(lambda t: t ** 2, [1,2,3,4])
or even: [(lambda t: t**2)(x) for x in [1,2,3,4]]Why would you have that lambda?
[x**2 for x in range(1, 5)](As a prerequisite to making its point, the article teaches many intricate details regarding two dynamic languages which even most of their practitioners probably never think about, then dives into more intricate details regarding a language which most of us have probably never even used, let along grokked all the details of. I sense that there's something important here for me to learn, but from where I sit this article is a lot to bite off all at once.)
A tight summary by someone who understands all this would be appreciated by many readers, not just myself.
A tight summary of all of it would be as long as the article. However, the point about semicolons in Rust is:
1. ; is a separator, not a terminator.
2. a;b separates a and b.
3. a; is a special case which means a;nil(or whatever is the equivalent in Rust)
4. The last expression in a function will be the return value of the function.
5. If the last line in a function is "a", it returns a. If it's "a;", it returns nil(from 3)
Why all this talk about Python and Ruby? Rust is not really competing with those languages-- it is quite specifically designed as a C++ replacement. They should be comparing themselves against C++, Golang, or D.
The Golang solution would just be a function which returns the next element or nil if there are no more elements. That function might be part of an interface, if you wanted to generalize it across several types. To me, this is a lot simpler than the other stuff that was discussed.
To go back to Rust specifically, rather than having magic semicolons, why not make the "break" statement take a value, which becomes the return value of the block? Despite all the tl;dr there was not much discussion of design alternatives.