Language Design: Use 'ident: Type' not 'Type ident'
soc.me
soc.me
A * B;
is a multiplication or a variable declaration, or whether A<B, C> D;
is a two comparisons with a comma operator or a templated variable declaration, without first knowing the names of all declared types.This means in practice that you have to declare types before they are used in a file, which means forward-declarations if they are defined later, that you can't separate lexing and parsing because the parser has to provide constant feedback to the lexer, and that misspelling a type name can lead to a syntax error! C++, not content with merely inheriting C's problems, throws in the “most vexing parse” as a bonus.
struct foo { int x; int y; };
which is easy to parse. Then came typedef. With user-defined types, foo*bar;
isn't parseable with a LALR(1) parser until you've seen the definition of "foo".
Even the ordinary case foo bar;
needs more than one lookahead to parse.
This is a headache for compilers, and a huge headache for anything that wants to work on single source files without seeing the included files.It's a big win if files are parseable without their dependencies. The Pascal/Modula/Ada family all are. I think Go is. Not sure about Rust, but it probably is. C and C++, no.
(The hacks for template syntax in C++ are painful to think about.)
An easy solution to the problem would be to do what e.g. Haskell does where case-sensitivity determines a type name. For better or worse, C didn't do that.
[string] $foo = "foo"
Rust has made a point of being parser-friendly, to support tooling. The entire reason the turbofish ::<> exists is to avoid ambiguous grammar. (This operator seems to vex people, I don't get the hate.)
One thing to avoid is depending on the existence of the symbol table to parse correctly. (C++ has this problem.) Needing a symbol table makes it hard for parsers like source code formatting and syntax highlighting.
Another is to avoid raw string literals that reach back and try to unwind the results of previous "phases of translation".
Redundancy in syntax can be a bonus.
I agree. I even wrote an article about that:
Amusingly, the Delphi dialect of Pascal has an ambiguity in its grammar from how it made function pointers work:
function g: Integer;
// ...
f(g);
Depending on the definition of f, this could be passing the function f by reference (as a function pointer), or passing the result of calling f.Delphi looks at the argument type to resolve the ambiguity, but it stays awkward in overload scenarios, when there's a choice between Integer or function pointer arguments.
At some point around about 2009, I added the ability to specify explicitly that you want the call, to eliminate the ambiguity:
f(g()); *c;
at the beginning of a translation unit.Didn't K&R have a tiny, tiny C declarations parser (printing out 'human readable' equivalents) example in their book? I think the caveat was that it needs to assume it's dealing with a declaration...
I'd be very surprised if Java's grammar were context-free. Do you have a source for this? I wasn't able to find one with a quick search.
This isn't true of all programming languages.
In C you cannot encode a Turing machine that is executed by the compiler at compile time. In Brainfuck you cannot encode a Turing machine that is executed by the compiler at compile time. In C++ you can encode a Turing machine that is executed by the compiler at compile time.
That is the difference we are discussing here.
But you can get quite close with macros
Yes, you can do a lot with the C preprocessor. You can also do a lot in languages that only have bounded loops and are therefore not Turing complete. You can either express nonterminating computations (Turing completeness), or you can't (still powerful, but dramatically less poweful). This question is binary. There is no fuzziness, there is no approximation, there is no "quite close".
Well-formedness includes type checking. But even without full type checking that can be done later, it also includes things like being aware, in C, of whether a given identifier is declared as a typedef in the current scope. So while C has a nice context-free "rough syntax" formally specified in the standard, its actual input language is context sensitive.
As for Java, the first example that comes to mind is that constructors must have the same name as the class they belong to. This "choose whatever identifier you like, but at some later point repeat that exact same identifier" is a very typical example of something that is not context-free.
You might disagree whether this constraint is part of what you consider Java's "grammar". So the answer to your question depends on what language level you are thinking of. But whichever level you apply to Java, you should apply the same to C++. C++ also has a context-free "rough syntax" in its standard.
Context free means something different to what you're saying here. This is a good discussion of this topic, I never knew C++ was so irregular and informal.
i.e. without extra rules you can't tell either way, so it comes down to the extra rules. `A * B` is unambiguous for type-before-name if you require "pointer-int" since it would need to be `* A B` for A to be a type.
(quite possibly there are counter-examples when you get into weirder corners, but my point is that it's not as simple as presented)
A * B; // declaration or multiplication?
in an unusual way. If it was a multiplication, the result would be thrown away and be pointless. Hence, it is a declaration.But what if A overloads the * operator and the side effects are desired rather than the result? D has a philosophy that arithmetic overloads should be for arithmetic-like operations, not I/O, template metaprogramming, or other nonsense. Hence if you try it like that, too bad, so sad, it'll be treated as a declaration.
(There are some other things in D that are designed to discourage trying to overload arithmetic operators for non-arithmetic purposes. For example, < <= > >= cannot be overloaded individually, only as a group.)
In 2020, not sure we should care that much about how hard compilers have to work to achieve this. Computers and software are here to support us--we're not here to support them.
That's not right. Long compile times are a real issue for some programming languages even today (Rust and C++).
> Computers and software are here to support us--we're not here to support them.
A false dichotomy, and not a perspective that offers any insight. Sometimes a low-level language is appropriate, and sometimes it is not.
I feel part of the problem with rust is it doesn't have quick and dirty mode thats fast and a production ready mode that is slow but does all the checks. I do this with my C programs, using various formal analysis tools which are slow to vet the code before releasing it.
The thing that will massively speed up rustc is the current re-architecting of it. We’re at the point of “few percent here, few percent there” with the current design. These add up over time, of course, but batch compilers are inherently slower than the newer style ones (after an initial compile).
Most editors did this on save before LSP came along. For logic, this is fast enough because the type system catches enough mistakes and I don't need to run tests all the time. For UI though, a faster iteration cycle would be nice...
Not sure what you mean on the dichotomy. If someone says that a language needs to have X because that will make things simpler for the computer, I say that they are wrong. The goal, the only reasonable goal, is to make things better for humans.
The compile time could be long because the implementation is poor. But it's also possible that the specific requirements do not allow for a significantly faster compile time.
That's why the requirements matter. They determine the space of possible implementations. If the requirements eliminate all "fast" implementations, then the resulting user experience will be poor because of slow compile times.
Its verifiers are awfully finicky, and tuning the parameters (including selecting the most appropriate verifier) can mean the difference between successful completion in a few seconds, and outright non-termination/timeout-with-failure.
It's true that the answer is to have better verifiers, but that's not just a matter of tweaking the verifier code, it's a serious research challenge. One of the most serious problems with formal methods is the ability to scale.
Declare a local named neko of type Kitten, passing a value to its constructor:
Kitten neko(42);
Declare a local named neko of type Kitten, without passing a value to its constructor: Kitten neko;
Declare a function named neko with zero parameters and with return type Kitten: Kitten neko();
https://en.wikipedia.org/wiki/Most_vexing_parse , https://stackoverflow.com/a/620149/ Kitten neko(Felis<0xCA7>::catus);
where catus may be either a typename or a variable depending which template instantiation you landed on.There's probably some way to say:
Kitten neko(Felis<is_function_type<typeof neko>::value_of>::catus);Take the first example you've given, and let's just talk about variable declaration:
int* a;
int * a;
int *a;
That's all the same. That's wrong. Obviously, you can write a lexer and parser that doesn't give a crap, but the human mind does. If you think it doesn't, that's only because you've internalized the various cases.It should be this:
int* a;
The type we're talking about is a pointer to a variable of type int: in other words, "a" is an int-pointer. If you were to create a macro, you'd do something like this: #define int_ptr int*
Do you see what I'm getting at? The star in this case is a suffix, equivalent (in our minds) to "_ptr". Conceptually, it doesn't belong anywhere else than attached to the type. It's a compound type, conceptually.Now, take the star being used in a different context:
int b = 10;
int* a = &b;
printf("%d\n", *a);
There, though we see the same character, it's a completely different thing. It's a dereference operator. Conceptually, it belongs attached to the pointer variable it is dereferencing; and since the star is already used in one context as a suffix, here it should be used as a prefix.This is no good:
printf("%d\n", * a);
It doesn't matter that "this compiles." That's not what this is about.Many C programmers (and programmers in other languages, even Python) are used to writing things like this:
int c = x*y;
That's wrong. Sure, the lexer and parser don't care. But that makes the language worse, for the human operator. "But it saves space!" Spare me.The thing with this one example, using the star, is that what we have is the equivalent of a homonym. We have one sign that is actually three different words. Mandating spacing removes the ambiguity you're complaining about.
C is what it is, but if we imagine someone were going to write it today, they should incorporate the above and mandate spacing. For the sake of the humans. "ident type" is not the only solution.
> int* a;
I must disagree... the following sends the wrong message the reader:
int* a, b;
Though I also agree that dereferencing is best without a space. Perhaps the correct suggestion is to not declare pointers and instances on the same line, but that's a convenience that many seem to enjoy int *a, *b;
In D: int* a, b;
The use of whitespace makes the distinction clear, although the parser doesn't care.I have been thinking about making a toy compiler (I wanted to write a borrow checker) that treats bad code as an error, solely aimed at numerical code - I have recently had "Scientific Programming in Python for physics etc." inflicted on me.
Slight tangent, but I think if Haskell enforced some kind of whitespace a la Python it would be much more approachable in real codebases (Haskell is usually quite readable if you are just translating mathematics into code but it - to me at least - feels dreadful as a productive language because of the way a lot of functions seem to be dumped onto into the text editor in a lot of the code I have read)
I use python at $dayjob, and I run into the limitations of whitespace as syntax all the time.
https://www.python.org/dev/peps/pep-3103/
It's quite frustrating as there are a lot of situations where a simple switch would be so nice...
It's been a long, long time since C required a programmer to declare his variables at the top, and best practice argues that you declare a variable as close to its use as possible. So, there really isn't the same case for multiple variables on one line, whereas it may have been a little more forgivable, once upon a time.
Moreover, when I was first learning C, I learned from an O'Reilly book by Steve Oualline. I remember him saying, when it came to operator precedence, that coding style that relied on the rules was a really bad idea. I think he said something like, "Multiplication and division come before addition and subtraction, and use parentheses for everything else."
My bottom line is C has a little too much "convenience" to it. (Granted, to my taste.) It goes back to my van Rossum comment. I'm against special cases and loosey-goosey stuff.
A pointer to an int "should" be &int, not int*. That we use *, the
dereference operator, to indicate pointers is wrong. * in a type means it's an
address, but * in a value means it's not an address. That's nuts! Make it
consistent and use & both places. If you need to have a distinction between
refs and pointers, it should be that pointers are nullable refs: &int? or some
such.
Edit: How to escape asterisks in HN?I did not want to stray too far from what C does now, for the sake of argument. I just wanted to say that the ambiguity in C is often because it's so loosey-goosey with whitespace.
int* a,* b; // ??
And that's usually why most C style guidelines use int *a, *b;
Now, if we treated "int<asterisk>" as the full type name in the syntax then you get a much nicer result: int* a, b; // much nicer!
Though that does have its own tradeoffs. But that would require actually changing the C syntax, at which point you might as well make it the postfix syntax: var a, b int*; int *f
is declaring that *f
will be an int.This perspective addresses why the star belongs with the name, why it's star instead of ampersand, why you need to repeat the star for multiple variables, why the brackets goes after the name for arrays, and it will sort of get you where you need to go with function pointers (although there's an automatic promotion which means things will work at the call site that won't work for the type).
This isn't to say alternative constructions mightn't be a better choice in a new language, but it's much more parsimonious when considering C than memorizing a bunch of special cases.
If we have
typedef float *floatp;
then the following compiles: float x, y;
floatp xp = &x, yp = &y;
while the following does not: float x, y;
float* xp = &x, yp = &y;
You have defined a new type (which happens to be equivalent to an old type - note that we don't get nominal type-checking from typedefs); the above logic still holds. foo*bar; // multiplies foo and bar
foo* bar; // declares bar as pointer-to-foo
foo *bar; // declares *bar as foo (so bar is still a foo*)
foo * bar; // multiplies foo and bar
These are all unambigous and there are only two semantics between them. You also have: foo* bar,baz; // baz is a pointer
foo *bar,baz; // baz is a foo
foo *bar,*baz; // baz is pointer again
foo* bar,*baz; // baz is now a pointer *to* a pointer
This is all obvious - or at worst unambigous - to a person reading the actual code (without trying to correct for the idiosyncracies of a parser), so the language should either match that or spit out a warning about unsupported spacing.> [`int c = x*y;` is] wrong.
No, that's unambigously multiplication.
class Id<T>() {
Instead of the D syntax: class Id(T) {
Not using < > for template parameters eliminates all kinds of parsing problems.You can have an ambiguous grammar in a language that puts the type after. A * B could be ambiguous even if A is the identifier name and B the type.
All arguments in the "Language Design" notes are of the same nature.
1 - (sorry, formatting is messing this up) mandate the declaration:
'A*' or '*B'
without space (depends if you think the pointer is a type of variable or a type of type)2 - add some keyword like "type"
Though the compilers are able to deal with it, so it's not a big issue
For example:
val x: String = "hello"
The type interrupts the flow of data from "hello" to x, so one thing that pops into mind is that this is typecasting the value to a string before storing it. Nope.Another possibility I instinctively see this as is doing a comparison and assigning the result (either true or false in this case) to x. Nope.
And human-language wise, colon is "description: explanation" (or more generally: general to specific), which actually fits this syntax better:
val String: x = "hello"
...and at that point, just remove the extraneous stuff: String x = "hello" val x: String
x = "hello"
The type at this point is almost like a comment.For declaration and assignment though, I agree that reading "ident: Type" is harder for me.
Perhaps an interesting idea would be to have the type at the end of the expression. Like so:
val x = "hello": String
Essentially, you're making a type assertion on an expression. Since it's an assignment expression (the value of which would be the assigned variable) then it also type checks the variable. val x = "hello"
Standalone declarations are the most important problem to solve here. let x = "hello": String
let x = "world": String
If 'let' is a keyword that implies single assignment, then problem is solved.Alternatively, using pattern matching alarm Erlang would be even better:
x = "hello": String //x is bound to "hello"
x = "hello": String //equality check is successful
x = "world": String //Error: values not equal> val x = "hello": String
Rust has had exactly this (so-called type ascription) for a while in nightly, but it's not stabilized yet due to some unresolved issues.
Breakfast: eggs and bacon
Lunch: falafel sandwich
Dinner: BBQ pork and slaw
The type of the variable is the "explanation" here. val x: String
"x, which is a String" val x: String = "hello"
"x, which is a String, is initialized with 'hello'"If anything you've just argued that
String: x, y
is a preferable declaration style.
val breakfast: “eggs and bacon” = String
and not val breakfast: String = “eggs and bacon”The use of a declaration with initialization as the example here muddies the water. Whatever syntax you use has to work for uninitialized declaration of variables, function arguments, and structure members:
val breakfast: String
fun serveBreakfast(breakfast: String)
struct MealPlan {
breakfast: String
}
Declaration with initialization just needs to be consistent with these.> Another possibility I instinctively see this as is doing a comparison and assigning the result (either true or false in this case) to x. Nope.
You can write it as
val x = "hello" : String
if you prefer. In fact that's a great advantage of this syntax: any expression can be optionally ascribed with a type. If you write the type first then it becomes too intrusive (and too much like a typecast, which absolutely should be intrusive).> And human-language wise, colon is "description: explanation"
True enough, but what other syntax would fit in postfix position? In human language we'd probably use commas ("Bob, chef"), but that seems a bit too ambiguous in a programming language.
But here's an idea. The OP wants meaning to flow left to right for lambda declarations. Why not have assignment go that way too.
"Hello" => x: string
Might work. Computer science pseudocode uses notation kind of like that.Also, I understand now why Scala and Rust always seemed clunky to me.
Does it? In my time at the university, assignments in pseudocode were usually written as
x <- "Hello"In Scala, in particular, types are not the assigned type like in C (where they also serve as the storage specification) -- they are assertions, that the compiler will check are compatible with the code.
So `val x: int = "hello"` is no good and the compiler can cut it short right there; this is especially useful as call-site documentation.
It's also the case that there are plenty of languages where all types imply storage specification and they support plenty of type elision.
let x = 5i32;
Sadly it's nowhere near as elegant for string literals.What do you mean? For string literals you just do
let foo= "bar";
and that's it.
BTW, in Rust you can omit most variable type annotations since the compiler is able to infer them. You have to give type annotations to functions though.
let foo= "bar".to_string();
It's on purpose though, because Rust likes to make heap allocations explicit, and I find it fine to be honest, but you mileage may vary.
> I’m more fond of using .into(). It requires adding type hints in some cases, but for most cases it is shorter than the alternatives. Especially when passing a string literal to a function that requires String.
> Now that specialization for str::to_string() has landed, we can safely say that to_string() has the same performance as to_owned(), and thus to_string() should be used since it’s more clear
> I now strongly prefer to_owned() for string literals over either of to_string() or into().
yay language complexity
You could argue that Rust strings are complex, because they are : having 8 "string-like" types (owned strings vs string slices (references), cstr/cstring, path and pathbuf, osstr/osString) is complex, but having three different methods doing exactly the same thing isn't.
Yes it's redundant, but what would you rather have : str as the only easily convertible type with no to_string method? Or the only reference type without to_owned? No ability to use into for strings while it works everywhere else? Obviously, redundancy is better than theses alternatives.
For example, knowing something is a float, double, int, or string can make an ident named "releaseTime" mean different things.
I also find that whitespace is more consistent when using Type ident, you get rivers where the spaces all line up, so all the type declarations AND ident declarations align. Whereas with ident: Type, I find it much more difficult because of the variable length of identifiers. (Yes, one could fix this by using tabs, but if idents vary in length by more than one tab stop, it becomes difficult to read horizontally.)
> This means that the vertical offset of names stays consistent, regardless of whether a type annotation is present (and how long it is) or not.
Why is this necessarily desirable? Strong typing systems have very expressive types, to the point where if something is typed correctly, most of the time my property names are just an alternative casing of the type. Types can be just as expressive or even more expressive than variable names.
> The i: Int syntax naturally leads to a method syntax where the inputs (parameters) are defined before the output (result type), which in turn leads to more consistency with lambda syntax (whose inputs are also defined before its output).
Maybe this is nice in theory? But `Int` really isn't an output here, and the value being assigned isn't either. Rather this seems more like `f(i, Int, value) -> assignment`. It seems just as arguable that `f(Int, i, value) -> assignment` is appropriate.
It seems like some of these are rooted in a "pure mathematical" approach which I can surely appreciate, but ultimately lambda calculus is as much a language as any other programming language, saying "lambda syntax does it this way" doesn't convince me very much.
As other posters have stated, this order makes parsing easier. But I also suggest this benefit extends to your own brain's parsing ability as well. The old order is indirect and suboptimal and makes you think harder.
Making parsing easier for the compiler is a convincing benefit, would have been nice to see that mentioned in the article. I think that's a substantially stronger reason to prefer types after names. I'm not sure if I parse either faster or slower though.
As you yourself noted, personal anecdotes are really not an argument. Someone could say they find Java easier to skim than Rust and we'd be nowhere. Like arguing which end of a boiled egg to crack first.
> As other posters have stated, this order makes parsing easier.
Programming languages don't exist to make itself easier to parse. They exist to make it easier for programmers to program. Otherwise, we wouldn't have such things like syntactic sugar. Hell we would just write in machine code and do away with assembly and higher level programming language. And parsing is a simple and superficial one time step. Being a tad bit more difficult is not a convincing argument.
> But I also suggest this benefit extends to your own brain's parsing ability as well.
Based on what evidence?
This is the problem with tech evangelism. It has the same problems as religions, lots of claims, no evidence.
Then what is? If you're looking for a randomized sampling of programmers with sufficient sample size, you're not going to find it here.
> Programming languages don't exist to make itself easier to parse.
No, but a fine example is that of C++: the difficulty in parsing means that if you make a typo, the error message you get might be bizarre and confusing. A compiler for a language that's easier to parse will have a much better idea of the programmer's intent and can provide a much better error message. I find it astounding how often rustc can figure out exactly what I wanted to do and suggest it as a note after the error message.
I would think that more-useful error messages pass your test of "make it easier for programmers to program".
While we're talking about making it easier to program, "name: Type" make it possible to avoid typing out "Type" at all, and letting the compiler infer it (no, this isn't good and readable to do in all situations, but often it's fine). If you have "Type name" style, and try to add the ability to infer types, you end up with Java's "var" abomination.
Regardless, I'm in agreement: I find "name: Type = blah" much easier to read. I read it as "name is a Type that is equal to blah". This also is an improvement in parameter lists, when they're lined up vertically:
def foo(bar: String,
baz: Int,
quux: Foo)
I find that much easier to mentally parse to determine parameter order than void foo(String bar,
int baz,
Foo quux)
Worse, imagine that all three parameters were of the same type, requiring a scan to the right to read the names. The important information to me at a glance is the name of the parameter, not its type.As someone who cut his teeth on C and later Java, much later learning Scala and Rust, I immediately liked the style of the latter two much better. Lately I've been doing a lot of Java and get constantly annoyed at the "backwards" order.
> This is the problem with tech evangelism. It has the same problems as religions, lots of claims, no evidence.
I suppose you could argue that what I've written above is just personal preference, but I see it as a bit stronger than that.
Evidence. Maybe a study showing programmers have a natural preference? Or scientific evidence? Anything more convincing than "Rust evangelist" anecdotes.
> No, but a fine example is that of C++: the difficulty in parsing means that if you make a typo, the error message you get might be bizarre and confusing.
Difficulty parsing? If it didn't parse and found an error, then it means it didn't have any difficulty parsing. That has more to do with the complexity of the language itself than parsing. Parsing is a very simple matter. Or maybe the compiler for one language is better? Also, I thought we were comparing Rust to Java?
> I would think that more-useful error messages pass your test of "make it easier for programmers to program".
It does, but once again all you've done is provide anecdotes without any examples or evidence.
> Regardless, I'm in agreement: I find "name: Type = blah" much easier to read.
I don't. The most important part of "name: Type = blah" is the Type. So it's nice to have it first. But then again, there are people who love dynamic programming languages. So once again personal preferences and personal anecdotes aren't convincing arguments.
> As someone who cut his teeth on C and later Java, much later learning Scala and Rust
Yeah, I too fanboy over new languages I learn. But then I get over it and move on with my life. My guess is you just wrote toy programs in scala and rust and nothing substantive.
> I immediately liked the style of the latter two much better. Lately I've been doing a lot of Java and get constantly annoyed at the "backwards" order.
So then use Rust? Why are you using Java?
> I suppose you could argue that what I've written above is just personal preference, but I see it as a bit stronger than that.
I don't have to argue it. All you've provided is personal preference. "I find "name: Type = blah" much easier to read. " is personal preference. It's no more a convincing argument of anything than you prefering chocolate over vanilla shows that chocolate is better than vanilla.
The point is that you want variable declarations and function signatures to be consistent, so you either write
val i : Int
def f(x: Int) : String
Or Int i
String f(Int x)
And if you do the latter then you have a confusing syntax because the output type comes before the input type, and it's very hard to do lambdas in a way that looks consistent.Because rythm makes text easier to parse for the human eye
>most of the time my property names are just an alternative casing of the type
String or int are very rarely appropriate variable names
I very much agree, but that only undermines the point if they are often appropriate type names. In the sorts of languages the GP was trying to restrict that sentence to, I don't think that's the case. I even have some doubts that it's true in C.
The moment `Type identifier` syntax encounters higher order functions and types, you end up with messes of parenthesis. Figuring out what a type means then involves bouncing back and forth across the type definition.
With `identifier: Type` complex higher order types still parse linearly left to right.
It's enough of a UI issue that people will end up avoiding higher order functions in `Type identifier` languages simply because they're a mess to express.
Syntax isn't unimportant, but don't waste energy on trivial matters like these. Just pick something and people will get used to it. Focus on the semantics of your language - that's what really matters.
One of the things that using ":" does is it makes what is "type" and what is "name" unambiguous to both human and computer.
This is, famously, one of the failings of C/C++. Determining what is a type and what is a name is excruciatingly difficult.
I'd be more open to this kind of discussion if there was an ounce of actual research behind what makes syntax more/less readable. As it is, it's just a bunch of people arguing endlessly about their very specific preferences. Just pick something sensible and move on.
Having said that, this article doesn’t advocate “ident: Type”, it advocates ”marker ident: Type”.
That marker is essential for ease of parsing and thus for autocomplete (it won’t try to autocomplete the ‘ident’ part by looking at variables in scope or function names, for example) and error messages (it could signal when name shadowing occurs, for example)
#104#101#108#108#111,[Space]world![Space][Space][Tab][Space][Space][Tab][Space][Space][Space][LF] [Tab][LF][Space][Space] [Space][Space][Space][Tab][Tab][Space][Space][Tab][Space][Tab][LF] [Tab][LF][Space][Space] [Space][Space][Space][Tab][Tab][Space][Tab][Tab][Space][Space][LF] [Tab][LF][Space][Space] [Space][Space][Space][Tab][Tab][Space][Tab][Tab][Space][Space][LF] [Tab][LF][Space][Space] [Space][Space][Space][Tab][Tab][Space][Tab][Tab][Tab][Tab][LF] [Tab][LF][Space][Space] [Space][Space][Space][Tab][Space][Tab][Tab][Space][Space][LF] [Tab][LF][Space][Space] [Space][Space][Space][Tab][Space][Space][Space][Space][Space][LF] [Tab][LF][Space][Space] [Space][Space][Space][Tab][Tab][Tab][Space][Tab][Tab][Tab][LF] [Tab][LF][Space][Space] [Space][Space][Space][Tab][Tab][Space][Tab][Tab][Tab][Tab][LF] [Tab][LF][Space][Space] [Space][Space][Space][Tab][Tab][Tab][Space][Space][Tab][Space][LF] [Tab][LF][Space][Space] [Space][Space][Space][Tab][Tab][Space][Tab][Tab][Space][Space][LF] [Tab][LF][Space][Space] [Space][Space][Space][Tab][Tab][Space][Space][Tab][Space][Space][LF] [Tab][LF][Space][Space] [LF][LF][LF]
yep, no important at all.
That is a big claim. Is very easy to believe (I do it before, when my knowledge of programming languages was about just 3 or 4. Now is more than 12). :
But is clearly false, and is easy to prove:
async/await
go chan
fn sort<T>(of:list<T>...)
try/catch
match
All the above are just small things that have a HUGE impact in how develop programs. Also, in matter of "small" stuff that could look insignificant: [1, 2, 3] + 1 = [2, 3, 4]
this one is a huge deal in certain niches, also, another "small" and insignificant thing: SELECT ... FROM source
source SELECT ...
All this are just small things. Not all that obvious at the time. Remember how before the times of GOTO the idea of more specialized control flow was unthinkable in the minds of many.Syntax MATTER MOST. Because, is OUR interface. The space of improvement is not super-big, truth, but it impact hugely.
Also, when done correctly, it make the semantics fit like a glove or not.
Another obvious example: Do concurrency whithout syntax help (just using threads). Or performant, safe, concurrency friendly, zero-gc, system-programming, etc without what rust and other langs have bridged.
The difference between Python, C++, Haskell, Common Lisp, Prolog, and SQL isn’t syntax. If it was, everyone would pick their favorite syntax and use it all the time. What matters is how well the semantics (and their potential performance implications) match your problem. The syntax just needs to be a decent enough interface to the semantics. Frankly, it seems to me like most of your “counterexamples” are about language semantics, not syntax.
Here’s the thing. Would I like every language to have a consistent, beautifully designed syntax backed by UX research and testing? Absolutely. But language designers have bigger fish to fry. There’s little value in wasting energy talking about syntax once it reaches a basic state of acceptability.
I do amend my statement - you’re right that it’s a big, unsubstantiated claim. There are no _demonstrated_ differences. I haven’t seen an ounce of evidence that it makes a difference beyond familiarity. Furthermore, even if it did, that wouldn’t make it top priority. It would just make arguments about it sensible.
Ok, let's try: Do SQL without the SQL syntax.
P.D: I don't think we are that in disagreement ("The syntax just needs to be a decent enough interface to the semantics"), is that the claim of "syntax don't matter" make it look is just an irrelevant aspect of the language. Can be argued how much relevant, but after years on this trade, go to the C++ community (for example) and tell them to change the syntax to lisp syntax and see how much it will succeed.
Syntax is 100% tied to paradigms, idioms, and such. Is intrinsic to the language we use.
My issue is with unproductive, endless debates about syntax minutiae like the original post. Syntax doesn't matter enough to be worth it, and such debates devolve into everyone shouting about their personal preferences anyway (see: many of these comments).
> Do SQL without the SQL syntax
I'm not sure what you're saying here. The syntax of SQL is completely arbitrary - I'm sure you could think of a completely different syntax that works just fine. Let me know if I'm missing something, but it seems extremely obvious to me that the biggest difference between C++ and SQL programs isn't how they look - it's how they behave. One wouldn't dream of replacing one with the other and that has nothing to do with their syntaxes.
> go to the C++ community (for example) and tell them to change the syntax to lisp syntax and see how much it will succeed
Obviously it'll fail - good. Even if lisp syntax was way better, they've gotten used to C++ syntax and have much more important things to spend their time on.
Select(`my_table`, [`column_A`, `column_B`])
.Filter(`column_C` > 53 && `column_D` == $varA)
.Sort_by(`column_C`)
There you go: you have the exact semantic as a traditional SQL query (1:1 mapping) and only the syntax is different.
Now, one may argue that the syntax is "ugly", less familiar, that the ` are hard to type or whatever, but this is just taste. One simply get used to it. The expressiveness and semantic are the same as in SQL> Syntax is 100% tied to paradigms, idioms, and such. Is intrinsic to the language we use.
I think then we don't have the same definition of syntax. The way i understand it is that the syntax is just the way to represent these idioms and paradigms visually. What the parent is saying is that these paradigms and idiom as what is important, but the exact way they are written, not as much (as long as it is within reason)
However this:
> but the exact way they are written, not as much (as long as it is within reason)
Then what is "within reason?". Is more logical to only have GOTO than IF, is better to have ELSEIF or nest IF?, what happened if my lang say that null is the same than Option.None?, what if generics use [] and not <>?.
Whitespace matter, yes? no?
Allow unicode?
CamelCase, snake_case or what? What if all const are lowercase, types mixcase and the rest UPPERCASE?
For some, APL syntax make more sense than algol.
Talk about why, that is the point of this kind of talk.
Is VERY easy to rug this kind of stuff. VERY. I WAS in that camp before. But now, I try to build my own lang (relational), and DAMM, it start to be much clear why syntax matter, even "the exact way they are written", because switch this to that and suddenly, my lang is ANOTHER paradigm (or worse, will be CONFUSED as be).
Naming, is one the hard things in computer science.
---
I understand why is easy to dismmis this as irrelevant. Sometimes I don't see why some people are so upset about typography and font selection, or why my profesional brother complain about framing in photograph. But go and SEE what the DESIGNERS of lang say about this stuff and you will note that for them, even this apparent less-significant thing matter. you can even get a prize on the field for show the importance of syntax (http://www.eecg.toronto.edu/~jzhu/csc326/readings/iverson.pd...)!
If that mean that most will not see, GREAT! That is the mark of good design.
The same should happen for types.
> the harder it is for a computer to parse, the harder it is for a human to parse,
I don't think this is true -- assembly (or bytecode) is very easy for the computer to parse, but much, much harder for humans to parse. English is much easier for humans to parse, but pretty difficult for computers to parse.
Design a language before you design a syntax.
I do prefer postfix because I think it flows very nicely "this is-a thing assigned-to that" is nicer than "thing called this assigned-to that."
In terms of the impact on the language, optional postfix annotation makes it a bit trickier if you want to make the identifier optional, and in languages that support it you tend to see special syntax to deal with that case (which breaks the author's fetish for self consistency).
Personally I think ordering of the trio of "alias" "thing" "value" should be consistent across the language, which extends far past variable assignment, and any one of the trio can be left out.
Example from Typescript:
const foo: 'A' | 'B' | 'C' = 'A'
Which states that foo must belong to the given union type. How would this look in a Type ident language?
const 'A' | 'B' | 'C' foo
That doesn't seem right. There's no clear barrier between the type and the identifier name.
Here's another contrived example:
const foo: () => Promise<void> = async (x) => console.log(x)
Here foo is of type "() => Promise<void>". How would this look in a Type ident language?
const () => Promise<void> foo = async (x) => console.log(x)
To me, this is unclear because it is hard to tell where the type ends and the actual function begins.
Last example.
const foo: { [string]: number} = {"hello": 3}
I believe says that foo is an object with string keys and number values.
What does this look like in a Type ident language?
const { [string]: number} foo = {"hello": 3}
I think all of the Type ident examples are more confusing because it's to tell where the type ends and the name begins (this is most clear in the first example). This probably makes syntax highlighting worse/parsing more complicated/is tougher on the user. With ident: Type, it is very clear that the type starts after the ":" and ends before the "=" sign.
Some of the things have been addressed (`extern crate`).
Many of the issues I disagree with: `Buf` is strictly better than `Buffer` (less typing, like `fn`). I have no issue with mixing `CamelCase::snake_methods`, and actually find it to be quite beautiful. The good parts of being Pythonic.
I would like to see the alternatives to turbofish. What exactly is the author suggesting? And what's wrong with `println!` and `format!` ? It isn't articulated.
`[]` misuse is bad, semicolons aren't consistent, `PathBuf` is inconsistently named, etc. Agree. `io::Result`, ...
Maybe there will be some cleanup in a future language edition.
This is also an argument for using keywords of the same length for introducing a variable and a constant. If that’s desirable, it rules out the obvious choices `var` and `const`.
Possibilities include `var` and `val`, which may be too similar-looking, and `var` and `let` - but are people used to (from JavaScript) `let` being mutable? Any other options?
As to short, equal length options for ‘let’ and ‘val’: one could consider using punctuation. Forth uses colons instead of ‘fun’, and I think, in a concise language, one could get used to using, say, ‘!’ for immutable and ‘~’ for mutable. Unfortunately, they aren’t easy to type. An alternative could be to always assume immutability and only use ~ in the rare cases where one needs to mutate.
So, a simple
foo = 3
or, if one wants to simplify parsing: = foo 3
introduces a new immutable variable, and ~ foo = 3
or ~ foo 3
a mutable one. If we allow leaving out spaces: ~foo 3
that starts to look like using sigils to indicate mutable state. I think that might be a good option in a mostly immutable language.I think I would use Forth’s colon instead of ‘=‘. That would make ‘=‘ available for equality testing, allowing us to get rid of ‘==‘.
The major downside of the 'Type ident' approach, is that if 'Type' is optional, then the parser can't be sure if its parsing the 'Type' or the 'ident' when encountering the first token. In practice this isn't too hard to solve however, it can be handled with some backtracking.
In my language, Winter, I have chosen the 'Type ident', approach, mostly due to similarity with C, C++ and Java. I do sometimes wonder if I made the right choice however. Maybe it could be an option? :)
Not all IDEs are the same, though, and I'm not sure how sophisticated this feature was to implement.
If this were true, we'd have to conclude that speakers of name-then-honorific languages like Japanese ("Graham-san") are better at remembering and focusing on people's names than speakers of honorific-then-name languages like English ("Mr. Graham.")
But there's no evidence of that, is there?
I can't really agree with this at all. With type aliasing the new typename can render the variable name pretty much redundant.
"This is the standard we should all adopt!" https://xkcd.com/927/
val x: String = "hello"
String x = "hello"
The first line reads: "value X is of type String and contains hello"The second line reads: "String x contains hello"
val and : are fluff and add nothing. Arguments about it being tougher to parse would have some merit if this wasn't all figured out almost 50 years ago.
I agree with the sentiment, but that apostrophe is bugging me.