Language Design: Use [ ] instead of < > for Generics
soc.me
soc.me
This fixating on relatively inconsequential issues, while totally missing the forest for the trees, is a dangerous attribute I've found in a small subset of programmers.
Software engineers, on the other hand, worry about how to efficiently write non-trivial programs. For that problem, this change doesn't move the needle whatsoever.
Most non-junior programmers think more like software engineers than like computer scientists. And even many who think like computer scientists can see the forest, not just the trees.
It might not be the author's preference, but the argument made about it getting abandoned as the language ages anyway is demonstrably false with the most popular languages in the world (Python, JavaScript, Java, C, C++), and in languages with operator overloading it could never be true no matter how far the language evolved.
> someList(0) /* instead of */ someList[0]
This is worse. How do you disambiguate that from a function call? It's now just abusing ()
The article proposes using basic function call syntax. That makes no sense to me. And with .(), you now have 3 characters to type instead of 2 with []. Obviously all of this is pedantic, but then again so is the article.
I've written a decent amount of Scala, it's fine.
I've also written a lot of Common Lisp, where array access is done using aref. That was also fine.
fn someArray(offset: usize) -> T* {
self.base + (offset * T.size) as T*
}That is, the cardinality of the type Array<T> is exactly the same as the cardinality of the type Function<Integer, T>. Any pure function that takes an integer and returns T can be replaced with an array, and vice versa.
Same deal with Map<K, V>, that's logically equivalent to Function<K, V>.
Two problems:
1. People don't think in category theory, they have containers to contain thinks and functions to calculate things.
2. People are right, the function call really is doing something different than an array index. If a thing is different, it should look different.
What is fundamentally different?
About the only reason I can see for array access syntax existing is that in the days before optimizing compilers, people would have lost their mind if array access called into a function each time (with good reason). It took manual programmer intervention to hint what compilers couldn't cleanly infer themselves. We don't need that anymore.
While they are mathematically equivalent, in these languages they are doing different things. Function application and array access are two different concepts the symbols are trying to convey to a programmer.
This is why, e.g., in Haskell where a string is literally a list of characters, they nevertheless have a different syntax for the two.
> Haskell provides indexable arrays, which may be thought of as functions whose domains are isomorphic to contiguous subsets of the integers.
https://www.haskell.org/onlinereport/haskell2010/haskellch14...
The fact that strings are different seems to have more to do with the ergonomics of strings rather than arrays.
If we limit ourselves to the math, we can agree that a type is a domain that contains certain values and excludes others. But we have to throw out mutability because values, by mathematical definition, are immutable.
And once we've got a clear notion of what values are, we can then talk about how many values there are in a type. And that's where I'm coming from in claiming functions are equivalent to containers.
If I'm adding shifting to a new language, I'm inclined to use functions just so it's clear what's going on, and to put operations like rotation on an even footing. Also, I don't think people have an intuition for the precedence of shift operators, which makes them less useful.
foo = [10,20]
Func foo(i) = switch(i)
case 0: return 10
case 1: return 20
default: return nilYou don't. If you apply an integer to an array, it will return an entry. To me, it makes perfect sense to use a(i) instead of a[i]. In fact, in case you need some more complicated data structure to organise your data one day, you could switch to a function call for lookup without rewriting the code that uses a(i).
The question is not 'how do you disambiguate?', but 'why do you want to distinguish?' Because an array is just a special case of a function that maps an integer to something else. It is logically not different from a function with a large contiguous switch statement.
I disagree. Square brackets links to memory, paren just returns data. Of course you could implement the paren function to return a reference, but it makes the code a lot harder to read since that is not the normal case. See comparison:
a[i] = 3;
or a(i) = 3;The alternative suggested is also non ideal, since array declaration/indexing is semantically different than a function call, not all languages have method calls, not all languages have postfix method calls, I can think of a few ways to break that syntax (what if the type has a call operator and indexing operator?), and it is hypocritical.
The problem with <> is that they are used elsewhere. If you use [] by sacrificing arrays, () will cause problems because they're used elsewhere.
The lowest friction solution would be to introduce a new two character bracket. How about <: or :>? (: :)? I don't know but writing about it in that tone won't get anyone to do something different.
D uses !() for generics. It looks odd at first, but soon becomes completely natural:
struct S(T) { T t; }
auto a = S!(int);
auto b = S!int; <= when only one argument is needed
Since ! is not a conventional binary operator, this:1. parses without lookahead and ambiguity issues
2. stands out in the code as being a template instantiation
Well over a decade of experience with it (and D code typically makes heavy use of templates) with no issues shows that it works.
Meanwhile picking either <> or [] feels like that decision is based on personal preference. Obviously, you can justify choosing the most popular syntax because of familiarity but then [] was never an option to begin with.
I'm ashamed that realization was slow in coming to me. But once it did, it was key to greatly simplifying the D template syntax.
1. Angle brackets are also not uniformly superior. Technically, I think it’s fair to consider them uniformly superior, but social aspects are important too; there was reasonable concern that Rust had exhausted its weirdness budget. (Of the most common languages people may be familiar with when coming to Rust, I think Scala is the only one that uses square brackets for generics; the likes of C++, C and Java are much more common and use angle brackets.)
2. For Rust specifically, it wouldn’t actually have been just a syntactic change—if you want to change array indexing to use function call syntax, you’ve got to sort out more technical challenges there, things that would be approaching trivial in full-GC languages but which are actually rather difficult in Rust. Specifically, function calls are rvalues only (`… = f()`), but array indexing can be an lvalue also (`a[i] = …`). So you need to more or less unify the Index and Fn trait families, auto-ref/-deref might cause trouble, and placement might make an appearance in it too for best results (and that’s something that still isn’t resolved). Now I’m inclined to believe that these changes would have been a really good thing and resulted in a more coherent and incidentally slightly more flexible language (and not in a dangerous way), but it would now be even more difficult to achieve (though not impossible; the edition system could be used for the syntax side of things, and Fn/Index implementations are still mutually exclusive, because you can’t implement Fn manually on stable, so you could blanket-impl Fn traits for types implementing Index traits without breaking backwards-compatibility).
I’m simplifying the story a bit. Others may fill out more history and technical detail if they desire. I’m going to bed.
More information on this and other Rust-related stuff, from a couple of years ago (thus, well after Rust 1.0): https://old.reddit.com/r/rust/comments/6l9mpe/minor_rant_i_w...
Though I do agree that array/object literals using [] and {} are probably not necessary.
As for the last paragraph, I'd implore language implementors to continue using < > lest we get that awfully ambigous nonsense.
The primary argument against < and > as generic brackets is that the ambiguity can make for some confusing error messages. It also prevents you from pre-matching your braces before the parsing phase (a technique that enhances error recovery).
Let's see if this will work.
<
Nah.
Some of the modifications done to titles are indeed weird and contraproductive IMHO.
Input sanitizing and escaping is one aspect of the HN code I've never dug too much into. There are a lot of corner cases that don't work, but it's never been a high enough priority to fix them.
As for HN not being able to handle having them as characters in titles, one would hope that would be a “simple” matter of escaping the characters at time of submission and parsing them at time of rendering.
It’s all about the context; I think it would actually be more confusing to go with the author’s suggestion.
Furthermore, I can’t think of a place in code where the generic and comparison use cases of < or > would be ambiguous or even adjacent.
Most languages that use angle brackets for generics (e.g. C++, Rust, Swift) also allow you to explicitly instantiate generics, which leads to a syntactic ambiguity, e.g. is
a < b > (c)
a function call or two comparisons? The usual way to disambiguate in these cases is to check whether b is plausibly a type and treat it as a generic, but this is ugly on a few different levels. a::<b> (c)
Is how you instantiate a generic. It's not that ugly. Bigger problem is value type generics/const generics where you want to do some logic like a<b > c>How does Rust solve the ambiguities with const generics?
fn function<T, { ambigous const generic expression }>(arg1: T) -> T {Most compilers don't even go that far: using type information parsing is a huge no no for almost all languages that aren't C++. So for example, Typescript will flag:
a<b>(c)
as an error if a, b, and c are untyped, meaning it automatically resolved the generic application in the parser before type information was processed, while this is fine: a<b>=(c)
Since the equal sign means > is part of a >= token. The parser is automatically determining if generic application was meant or not, before any type info is considered.The grammar ambiguity is resolved as follows: In a context where one possible interpretation of a sequence of tokens is an Arguments production, if the initial sequence of tokens forms a syntactically correct TypeArguments production and is followed by a '(' token, then the sequence of tokens is processed an Arguments production, and any other possible interpretation is discarded. Otherwise, the sequence of tokens is not considered an Arguments production.
When I wrote https://chrismorgan.info/blog/rust-artwork-owl/, which has the title “<_>::v::<_>”, I thought carefully about how it would be handled by feed readers, &c. Fortunately I had already been very careful in the site implementation (e.g. I support HTML in my titles and can strip it or have a different plaintext title, so angle brackets were definitely handled correctly) and had made the deliberate decision that my feed was Atom and uses <title type="html">, so feed readers should all get it right (though doubtless some will ignore the declared semantics and butcher it); if it had been an RSS feed (which doesn’t support declaring the type of a title), some feed readers would have stripped it to “::v::” (and been justified in so doing). Reddit handled it just fine, but maybe it’s just as well I didn’t submit it to HN!
Foo < a , b > x;
This can be either a template installation or a sequence of comparisons.It's kind of a stupid thing to rant over. Whether a language uses <> or [] is the least interesting thing about it, but that doesn't stop everyone from bikeshedding.
auto add(T)(T lhs, T rhs) {
return lhs + rhs;
}
and to instantiate it you do the following: add!int(5, 10);
add!float(5.0f, 10.0f);
// won't compile; Animal doesn't implement +
add!Animal(dog, cat);
Personally, as a C++ adventurist I find the instantiating syntax a bit confusing, but quite tolerable.I wonder whether anyone attempted to use "|T|" as their template syntax...Ruby uses it in iterators (I think) and if I'm not mistaken Rust uses it with closures.
Wouldn't be a bit cleaner to have something like the following?
|T| something (T foo) -> T {
// do something in here
return foo;
}
Personally I prefer it.Ruby and Rust both use them for closures (I assume Rust took them from Ruby). Nearly all of Ruby's iteration constructs are based on closures, so it makes sense that you understood them as for iterators.
Now people are proposing digraphs and trigraphs as alternatives. Have we learned nothing from C?
I'm annoyed that (even in 2020) every programming language restricts its syntax to ASCII characters (1967) which can be typed on an IBM Model F keyboard (1981). Unicode has dozens of unique styles of brackets. Everybody's using an OS that supports Unicode, and an editor/IDE that autocompletes most of their source code anyway. We're even using programming languages which allow Unicode identifiers, so ASCII-only viewers have been in trouble for 25 years already. People are even using emoji in commit messages. That ship has sailed. Unicode is safe to use.
It could be List⟦Int⟧ or List⟬Int⟭ or List「Int」 or dozens of others. They're big, they're clear, they're easy to parse (one char, no other uses). All you need to do is pick one and make it part of the language, and update a few editor modes to support it in some templates.
imagine being a beginner programmer and not even being able to find the character you need to make your program work without learning about Unicode...
Array(1, 2, 3)
someList(0)
array(0) = 23.42
map("name") = "Joe"
All of this is valid MUMPS syntax and does what you'd think.I'd loooove to hear the author's thoughts on the above then extending weird parentheses syntax crap to function calls, MUMPS style too:
USER>s foo(1) = "hi,there"
USER>s $p(foo(1), ",", 2) = "you"
USER>w foo(1)
hi,you
USER>
Maybe I'm warning of a slippery slope here, but this is a direction that I'm not sure you'll want to go down.Also a fun fact, MUMPS has very dynamic typing and thus very much no type annotations.
- https://docs.scala-lang.org/overviews/collections/arrays.htm...
- https://docs.scala-lang.org/tour/generic-classes.html
Unfortunately, scala caused parsing problems by allowing alphabetic function names as operators, grabbing implicit arguments from all over the place, and leaving out () in function calls (I think they reversed the latter decision)
But, you could curry generic parameter application, and then have a simple operator, so it'd be Type ! Param ! Param. Thus:
Map ! Str ! Int a = 5
Not sure I like that.