False Cognates and Syntax: The Importance of Syntax in Programming Languages
txt.io
txt.io
Then I started using CoffeeScript. Now, don't get me wrong; CoffeeScript has lots of little features that are very nice, like concise comprehensions and bound functions, which are not mere details. But if you took away all of that and left only the JavaScript with different syntax details, you would still have something much more pleasant to read and write than JS. -> function syntax, postfix conditionals, unless, on/off/yes/no, @, and "string #{interpolation}" are all tiny "syntax details", and I simply cannot overstress how much nicer they are than the JavaScript they trivially transform into.
is_greater_than(X, Y) ->
if
X>Y ->
true;
true -> % works as an 'else' branch
false
end
(Code from http://www.erlang.org/doc/reference_manual/expressions.html section 7.7)See also: http://erlang.org/pipermail/erlang-questions/2009-January/04...
if x > y -> true; otherwise -> false; end
Using otherwise makes it a loss less ugly, and I guess fixes the false cognate.
case (x>y) of {True -> "Yup"; False -> "Nope");}
or if (x>y) then "Yup" else "Nope"
An example that uses otherwise would be fizzbuzz x | x `mod` 15 == 0 = "Fizzbuzz"
| x `mod` 5 == 0 = "Buzz"
| x `mod` 3 == 0 = "Fizz"
| otherwise = show x -- This is x.to_stringFor instance in Lua:
function abs(x)
test = {
[true]=-x,
[x>0]=x,
[x==0]=0}
return test[true]
end
> print (abs(-3))
3
To understand this code, think that `test[true]` is overwritten at initialization time by the last predicate to evaluate to True.It's not the best way to do this, but maybe it'will help understand the above construct.
test[true]()(defun is-greater-than (x y) (cond ((< x y) T) (T nil)))
As with the erlang example, I need a T test to get an else clause, although in this particular example, the result of the cond is nil if nothing matches anyway.
Technical note: even in languages where every expression has a value, it is possible to have an `if` form that doesn’t require an `else` value: just return nil. See, for example, Clojure: http://clojure.org/special_forms#if
A typical example of a confusing return in haskell is this. Note that this always prints "hi".
foo bar = do
if bar > 5
then plugh
else do
xyzzy
return ()
print "hi"
Here xyzzy and plugh return different types; since the if statement needs the same type on both if branches, return has to be used to return a dummy value of the same type as plugh (here assumed to be ()). But it only returns it to the outside of the if; haskell's return does not influence control flow.You do get used to this pretty quickly, but there's also a tendency to move away from that style of haskell. I'd write the above more like this, using guards rather than the if, and using void to force both plugh and xyzzy to return the same type.
foo bar = go >> print "hi"
where
go
| bar > 5 = void plugh
| otherwise = void xyzzyReturn does influence the control flow in the sense that it signals the end of the monad - it's just that Haskell allows you to define and call a monad ('function') inline, which is not idiomatic in most procedural languages.
It's not too far removed from the following Python code:
> (lambda x: 5)(5)
Which, naturally, returns '5'. Lisp, of course, treats lambdas similarly; however, writing a series of statements in Lisp (like progn) is not considered idiomatic/'good' Lisp, whereas writing monads in Haskell is absolutely necessary.
In [ x.y() for x in X], it's immediately obvious that x.y() is doing some work. Not so obvious in [ x.y for x in X ].
This is a very practical concern - in Django code (both mine and other people's), I see lots of unnecessary SQL queries all over the place because of this.
Of course, Ruby, Scala, etc, are not immune to this criticism.
http://en.wikipedia.org/wiki/Uniform_access_principle
It does put more onus on the API implementer to be careful about hiding non-trivial work, like you say.
The issue being that attributes can be accessed at all of course.
> Of course, Ruby, Scala, etc, are not immune to this criticism.
Well they are — or at least Ruby is — in that they don't allow attribute accesses from third-parties at all. Just consider that Python is the same (it is).
Hell, in Python `.` is already a method call. In fact it's a whole sequence of method calls.
Python is different because the style guide explicitly discourages [1] computationally expensive accessor methods (as well as accessor methods with side effects) specifically so that programmers can treat accessor methods the same way they would treat a data field. Assuming that x.y is a field access in Python is not supposed to lead to problems, and if it does, it's the fault of the class implementer and not the user.
[1] http://www.python.org/dev/peps/pep-0008/#designing-for-inher...
Yes it is - that's always a method call: http://docs.python.org/reference/datamodel.html#invoking-des...
The problem is that you're unclear whether it's a mutating method or not, as well as whether it's an expensive operation. But that can be solved a number of ways - immutability and lazy evaluation would be one approach, though unfortunately neither Python nor Ruby enforces immutability, and both use absurdly eager evaluation.
(Regarding that last point, try doing [x.y() for x in foo][0] and you'll see that y is called for every x!)
The else clause is also mandatory in C, java, ... in if-expressions:
int foo = <test> ? <then-clause> : <else-clause>;
They even have to have the same type!I personally do not feel like this is squinting and rotating your head 47 degrees: if you look at simple Haskell code examples the "return" function is used in the same way as it would at the end of any normal C function. The author of this article sees it that way, but that is just an opinion in not backed up in the argument by the "false cognates" premise.
On the other hand, I think the article's point that it's tricky is valid. The way it is usually used makes it appear that it is causing the procedure being defined to exit. If a user thinks that's what it's actually doing (which they tend to do when coming from other languages), they will try to write things like "when (i == 0) (return x)", which does not mean what they think it means.
In philology / historical linguistics, "cognate" is a term relating strictly to etymological origins (in which context English and German "gift" really are true cognates). In typical middle/high-school foreign language classes, though, "cognate" is usually used in the related-but-rather-different sense of a word that ought to be easy to remember for the vocab test because it's sound and meaning are both similar to a word in your own language, and a false cognate is a word that looks or sounds familiar but means the wrong thing (in which context English and German "gift" would be considered "false cognates" for pedagogical purposes). Most true cognates in the technical sense would never be presented as cognates to a beginning foreign language class.
do if condition
then do putStrLn "bailing out early"
return ()
else putStrLn "carrying on"
putStrLn "Launching missiles..."
This does not do what a C/Java/C# programmer would expect.`return` in Haskell is a false cognate and it causes beginners difficulties. Some of them post to StackOverflow asking confused questions about it.
if-then-else in Haskell is also a false cognate (and there are formatting issues to trip up the unwary as well). It's actually the same as C's/Java's/etc ?: operator, so it's unfortunate the standard library doesn't contain something similar to that instead of adding if-then-else to the language.
Having said all that, I don't think being a false cognate should automatically be a disqualifying attribute. Haskell's new <> operator (a synonym for mappend) looks like Basic's/Pascal's not-equal-to operator. But I think few Haskellers come from that immediate background.
Anyway, for me, a good example of this phenomenon is when people come to C++ from a language like Java. In Java, you have to use "new" every time you want to create an object, so these people tend to go around putting "new" everywhere.
"new" simply has a different meaning in Java even though the syntax is similar. If I were designing Java, I'd leave out the "new" keyword altogether. I see it as unnecessary to the semantics and confusing to C++ programmers, but that's just me. =)
I agree that the post was a plea to language designers, but such a plea is probably pretty useless and I didn't find it worth discussing. However, it can be fun to relate to his frustrations by sharing a story from your own experience.
Are you a Java developer? What's the most annoying behavior you see coming from converted C++ developers?
As for me, I don't call myself a [insert language]-developer. I use whatever language suits best to design software. So far I have experience in many common procedural, OO and markup languages (C,C++,Java,Obj C,CUDA C,bash,python,VB,Matlab,Lotusscript,html+css) and I try to get experience in functional languages as soon as I can get some time for that. What's characteristic for me in Java is mainly its VM architecture. This makes it useful when you need the flexibility of a 3rd generation language for easily portable code (e.g. many business applications), however it has some disadvantages that have hindered its success for consumer applications. The main disadvantages IMO are the non-native feel of the GUI, maintenance and compatibility problems of the Java VM and Oracle's business strategy as of late. Java's syntax is just fine, I don't have any grudges about it. Knowing Obj C well, I know what ugly syntax looks like, however I still like that language as well (performance and frameworks are quite good, especially compared to other mobile frameworks).
A good example I can think of right now is Imaginary numbers and complex numbers. The legendary Gauss said they would have been better named Lateral Numbers.
But for programming languages avoiding false cognates is even harder as they not just naming things, but also deciding structure and must draw from the same pool of words as hundreds of other languages.
You can use sequences of punctuation characters, e.g. <:= and =:> for brackets, or =>> for an operator, or :*; for a separator. Or use punctuation characters not in ASCII: there's hundreds of them in the symbols blocks of Unicode.
The unwillingness of whiners is the number one reason for software to suck. Why does a C11 compiler still compile Ansi C? It is utter crap, and whiners are to blame.
edit: Conservativeness is the reason young developers don't learn C, why C# beat Java and why everyone will have to learn a new language every few years to keep on top of the curve.
If you see a flaw in a language you use, do you report it, contribute to a thread, or do you just learn to live with it?
(sorry I forgot to mention I had added 2 paragraphs)
(edit: The author of the comment I responded to added another two paragraphs.) C# beating Java is not something I think is as clear-cut as you seem to believe (from my vantage, C# is used by Windows developers, and Java is used by everyone else: the dividing line has little to do with syntax preferences and much more to do with the quality of the IDEs and the integration of the runtimes; before MS was legally required to stop distributing it, J++ was gaining ground), but to the extent to which it is true you have to remember that C# also is not making breaking language changes, and in fact didn't even make drastic changes to Java (which one might argue is the legacy that C# continued). In fact, C# was so compelling to a lot of us who adopted it early because it had such amazing backwards-compatibility in the form of interop with native C and C++ code with the ability to nearly natively interact with our existing COM objects.
If a compiler removes a wart, why wouldn't you s/wart/scar/? Perhaps it's a bit of a pain every compiler release, but in the end isn't your software a bit more future-proof?
There is no need to stop making progress, this is why good software is open source and on github.
New generations of whiners are better judges of what is and what is not useable than old farts. The new generation has a glimpse of the future.
edit: I don't think native code interop is backwards compatibility, that would imply C/C++ is somehow a thing of the past. My point was that as long as C/C++ are being adapted to the future, they will never be legacy and retain relevance where their use is warranted. C# perhaps is not a winner in adoption, but it is gaining ground on linux and osx. My observation that C# beats Java is more in the syntax area.
As for your comment about open-source... that doesn't help the problem of "all the developers are spending their time updating and re-testing working code rather than writing new code that relies on it", and specifically looking at GitHub makes no sense. Finally, the "old farts" you are now denigrating have more knowledge and experience of the kinds of failures you will run into, and so far in my experience the people rebuilding systems end up learning, the hard way, the same lessons that they could have just inherited (a rather visceral lesson as I've watched my friend Yehuda work on Bundler over the years, as he got to rediscover all of the things that APT solved over a decade ago).
The key is that the examples it flashes exactly try to cover the code mistakes/syntax errors (even things that don't compile!) that you would normally create in the first few minutes or hours - before you're 'in the zone' when coming from one of the other languages in the list. Things like leaving off a semicolon when coming from Python to C++. You can be an experienced C++ programmer, but after long hours of python, you need a period of adjustment. Isn't this the most dangerous time to code?
So if you put that you want to get in the zone with C++, but in your profile Python is listed, then some of the examples will be missing a semicolon while being indented properly. If you put a language that uses eq instead of ==, . instead of +, this is brought up a couple of times. All the things that separate languages - so that in the first few error-prone hours of transition you leave them out or forget them at times - are brought right to the front so that you can produce much higher-quality code in your 'target language' after 'zoning up.' What do you guys think?
Pedagogically, it may be better not to flash incorrect code, but instead make you write the code. But ask you to write code such that it explicitly tests something that may trip you up coming from one of your other languages.