Why Does “=” Mean Assignment?
hillelwayne.com
hillelwayne.com
In the mid 60's, when I started programming, most programs had to be keypunched. (There was paper tape entry and Dartmouth's new timesharing system using BASIC, but these weren't in very widespread use.) Until 1964, the keypunch machine in use (IBM 026) had a very limited character set. This is the reason that FORTRAN, COLBOL, and LISP programs were written in upper case. Even the fastest "super" computer of the time, the CDC 6600, was limited to 6-bit characters and didn't have lower case letters.
Naturally, symbols like "←" or "⇐" weren't available, but even the characters ":" and "<" were not on the 026 keypunch keyboard and so ":=" or "<-" were much harder to punch on a card.
These early hardware limitations influenced the design of the early languages, and the early programmers all became accustomed to using "=" (or rarely ":=" or SET) as assignment even though "⇐" might have been more logical. The designers of subsequent programming languages were themselves programmers of earlier languages so most simply continued the tradition.
This explains why early ALGOL implementations (e.g. Burroughs, the only dwarf who went all in on the language) and dialects (ALGO, JOVIAL, MAD, NELIAC) actually used ‘=’ for assignment. This was perfectly legal according to the language definition, which distinguished between the ‘reference language’, which used ‘:=’, and its ‘hardware representation’, which could be anything. (Naturally there was no portability.)
The limitations define the style. And then dogma perpetuates it, as the new generation continues to live with the results of once-great but now outdated thinking.
Take Erlang for example:
1> X = 1.
1
2> X = 2.
** exception error: no match of right hand side value 2
Notice variables are immutable (not just values themselves). Once X becomes 1, it can only match with 1 after that. You might think this is silly or annoying, why not just allow reassignment and instead have to sprinkle X1, X2 everywhere. But it turns out is can be nice because it makes state updates very explicit. In complicated applications that helps understand what is happening. And it behaves like you'd expect in math, in the sense that X = X + 1 doesn't make sense here either: 1> X = 1.
1
2> X = X + 1.
** exception error: no match of right hand side value 2
3>
It does pattern matching very well too, that is, it matches based on the shape of data: 1> {X,Y} = {1,2}.
{1,2}
2> X.
1
3> Y.
2
4>
In other languages we might say we have assignment and destructuring but here it is rather simple it's just pattern matching. X = Y,
X = 5,
%% Y = 5 now as well
-- ?- append([1,2], Y, [1,2,3]).
Y = [3]. ?- append(X, Y, [1,2,3]).
X = [],
Y = [1, 2, 3] ;
X = [1],
Y = [2, 3] ;
X = [1, 2],
Y = [3] ;
X = [1, 2, 3],
Y = [] ;
false.
(a few newlines added for clarity)Since the equal sign is an assertion of truth, and since assertions cause a crash on failure, you can pretty much derive the rest of the language/BEAM characteristics from that.
I like the '=' as an assertion idea. That is, it's like running code with assertions turned on. Never thought about it that way but it makes good sense.
Also stealing your idea of "rebootable code" :-)
I almost submitted the talk to That Conference this year. I should start shopping it around more.
Also, there's a simple counterexample. You can have static single assignment that has dynamic multiple assignment (e.g. within a loop construct). Under the pattern-matching semantics, evaluating the assignment (or rather the pattern match) a second time could fail, whereas an assignment always succeeds.
What is SSA here?
iex(1)> x = 1 1 iex(2)> x = 2 2 iex(3)> ^x = 3 (MatchError) no match of right hand side value: 3
Notice the pin operator (^) to achieve the same effect.
2 = X.
{ok, Y} = X = foo().
Oh, and here's a fun one: 1> {X, X} = {1, 1}.
{1, 1}.
2> {Y, Y} = {1, 2}.
** exception error: no match of right hand side value {1,2}
If you're wondering what the use of that is, here's an Erlang "drop all occurrences of an element X from a list" function. (Prerequisite knowledge for this: Erlang functions have several clause-heads; which one is executed on each call depends on which one is able to successfully bind to—i.e. = with—the arguments.) drop(X, List) -> drop(X, List, []).
drop(X, [], Acc) -> lists:reverse(Acc);
drop(X, [X|Rest], Acc) -> drop(X, Rest, Acc);
drop(X, [Y|Rest], Acc) -> drop(X, Rest, [Y|Acc]).
In other words—first, set up an accumulator. Then, go through the list, and if X can bind in both positions, skip the element; otherwise (i.e. if Y != X) then shift Y into the accumulator. At the end, reverse the accumulator (because you were shifting.) drop(X, List, Result) :- drop(X, List, [], Result).
drop(_, [], Acc, Result) :- reverse(Acc, Result).
drop(X, [X|Rest], Acc, Result) :- drop(X, Rest, Acc, Result).
drop(X, [Y|Rest], Acc, Result) :- dif(X, Y), drop(X, Rest, [Y|Acc], Result).
I only had to make a few syntactic changes to the original program to obtain a Prolog predicate from it. The most significant change is that I use an additional argument to hold the original function's return value. An important consequence of the resulting relational nature is that this can answer quite general queries.For example, we can ask: What are possible solutions Ls that arise when we drop X from the single-element list [A] ?
?- drop(X, [A], Ls).
X = A, Ls = [] ;
dif(X, A), Ls = [A].
We see that there are two possible answers: One where X is A, and hence the result is the empty list []. And the other where X is different from A, as indicated by dif(X, A).Even more generally, we can systematically enumerate all conceivable solutions for all list lengths:
?- length(Ls0, _), drop(X, Ls0, Ls).
Ls0 = Ls, Ls = [] ;
Ls0 = [X],
Ls = [] ;
Ls0 = Ls, Ls = [_586],
dif(X, _586) ;
Ls0 = [X, X],
Ls = [] ;
Ls0 = [X, _598],
Ls = [_598],
dif(X, _598) ;
etc.
This uses iterative deepening to generate all possible answers. ?- drop(X, [a,b,c], [a,c]).
X = b ;
false.
As another example, if both Ls0 and Ls are [a,b,c]: ?- drop(X, [a,b,c], [a,b,c]).
dif(X, c),
dif(X, b),
dif(X, a).
This means that X must be different from each of these elements.More generally, which elements can be dropped at all from the list [a,b,c]:
?- drop(X, [a,b,c], _).
X = a ;
X = b ;
X = c ;
dif(X, c),
dif(X, b),
dif(X, a).
Interestingly, the predicate definition does not even use "=" in its clauses. In Prolog, (=)/2 is a built-in predicate that means unification. In the definition above, we use implicit unification instead of explicit unification.On the other hand something like
drop _ [] = []
drop x (y:xs)
| x == y = drop x xs
| otherwise = y : drop x xs
is often more readable at a glance to me since I don't have to parse variable names to see control flow structures.Also, in statically typed languages what happens if equality doesn't work for a value.
I'm not familiar with this term "destructuring." Can you elaborate on the concept? Might you have an example of a destructure operation and a language where its's used?
> let {a,b} = {a: 1, b: 2}
> a
1
> b
2Basically it means that you can assign several variables at once, by asssigning a complex value to an aproprietly structured literal with variables as placeholders.
In pseudojavascript, let's say you got this from some API:
var someList = [1, 2, 3, 4];
var someMap = { name: "John Smith",
age: 20,
cars: [{ id: "AAA-1111", model: "Ford T"},
{ id: "AAA-1112", model: "Mercedes Benz"}] }
And you want to extract useful data for your code quickly. In languages with destructuring you can do sth like this: var [x, y & rest] = someList; //x==1, y==2, rest == [3, 4]
var { "name": name,
"age": age,
"cars": [ {"id": firstCarId, "model": firstCarModel}
& restOfTheCars ]
} = someMap;
And the language will put the correct values in your variables name, age, firstCarId, firstCarModl, restOfTheCars. > let a = [1,2,3]
undefined
> [x,y,z] = a
Array(3) [ 1, 2, 3 ]
> x
1
> y
2
> z
3In nearly every other language, when you call a function, the top of the function typically involves a fair bit of analysis of the arguments to decide what needs to be done.
(That's being generous: typically the entire body of the function revolves around making decisions about what to do, such that you don't know what's going to come out of the function until you've traversed multiple branch options.)
In well-written Erlang, the function itself branches at the top, taking advantage of pattern matching, destructuring, and function heads.
Let's say you're passing around arbitrary data in the form of a tuple that you need to display. Or not, depending on the value of some debug flag.
In Python you might see code like this, assuming that an integer is stored in your program like ('int', 3) and a string might have some arbitrary label, so it's ('str', 'label', 'real string'):
if (debug):
display(x)
def display(x):
if (x[0] == 'int'):
display_int(x[1])
elif (x[0] == 'str'):
display_string(x[1], x[2])
In Erlang, you can branch and destructure at the function head. You don't check the debug flag in the calling code, just pass it along for the function to decide what to do with it; if the flag is false, the function just returns false. (Variables in Erlang are capitalized; the lower-case "strings" in this code are atoms, also called symbols in some languages.) display(_, false) ->
false;
display({int, N}, _) ->
display_int(N);
display({str, Label, String}, _) ->
display_string(Label, String).
Since you can choose the key data elements in your function arguments and branch accordingly in the function head, it becomes immediately obvious looking at the code how many different ways the code can evaluate, and the destructuring in the function heads means you don't have to waste code chunking the data apart before working with it.(Updated: forgot to include the debug flag in the 2nd and 3rd function heads. In all 3 heads, the _ pseudo-variable basically says "I don't care what value we have here". Another option would be to use _Debug in the 2nd and 3rd heads to indicate to the reader what the flag is, or use true since we're expecting that value.)
if you want to calculate the output of a function into a data table(which is what early computers were often doing) iterating a variable x over a f(x) = xn + b(with n and b being fixed constants) is exactly how you would do it on paper so it's probably a case of computers emulating applied math rather then simply being pure theory machines.
Here, you would write x' = x + 1 to give the recurrence relation x[n+1] = x[n] + 1. Or, more generally x' = f(x) for x[n+1] = f(x[n]). The reason for this is because in mathematics x = x + 1 is absurd (ignoring modulo arithmetic).
I think language designers were certainly consciously aware that they were designing something well defined, and chose to use this kind of syntax because its simpler, rather than an implicit mapping to the natural numbers via something like CPU cycles.
See Lucid[1]
>takes away the only thing that makes it a sequence
Unless everything is a sequence[1]
[1] http://www.cse.unsw.edu.au/~plaice/archive/WWW/1985/B-AP85-L...
We then have the sequence $x_n+1 = f(x_n)$. Often, when dealing with incremental algorithms, the notation x' = f(x) is used. Here x' (pronounced x prime) stands for "the next value of x". It's a nice balance between the correctness of using indices and the conciseness of leaving them out.
Going to a higher level, the sequence x' = f(x) is essentially trying to find a fixed point of the function f. To look at this in an actual for or while loop, you need to consider the stopping condition of the loop as part of the function.
A. An arithmetic formula is a variable (subscripted, or not), followed by an equals sign, followed by an expression.
B. It should be noted that the equals sign in an arithmetic formula has the significance of “replace”. In effect, therefore, the meaning of an arithmetic formula is as follows: Evaluate the expression on the right and substitute this value as the value of the variable on the left.
> A notorious example for a bad idea was the choice of the equal sign to denote assignment. It goes back to Fortran in 1957 and has blindly been copied by armies of language designers. Why is it a bad idea? Because it overthrows a century old tradition to let “=” denote a comparison for equality, a predicate which is either true or false. But Fortran made it to mean assignment, the enforcing of equality. In this case, the operands are on unequal footing: The left operand (a variable) is to be made equal to the right operand (an expression). x = y does not mean the same thing as y = x. Algol corrected this mistake by the simple solution: Let assignment be denoted by “:=”.
> Perhaps this may appear as nitpicking to programmers who got used to the equal sign meaning assignment. But mixing up assignment and comparison is a truly bad idea, because it requires that another symbol be used for what traditionally was expressed by the equal sign. Comparison for equality became denoted by the two characters “==” (first in C). This is a consequence of the ugly kind, and it gave rise to similar bad ideas using “++”, “--“, “&&” etc.
From Good ideas, through the Looking Glass by N. Wirth:
https://www.semanticscholar.org/paper/Good-Ideas%2C-through-...
And if the goal of mathematicians and programmers were the same, this might be a compelling argument. In no sense do I see my programs as giant algorithms, they're machines, and as such require different syntax to construct.
> But mixing up assignment and comparison is a truly bad idea
Then := is also a bad decision. You should use the fully historically supported form of 'let <variable> = <value>'. Both := and == suffer from the fact that if you omit a single character, you change assignment into comparison or vice versa.
It strikes me as a nitpicky argument for it's own sake.
But then we get to this line:
> Because it overthrows a century old tradition to let “=” denote a comparison for equality
In what sense is "=" for comparison a century old? Hasn't it been used for "things that are equal" for multiple centuries? If so, and if that's distinct from "for comparison", then isn't for comparison also a new, non-standard use?
Does anyone understand what he's on about in this quote?
It also, in science and engineering, denotes a method for obtaining lhs when you have values for the terms on rhs. That is the point of converting y = ax² + bx + c to x = (-b ± √(b² − 4ac)) / 2a. Early languages like Fortran explicitly followed that usage.
Coincidentally in async languages like Verilog there is a another assignment operator '<=' which means that the variable won't take the new value until the next clock cycle (or more explicitly the next evaluation of the process block). '=' exists also and has the same meaning as with traditional languages.
a <= 1
Should set a to the minimum of a and 1. I.e. after that statement the variable will indeed be less than or equal to the RHS.Ken Thompson on why '=' is assignment and '==' the equality check.
It's this kind of mindset that puts me off Go. But I can totally see that many who want a better C getting into Go exactly for reasons like this.
If you declare a variable, then x = 2; works in Go.
Syntax that looks pretty in isolation is a poor design guideline for programming language design. It ignores the problems of writing and debugging this code. I don't know how many bugs went unnoticed because of the "if (a = b)" mistake but it certainly weren't few. Sure, nowadays the compiler warns you about that but that took surprisingly long to be implemented.
The argument for saving a few keystrokes for potential longer debugging sessions is not a good argument. Code is much more often read than it is written.
But I didn't criticize Go's verbosity (after all, they "fixed" the assignment thin) but the mindset of doing or not doing things for specific reasons that look quite backwards for today (or let me be frank, just plain stupid) but the community then fights with vigorously for this. There are examples for that in the design of the language but the most egregious example of this is the package manager. This went from "we don't needs this" over "do it in this problematic way" over "let's do it like anybody else" to now "no, we are special, we need to do it completely differently".
IMO, Go would have been a fantastic language to have in the 90s. But looking at it from today's perspective, it looks outdated in many places. But compared to C which is a language of the 70s it is still great and therefore I understand its appeal for programmers who haven't found another language to replace C with (going from my own experience, most programmers have replaced C with multiple languages instead of just one).
Since in procedural language, assignment is a much more common operation than equality check, it is reasonable to favor "=" over ":=" or even "<-".
x: 2
These 'accessors' also work on local variables, there is no special syntax for assignment. So when you write x it means send the message x, which starts its lookup in the local environment and works itself outwards. When you write x:2 it sends the message x: with argument 2, also starting in the local environment.
The accessors are automatically generated in pairs for slots. It sorta works, but it seems a bit too convoluted just to say "hey, we can do everything with just messaging".
Fun fact: various dialects of BASIC let you optionally put LET before your assignment statements. So the following are equivalent:
X = 10
LET X = 10As someone who learned first on a ZX81, I'd always assumed that "LET X=5" was the canonical form, and that "X=5" for assignment was just a shortening for convenience.
I wonder if you, like me, were also completely thrown when you first encountered a programming language where you didn't also have to enter line numbers...
Haha, yes, totally! For me this was AMOS on the Amiga (a Basic variant). I was so thrown I started by using them anyway as they were supported by the interpreter - it just treated them as labels.
I soon stopped though when I discovered it didn't reorder based on number, so you ended up with
10 print "world"
5 print "hello"
20 goto 5 018 LET P=P K * B -
and 135 LET Q=1 F G / 1 - REC 0.00001 * - Q *
but also the sort-of traditional 016 FOR X=1 STEP 1 UNTIL T
I guess they had a Basic that only allowed expressions with a single operator or a single function call (some early systems with very limited memory did that) to the right of =, and didn’t have the memory for an infix parser.> How can a = a + 1? That’s like saying 1 = 2.
It looks like a lot of us don't have (never had?) trouble accepting a statement like `a = a + 1` and, upon stepping back, thought maybe it had to do with our internal dialog (which is now probably intuition as we've been processing statements like this for so long).
But yes, most compilers will catch this case.
while (v = next()) { ... }
or, perhaps more usefully: while ((v = next()) != ERROR) { ... }
EDIT: just did a quick check, and LLVM did in fact warn me about this: if (x = 4) { ... }
To silence the error, you have to double up on the parenthesis, which I guess is a reasonable enough solution: if ((x = 4)) { ... } # bad
if v = array.grep(/foo/)
do_something(v)
# some code
end
# good
if (v = array.grep(/foo/))
do_something(v)
# some code
end
Caught a number of nasty bugs that way.[0]: https://github.com/bbatsov/ruby-style-guide#safe-assignment-...
if (2 == x)
int x = 10 int y = x
would read in English as x gets 10, y gets x.
Shortly after that class I took a break from any coding. When I went to another school I saw Softmore and Junior level programmers still struggling with this. I retained the habit of using 'gets'and never had a problem with this.
Not sure why/how I picked it up, but similar principle, I suppose.
EDIT: It's since been changed to "let", "set", "equal" (but still lower-case).
I think that Lisp doesn't even fit well into that table, since putting LET in the first column would implicitly bring in declaration as well, much of the time SET in the second column would be misleading since mutation is so often done implicitly by a looping construct like DO, DOLIST, DOTIMES or whatever (no SET in sight), and there are all sorts of equality tests one could use.
Edit: ok, I see the author has both updated it, and is specifically talking about Lisp 1.5 which does not have the looping constructs. On the other hand, it at least has a bunch of alternatives to SET that probably should be listed as well, like RPLACA, RPLACD, ATTRIB, and so on. I think that even in Lisp 1.5 variables could be declared in different ways to using LET (like rebinding a variable passed in to a function, so the assignment was in invocation).
It just happens that there was no consensus yet on how it should work, yet, either!
Then again, neither of the last 2 (modern) programming languages I've used have had FORTRAN/ALGOL-style assignment semantics, so I think the jury's still out.
The operator " = >" (pronounced "heffalump") is
convenient for referencing structures that are
accessed indirectly. The expression
a=>s.x is equivalent to the expression (@a)»s.x.
Source: https://archive.org/stream/bitsavers_xeroxaltobualSep79_5187...≔ U+2254 COLON EQUALS
and there is also
⩵ U+2A75 TWO CONSECUTIVE EQUALS SIGNS
if you find typing two = characters fatiguing...
I was told once by a math professor that it is a habit inherited because we use Arabic numbers/maths which were really meant to be read from right to left. Don’t know if the theory has any merit.
Imperative programming corresponds to the imperative moodin English, where English normally has verb-object word order with the subject (the entity being commanded) ommitted. The subject of the command to set the value of x to the result of the addition of two and two is the computer/runtime running the code, not the variable x which is the direct object.
English has SVO order for declarative sentences, which correspond to declarative programmig, which tends to feature definition or binding rather than mutating assignment.
I sort of disagree with this. Many functional languages pull heavily from lambda calculus and other forms of mathematics. In math, "a = a + 1" isn't the same as "1 = 2". The issue isn't equality, it's that you're trying to rebind a bound variable, which isn't possible.
In other words, rebinding a bound variable is not the same as "1 = 2".
"=" means equality in math; a = a + 1 is the same as 1 = 2 because if you subtract a from both sides and add 1 to both sides you get 1 = 2.
Lambda calculus has the concept of binding variables, but it doesn't use "=" for that, it uses application of lambda forms. It's the same idea that's applied in some variants of Lisp, where LET is a macro such that (let ((x 1) (y 2)) ...) expands to ((lambda (x y) ...) 1 2).
The way it plays out is that rebinding is perfectly fine, because it's not really any different from binding in the first place. The same way that
(let ((x 10))
(f x)
(setf x (1+ x))
(g x))
can be re-written as ((lambda (x)
(f x)
(setf x (1+ x))
(g x))
10)
That can itself be re-written as: ((lambda (x)
(f x)
((lambda (x)
(g x))
(1+ x)))
10)
If you'd like to read more about this sort of thing, Sussman and Steele's "Lambda: The Ultimate Imperative" is a good starter: http://repository.readscheme.org/ftp/papers/ai-lab-pubs/AIM-...Rebinding is also needed for simple composition of operations. Given some f(x), the formula f(3) + f(4) involves x being simultaneously bound to 3 and 4 in different instances of the scope inside f.
Functions cannot "work" in math if arguments cannot simultaneously be bound to different values.
Recursion is not possible, and so defining computation in terms of recursion goes out the window.
Therefore using "=" to mean "bind variable" is a notation mismatch, which is the whole point of what you quoted, no?
Wouldn't it be easier to use = for everything and get rid of those weird edge cases?
Although even then it's nice to use different symbols because they are different meanings. I don't like it when a word has different meanings depending on context.
I don't think that's enough. Take the two following Python lines:
a = b = c == d
a = b == c == d
No expression is being used as a statement, yet you can't syntactically separate assignment from equality.But I agree with your conclusion, that statements should not also be expressions, so a = b = c shouldn't work (at least not the way we're used to, that the b = c is an assignment "statement" that also produces an expression value).
But in the end all this does is allow "a = b = c" to be a nonambiguous statement meaning, in more familiar notation, "a = b == c". Not exactly clear! So even though a language could use the same symbol for both assignment and equality-check, I don't recommend it! Although I still like to keep statements and expressions as strictly non-interchangeable constructs.
if x=y # obviously compare x=y # obviously assignment x=y=z # is that really needed? If so, brackets
a = b = (c == d)
and a = (b == c == d)
respectively. And I think it would make sense to have a syntax rule that strictly enforces parentheses around truthy expressions used in such contexts.A complex destructuring case would be something like this:
{cashMoney -> foo, stuff: {powerLevel -> bar} <- {cashMoney: 5000, stuff: {powerLevel: 8999}}
I think the arrows nicely indicate where data is flowing, and it's unambiguous what is a variable name and what is a field name.
Funnily, when Thompson was asked what he would do differently if he were doing it over again, he said, "I'd spell creat with an e."
Whereas Java is all about protecting developers from themselves, Ruby (for example) let's you get away without variable type declaration because at the end of the day, how often do you not know whether a particular variable is a string or an integer? Enough to justify enforcing declaring everything at the outset or receiving errors throughout your code?
That's not to say that those enforcements never make sense. I think that structured languages like Java are better for larger teams where enforcing standards is more important than a smaller team. But there are other ways to do that and a lot of time the rules make development a horrible experience.
I'm refactoring some data-pasta to more clearly declare types because we just have dicts of lists of whatever, so often enough that a random reader chimed in after 20 minutes.
In someone else's code? 100% of the time. I was once a diehard Ruby believer, and I still use it for small programs. But having had the singular displeasure of working on a huge Ruby codebase (pre-JVM Twitter), I have learned, in the hardest possible way, that untyped languages are absolute nightmares when more than one person is involved.
From a language, library, or framework I want to be in and out as fast as possibly. I want to have unambiguous usage that is easily discoverable. I want to avoid having to load a million things into my memory -- not least because I work with a huge pile of legacy code that all uses different technology. Rails is not the worst system I've ever used for this, but it's edging up there. Ruby, as a language, I don't mind at all, but the coding convention of people who primarily use Rails is something that I think can be improved dramatically.
Of course, IMHO ;-)
A nontrivial proportion of people who start trying to learn to program never get past that point, so it's worth taking them into account.
> at the end of the day, how often do you not know whether a particular variable is a string or an integer?
Strawman. Types are not about the difference between a string and an integer, they're about the difference between a user id and an order id, or a non-empty list and a possibly-empty list, or...
Can you explain - how is the difference between a string and an integer different from the difference between a user id and an order id? You're calling the OPs comment a straw man but I'm not understanding how your examples are not the same thing given type system that understood all four of those.
But really, this is such a pedantic discussion, trying really hard to not get sucked in.
a <- a + 1
a + 1 -> a # also works, but REALLY bad practice
But, so does
a = a + 1
Granted, there are a bunch of R haters (especially from people with formal CS educations), I think this convention makes a lot of sense. While most will disagree about the '<-' I like it from a code reading sense in that you know that it is an assignment right away. Coming from a mathematical perspective before learning to code, this makes a lot more sense in the 'assignment' fashion.
In case you are wondering, the difference between <- and = in R is in scoping. For example, in the following function call:
foo(x = 'value')
x is declared in the scope of the function, whereas:
foo(x <- 'value')
x is now declared in the user environment. Granted, that is not good practice, but that is why there is a difference.
for example, this assigns a ggplot to 'plot': df %>% na.omit() %>% ggplot(aes(x=x, y=y)) + geom_line() -> plot
That is really confusing in that the way most people would read it is that it is something to be plotted. However, the assignment does occur and is masked. having 'plot <- df %>%' as the first line makes it clear that a new object is being created.
We actually had to modify our style guide to prevent the '->'
To be clear, the assignment syntax is great, it's other things that are the target of R haters' hate! Like a high-level language without hash tables.
https://en.wikipedia.org/wiki/Relational_operator#Confusion_...
but Wikipedia says: "The reason for all this being unknown. [footnote] Although Dennis Ritchie has suggested that this may have had to do with "economy of typing" as updates of variables may be more frequent than comparisons in certain types of programs"
while the OP says: "As Thompson put it:
Since assignment is about twice as frequent as equality testing in typical programs, it’s appropriate that the operator be half as long."
Thompson and Ritchie designed C together.
Can anyone explain to me what this is supposed to mean? (I know Prolog, I understand cuts very well. But I don't know what the author means here.)
Fun fact: If (=)/2 were not available as a built-in predicate in Prolog, you could define it by a single fact:
X = X.
Thus, you do not even need this predicate as a predefined part of the language. PIP destination←source /switches
became this: PIP destination=source
or this: PIP destination_source
https://en.wikipedia.org/wiki/Peripheral_Interchange_ProgramWhen it reads as "let x equal y" it makes sense. I think it makes good sense for imperative programming. It makes slightly less sense for functional programming. For declarative configuration, I think a colon makes much more sense.
This topic has my curiosity peaked as well. What is a good place to start reading/understanding more about this?
From my understanding software development is a complex field that blew up in multiple places at the same time.
At the same time when reading
> How can a = a + 1? That’s like saying 1 = 2.
I think: Well, depends how you interpret it.
If the translation is „we state, that from now on, a is the previous value of a plus 1“ it’s totally ok.
Not a big difference from saying
a = b + c
So i‘m not sure if the usefulness if changing all languages to using := is that high that it’s worth to think about changing it in current languages - and even when inventing a new language I’m not sure if it’s helpful.
I also wonder if reassignment, beside counters that need to be changed with iteration/ appearance of the event to be counted, is a thing that should generally be avoided if possible. Ok maybe in general everything where data is transformed in iterations...
Drop updates the logical structure without necessarily acting on the physical storage in a way commensurate with "deletion". You can this drop a million tuple table in a ms (less I expect) whilst deletion would take far, far longer (of the order of a million-times longer).
/guessing
DROP doesn't mean DELETE.
DROP <object-type> <object-name> is sort of like a macro for, loosely
DELETE FROM <catalog-relvar-for-object-type> WHERE name = <object-name>
Except that real RDBMSs don't usually have DDL that is really equivalent to DML against system tables, and particularly (esp historically, but still in many years DBs), DDL has a different relation to transaction processing than DML, so it's a very good thing for clarity and developer intuition to not overload DML keywords for DDL operations, despite the loose similarity between CREATE/DROP in DDL and INSERT/DELETE in DML.
So this is a dilemma I have while working on a new language. I'd like to go with `:=`, but `=` is absurdly popular, and I'm trying to keep the language as approachable as possible.
I don't think the clarity of `:=` is so compelling that it outweighs the `ew, why are there colons in there` reaction that I think most novice coders would have.
(Remove redundant commentary about database stuffs.)
My advice: don't allow assignment in expressions. To me, it's like the case-sensitive issue: the language designers think it's a useful feature, but it actually works against most developers.
I don't think case-folding identifiers is helpful. The language has decreed fooBar is the same as foobar, and that handles the error where you spelled the same idea two different ways, but it fails silently on the error where you spelled two different things a similar way. Worse, there are some people who are very sensitive to case and will be confused, while others will happily type their entire code in all caps.
I think a linter is the best way to catch these issues, and those subjective rules are precisely the sort of thing that need to develop more rapidly than the core parser.
Conceptually, what is the difference between these two identifiers:
myObjectInstance
MyObjectInstance
?
And the key here is the reason for the difference: if it's a typo, then a case-insensitive language design will allow it and no-harm, no-foul. If it's not a typo, then who wants to work on a codebase littered with identifiers whose only difference is case ? :-)
In Haskell, one is a variable, the other is a type, and that's enforced by the language. It's the same, albeit by convention, in Java. There are a lot of cases where you want to describe a type and a thing, so apple = new Apple() is pretty reasonable.
When I think of case-insensitive languages, I'm thinking of Basic, LISP, SQL, and those don't have a lot of type declarations.
And consider two counter-examples:
my_instance vs myinstance
things vs THINGS
The first shows case-folding is only a partial answer to ambiguous identifiers. The second shows that differences in case can be very obvious to the reader.Those are motivators to me for pushing this off to the linter: there are a lot of subjective judgements in what should and shouldn't be the same, and having the language keep its rules for identifiers as simple possible seems like a good separation of concerns.
My final concern is metaprogramming and interoperability. In SQL, for instance, there are bizarre rules to work around case-insensitive identifiers. If another system asks you for "myObjectInstance" and "MyObjectInstance", it has to know your case folding rules to know those two identifiers are the same.
> If it's not a typo, then who wants to work on a codebase littered with identifiers whose only difference is case ? :-)
Ever worked on a Python project that interacts with Javascript, so it's snake and camel case?
I generally agree, I'd just prefer a gofmt-style utility that would just automatically resolve those and tidy everything up. I completely agree that just chucking error messages is a poor answer.
Finally, here's a challenge, if identifiers are going to be folded by the compiler: what locale should be used? In particular, how do you handle I, İ, i and ı?
Re: languages: Pascal/Object Pascal is case-insensitive, and is statically-typed.
Re: SQL: all implementations that I'm aware of use case-insensitive identifiers for all of the reasons that I've outlined. Any that don't are problematic, at best.
Re: locales: the way that this is typically handled is by a) restricting the allowed characters to the English alphabet (older), or b) by using Unicode (UTF-16) encoding for source files (newer).
Traditionally := was used for assignment, which makes sense since it is an asymmetric symbol for an asymmetric operation.
When I started, = was the norm for assignment, with == for evaluating equality. Later I started noticing that people were recommending <- based on scope concerns, but it was kind of subjective preference. Now I see articles like this saying that the "preferred use" is <-, and some people don't even know about =.
I agree = versus == can lead to tricky errors in code, but I still prefer = for various reasons in general in languages (although I get the scope arguments).
The reason is that = is shorter, and assignment is an definitional equality, which to my impression is the whole point of programming in general usually. As others have suggested here, between = and ==, = is by far the more commonly used operator in meaning, so to me it makes sense to use = for succinctness.
In math, there is an "equal by definition" operator, with three lines, so I could see that, but keyboards don't have that, so it's more steps. := is also a kind of standard "by definition" usage as discussed in the article, but again, it's more steps. I'd still prefer that over <- in R.
It really wasn’t until Google published their style guide and then later when Hadley published his that I started seeing = in heavy usage.
To your point about = being shorter I wonder if the difference was that a lot of the folks I interacted with were using Emacs with ESS which would interpolate _ to <- so they wouldn’t have noticed? Just a theory.
Either way I was taught from some of the R creators that <- was evil and to be avoided. It wasn’t until I stopped using R regularly that I switched, it became too mentally taxing to change assignment operators when I bounced between languages
x = x+1 is really confusing because of the way kids learn math. The notation x <- x+1 at least conveys the idea of taking a value and storing it somewhere.
I spent some time as a package author, that’s where I found <<- to be the most useful
That’s a common misconception (even the official documentation is misleading).
In reality, assignment `<-` and assignment `=` have the exact same semantics (except for precedence, and unless you redefine them, which you can). The confusion comes from the fact that the `=` symbol is syntactically overloaded: it is also used for named argument passing (`foo(x = 1)`). This leads people to falsely claim that, if you wanted to perform assignment in a function call (which occasionally makes sense in R), then you’d have to use `<-`. But this is false. You merely need to disambiguate the usage, e.g. with extra parentheses: `foo((x = 1))` works.
http://www.theasciicode.com.ar/extended-ascii-code/underline...
> Looking at this as a whole, = was never “the natural choice” the assignment operator. Pretty much everybody used := for assignment instead, possibly because = was so associated with equality. Nowadays most languages use = entirely because C uses it, and we can trace C using it to CPL being such a clusterfuck.
Articles, like desserts, aren't always right. I am sorely tempted to pull out Sammet's book tonight and start counting.
Results:
ν = ε 18 (AMTRAN, BASIC, COLASL, CPS, FORTRAN, FORTRANSIT, JOSS, JOVIAL, Klerer-May, Laning & Zierler, MAD, MADCAP, MAP, MATH-MATIC, MIRFAC, PL/I, QUICKTRAN, UNICODE)
ν ← ε 4 (APL, DIALOG, IT, LISP2)
ε → ν 2 (MADCAP, NELIAC)
ν := ε 1 (ALGOL)
ε * ν 1 (BACIAC)
So pretty much everybody used ‘=’. ALGOL was an outlier, though admittedly more influential than BACIAC.
Plankalkül for example was out a decade earlier and used →. EX:
P1 max3 (V0[:8.0],V1[:8.0],V2[:8.0]) → R0[:8.0]
https://en.wikipedia.org/wiki/Plankalk%C3%BClSo, there really was a lot of languages out there, but := was fairly common because it was easy to parse and type.
As far as := is concerned having to use the shift key for an assignment is less than ideal but there really aren't any better options on a modern keyboard. I think C's use of = is a bit braindamaged but it is nice and quick to type.
Rutishauser was, at least.
And most BASICs, though some (most?) required the keyword LET to introduce an assignment, as well.
But most BASIC versions would accept 'let a = 1'
Personally, I don't see why the big fuss. If your assignment and comparison operators are the same the only thing you lose is doing x = (y == 2); kind of expressions (since you would disambiguate from conditional expressions).
Now, assignments are done more frequently than comparisons, hence why it makes sense for it to be the simpler/shorter one.
LET was virtually always optional; an expression by itself on a line (like 'X <> 10') is a syntax error anyway, so removing LET doesn't introduce any ambiguity.
Plankalkül had the order and the assignment reversed relative to C, so to increment a variable Z1, you'd write:
| Z + 1 ⇒ Z
V | 1 1
S | 1·n 1·n 1·n
(The first line is the 'main line'; Z stands for Zwischenwert, i.e. "intermediate value". The second line is the 'value line', which contains the indices of the variables -- to store Z1 + 1 in a new variable Z2, you'd replace the second 1 with 2. The third line contains Struktur-Indizes.)[1] https://web.archive.org/web/20090220012346/http://delivery.a...
[2] http://www.cs.ru.nl/bachelorscripties/2010/Bram_Bruines___02...
Probably more important is that the Unix/C developers came from Multics, written in PL/I, which (following Fortran) used ‘=’ for assignment. And Fortran was still important; Unix had a Fortran compiler at least as far back as Second Edition, when C was being born. Kernighan & Plauger's Elements of Programming Style used Fortran and PL/I for its examples, and Software Tools used Ratfor. ‘=’ is simply what they were used to.
We read it out loud as "becomes the same as", this being what was actually happening. I've carried the habit over to C despite the bare =, it helps me reason and seems to aid in preventing the =/== error.
"x gets 7" is fine but
"x becomes the same as y" seems better than "x gets y"
Genuinely asking - my first thoughts are:
"transiently the same" is not the same as "the same", hence it is not "the same", and therefore "not the same".
No jokes on the meta-level please :o)
You’re kind of making the parent’s point.
There is a slight distinction. If you are creating a programming language, you can come up with whatever syntax you'd like for assignment and the user has to learn it. Some choices are better than others if you want your language to be used, but there is a specification for the language that you have written down, either as a human readable document or as the compiler/interpreter. With pseudocode in technical documents, the author typically lacks the space, time, and interest to generate such a specification and leans on mathematical notation to keep things precise.
let mutable x = 4 // x is equal to 4
x <- 7 // x is equal to 7I see your point, though. I was just bringing up an alternative consideration in the debate regarding equality.
There's no K in Thompson & Ritchie.
Kernighan seems to get inserted all over the place regarding Thompson's work. It's a bit disrespectful to the legend himself.
a := 5
You can also equivalently do
var a = 5
I almost never use the latter within functions (only at module scope).
I believe the former is forbidden except within functions.
var a = 5 a, b := 2, 3
Then you also get the operator inconcistency between "assigning to a new variable" and "assigning to an existing variable" and is it so important to tell the two apart from within an operator, of all things?
When writing Go, I have been way more inconvenienced by its error handling mechanics than having to tell that I want a new variable using a few more characters.
A language could reasonably use plain = to let the programmer state that two things are equal. In a debug build, this could be compiled into code that validates the condition. A debug build would abort if the condition is not met. For a performance-optimized build, the compiler could use the information for optimization. It could be like the __assume keyword that Microsoft supports.
assert(length < 10)1 => x
[ 10 11 12 13 ] => mylist
Reads and evaluates left to right. C seems backwards after that. (Also LISP-type lists without commas).
> 1^A
1
> A+A^A
2
> IF{A+A*A^B B}
8I really wish more languages had an specific assignment operator.
(Pascal is essentially the easy bits to implement from Algol68)
In pure algorithmic notation, people might use <- or <= instead.
I was super confuse when I saw:
i := i + 1
I'm not sure how I get used to it but when I discover Erlang I feel so much happy since we nolonger has that in Erlang.
I wish more language follow it.
Sometimes mathematicians us ":=" to be clearer and shorter.
In general, = is not rigorous. It's used as a stand-in for the word "is".
Another example is when someone writes x = 1, ..., N, meant to imply x is changing, not that x is equal to a tuple. Also the index in summation notation, which really is an assignment as in programming. I could go on.
I really don't think saying "a = 1; a = a + 1;" is the equivalent of saying "1 = 2" because "a" is a variable.
And as a teacher of programming to adult students, this causes a small amount of wasted time with the overwhelming majority of students, because they know what "=" means, even the non-mathematical ones. Usually not a huge obstacle, but a perceptible one.
And once they learn that it's a mutating assignment, some (small) fraction of students continue to forget, for an infuriatingly long time, which direction the mutation goes.
> C history ... The essence of BCPL was the array and pointer model which abandoned any hope of strong typing and (with hindsight) a proper mathematical model of the mapping of the program onto a computing engine. Even the use of := for assignment was lost in this evolution which reverted to the confusing use of = as in Fortran. Having hijacked = for assignment, C uses == for equality thereby conflicting with several hundred years of mathematical usage. About the only feature of the elegant CPL remaining in C is the unfortunate braces {} and the associated compound statement structure which was abandoned by many other languages in favour of the more reliable bracketed form originally proposed by Algol 68. It is again tragic to observe that Java has used the familiar but awful C style.
0: http://www.cl.cam.ac.uk/~mr10/
1: https://gist.github.com/seaneshbaugh/e09abd748ccc07c5463f253....
It's aimed at teaching programming to 10 year olds no prior experience. One of the first examples is implementing RSA.
So in "x = x + 1", x in its next state is mathematically equal to x + 1 in the state before the assignment
"takes the value", "is equal" and "becomes equal to" are not saying the same things.
What matters is inside the parser for the language, and expressions (as in spoken or written words) between people about the language. Confusion abounds in the latter, but a well designed language doesn't have moments of confusion in the former case.
Some of the choices incur backtracking cost parsing the input. Some don't but incur more keyboard presses per unit of code expressed.
2 + 2 = 4
<n terms> = <constant>
I don't recall teacher showing you equalness through stupid but important ideas: 2 + 2 = 2 + 2
2 + 2 = 2 + 1 + 1
all these are equal, and it's important later on when expression, with variables, aren't reducible to a constant but to another expression which is equal to the other one. IMO it would help a lot for people to try to simplify a problem into a known identity or interesting form (for instance the 'trick' of adding and subtracting one to get closer to binomial formula or other)Or perhaps <-
I must have misinterpreted dboreham as suggesting
destination -> source
Which is very confusing to me.A classic way to guard against accidental use of assignment in logic statements is to write it "backwards", i.e.:
if (2 == x) {...}
Since (2 = x) won't compile, while (x = 2) will not only compile but might appear to work for a very long time.Variable on left is conventional, it's how people sound it out in their head, so it's just easier to process mentally when the expression becomes more complex.
Yoda syntax catches that single error at the expense of making all equality expressions harder to write, read and maintain.
And, besides, if you're going to use Yoda syntax, you're going to want your linter to enforce it... at which point you may as well turn on -Wall anyway.
ben@burnination ~ $ cat test.c
int main() {
int a = 3;
int b = 2;
if (a=b) {
return 1;
} else {
return 0;
}
}
ben@burnination ~ $ clang -Wall test.c
test.c:4:10: warning: using the result of an assignment as a condition without parentheses [-Wparentheses]
if (a=b) {
~^~
test.c:4:10: note: place parentheses around the assignment to silence this warning
if (a=b) {
^
( )
test.c:4:10: note: use '==' to turn this assignment into an equality comparison
if (a=b) {
^
==
1 warning generated.And the errors result from typing "=" when you needed "==" has nothing to do with mathematics. It is a language ergonomics issue resulting from two operators being legal in the same context and one being similar to (a prefix of) the other.
The equals sign has been used for hundreds of years in mathematics and it's completely unnecessary to change this meaning.
That is an assignment. There is also a math symbol that specifically designates equivalence.
But if = in math notation denoted "assignment", what could ≠ mean?
Letting ~ represent an equivalence relation on a set X:
Equivalence relations have specific properties:
x ~ x holds for all x in X
x ~ y => y ~ x for all x, y in X
x ~ y /\ y ~ z => x ~ z, for all x, y, z in X
However, they do not have to be equal, merely equivalent by whatever our relationship requires (though they may be equal).A common example that I would expect most CS students to have encountered is the idea of "congruent modulo n".
We would say that `x ~ y (mod n)` iff `(x mod n) = (y mod n)`.
So over the whole of the integers we can see that every odd number is equivalent/congruent to each other over 2. But we would not say that they are equal to each other.
On the other hand, we would say that 1/4 is equal to 2/8 (or `1/4 = 2/8`).
http://www.math.harvard.edu/~mazur/preprints/when_is_one.pdf
The punchline:
> In appropriate deference to the manifold ways an object can be presented to us, objects need only be given up to unique isomorphism, this being an enlightened view of what it means for one thing to be equal to some other thing.
Your example
y = mx + b
could equivalently be stated as y - b = mx
What does it mean to assign something to ‘y - b’?Another example,
2 = x
means exactly the same thing as x = 2
What does it mean in your semantics to assign x to 2?[edit]
:P LOL at downvotes, HN is no place for silly annecdotes "this is serious things we dooo".
hmm, I don't agree with what you are implying defines acceptable comment content on HN. There are many if not a majority of comments highly rated on HN which are highly personal that necessarily talk about themselves, it's part of what makes those comments interesting, if we all commented like wikipedia articles it would be a dull place indeed. My comment maybe flawed in the eyes of other HN readers, but not for the reason you have highlighted.
In my case it's is both personal and relevant but to be honest - not interesting or insightful which is probably why someone mean enough gave me a downvote... but then afterwards I broke the "first rule of fight club" which on HN is equivalent to throwing yourself to the wolves, I accept that, and also can't help it, it's just my nature, I don't self censor for points.