A huge reason why C syntax is "not weird" is because its weirdness has been normalized. I didn't realize how absolutely bizarre some of the things in it were until I was an undergraduate teaching assistant for a C programming class for Math and Engineering majors.
The big problems I remember:
- Why {} braces?
- Why A = B? They're not actually equal. This can be a huge conceptual stumbling block.
- Why is this function returning "void"? This makes no sense, that's not a function.
- Why A.B()? Period gets read as a sentence terminator and that function is being called with no parameters, which makes no sense.
- Why do we use square brackets [] and not parentheses () for arrays?
- Why do we need a ';' at the end?
- Lots of confusion over `&&` and `||`.
- Problems understanding `+=` and similar.
Old Smalltalk editors did some things to make the language more understandable, such as turning `:=` into an leftward pointing arrow and `^` into an upward arrow. Once I saw this in practice, it made a lot more sense to me. Not that Ruby and Rust both have the `|vars|` syntax.
Pascal-family languages like Ada look "too weird" because most people learned a C family language early on. It'd be like saying Russian or Japanese is a weird language because you only know Romance languages.
Personally the distinction between subroutines and functions always struck me as overly math-centric thinking, because it's not an interesting difference as soon as non-local state exists, which is almost always the case.
> - =, +=
I don't quite get why this remains contentious; even in math these are somewhat overloaded ("=" in equations vs. "<=>"/"=" for equality of equations, for example); code simply is not math. Why would the symbols mean exactly the same thing in a different context?
> - Why {} braces?
Big braces have always been used to group stuff together (e.g. case distinctions, summaries etc.), so it kinda makes sense to use that for groups of lines, i.e. code blocks.
> - Why do we need a ';' at the end?
C/UNIX was originally written using actual teletypes, so not only does ";" make it easier to parse, but it also gives you a very explicit marker that you definitely reached the end of the statement. That matters when every line you're looking at when interactively editing a file uses up some paper.
It's related to the same problem with = as assignment, note GP's context: Math and engineering majors. They have been taught, assuming they paid attention, that = is not an assignment but a relational statement. If they've never programmed before, having = as a thing that changes the system state is still new to them. This is a common issue I've seen in tutoring over the years, where students' mental model is closer to Excel than to <standard programming languages>. That is, if you do this:
a = 10; b = 20;
c = 3*a+c;
printf("%d\n", c); // prints 50
They aren't surprised, but then they do this: a = 20; // this demonstrates part of their confusion, they half-get that it is an assignment
printf("%d\n", c); // prints 50 ?!?!?
They expected 80. So they haven't fully internalized that `=` means assignment and is not establishing a relationship between the thing on the left and the formula on the right. One of the key things that helps is teaching them to "read" a program, and to read `a = 10` as "a gets 10" or "10 is assigned to a".> Why would the symbols mean exactly the same thing in a different context?
When you have one mental model that you've been taught for 12+ years and are 18-20 years old and don't have a ton of experience, this is confusing. That it is not or was not confusing to you is either because you had a learning experience that specifically targeted this confusion, you made the connection more readily (maybe closer to the top of your class than average), or are looking back on your past and neglecting how far you've come since you were that age and had that level of understanding.
The problem with += (IME) is with the incomplete and incorrect model that students have, += is one of the situations that can force a realization that = in C-syntax languages is assignment and is not a relational expression. But if they haven't grokked that yet, it's jarring.
To make this a slightly more useful comment, I'll mention that I confirmed that even with -Wall neither clang-6.0 nor gcc-9 gives a warning about the second assignment to "a" being unused. This would often be a nice thing.
Not really, many programming languages generally considered math-centric are perfectly fine with functions returning (or taking) a unit type of some sort, category theory (and even its predecessor abstract algebra) has a lot of use for initial and terminal objects, etc.
It’s more like mathematicians used to be vaguely uncomfortable with non-referentially transparent things when computation as a mathematical object was new. They realized the error of their ways pretty quickly, all in all, but a realization such as this can take decades to make its way through layers of educational materials.
> I don't quite get why [ = for equality] remains contentious
Because it’s a huge problem when teaching, even if it isn’t for competent practitioners. Maybe it wouldn’t be if we put CS (and graphs, and games, and lots of other things utterly lacking in prerequisites) in elementary-school mathematics where it belongs, but if wishes were horses I’d have a racing franchise by now.
I mentioned the = thing because I only ever see this brought up by what seems like a very particular developer tribe. I used to tutor 1st semester CS students in programming C and Java - for many this was the first time they were programming - and I don't think I've ever seen someone confused about it. Lots of other things though; programming is hard.
Well, the word is taken from math, it seems pretty rude to take a word while preserving nothing of its original meaning. A function is a mapping between sets, it has nothing to do with abstraction or reusability or any of the dozen other keywords said when introducing the concept in programming courses, it's just an association between sets satisfying certain (relaxable) constraints.
From this perspective, a function that returns nothing is a hilarity, it collapses all elements of its domain to a single point. A function that takes nothing and returns nothing is a deviancy, where is the mapping then if nothing is taken and nothing is returned?
Keep in mind that the usage of 'functions' to denote callables* is a new thing that happened probably around the mid to late 1980s. Computer Science always called self-contained pieces of code '(sub) procedures' or '(sub)routines' or '(sub)programs'. Fortran, the oldest programming language, honors the difference until now. Pick any conference paper or book written before 1980 and I guarantee you that 'function' will denote a literal mathematical function used to describe some property of an algorithm or declare a certain constraint on a computational object, and that callable code blocks are always called one of the terms above or their variations.
>because it's not an interesting difference as soon as non-local state exists, which is almost always the case.
The thing is, there is no such thing as 'state' with functions in the first place. The word denotes something in a world that knows nothing about change or time, taking the word and redefining it to include those concepts is meaningless. Imagine if I took numbers, added physical units to them, then insisted that this is the mathematical definition of numbers all along. That's what happened with functions.
Also how exactly is the difference between pure subroutines and mutating subroutines irrelevant when non-local state is involved? it's the exact opposite, the difference is an irrelevant and opaque implementation detail if the mutated state is callee-local and never leaks to the caller, otherwise it's an outsized obsession of life that can mean the difference between a good weekend and a sleepless 24 hour and all things in between. It's exactly when it's non-local that mutable state is revealed as the horsemen of death it is.
* : see? even the verbs don't make sense: how exactly do you 'call' an association? it's a mapping, you either 'apply' it or 'refer' to it or 'look up' an element in a table listing (a sample of) its input-output pairings. Similarly, you don't 'return' from a function, you never 'went' anywhere to return, you had an element and used it to get another element with the help of a mapping between the 2 sets containing the 2 elements, this 'happened' instantaneously, timelessly, there is no instruction pointer in math breathing down your neck and telling you to return here and go there.
It's about command/query separation which is still a best practive to this day.
To give another PoV: I've taught javascript to absolute beginners (graphics design students coming from humanities studies with pretty-much zero programming or general "scientific" ability) and the syntax, fairly close from C, was very much not an issue for them, unlike the maths required ; I had to spend much more time explaining how transformations in a 2D plane, vectors, etc... do work, than correcting syntax mistakes.
In ~20 hours of class they were pretty much all able to write small programs with for-loops, conditions, etc. which did nice visual effects in p5.js so I'd say that teaching the basics of programming is an entirely solved problem in 2021 ; whatever the topic, you cannot expect to teach stuff faster to a whole classroom.
What C syntax is this referring to?
struct a {
void (*B)();
}
struct a A = { .B = ... };
A.B();What I should have written:
- Why is "Foo()" a function is being called with no parameters, which makes no sense.
- What is "A.B" Period gets read as a sentence terminator. I learned C++ and Java very early on in life and had never considered A.B to be an issue.
Especially in the 70s and 80s, that was a huge performance barrier.
Contrast this with C, which probably won because of its largest benefactor, Bell Labs and UNIX. Smalltalk in some sense was a competitor with UNIX itself.
Smalltalk's influence is everywhere, but the sad thing about it is that it's dev tools were so much ahead of time of anything else.
As a Python coder, I played with Pharo and tried several tutorials and read a book on Smalltalk-80. I get the gist of things, but have trouble actually writing the code as things aren't obvious and there is no StackOverflow or cookbook for common tasks. It took me forever to figure out how to do file IO and that sort of thing. The other problem is that deployment isn't just running a script, but deploying the image and runtime of some obscure language nobody has ever heard of in enterprise software and no IT department is going to be cool with it.
i actually meant pharo above, i never used genuine smalltalk, just what the pharo mooc gave us
https://www.cincomsmalltalk.com/main/successes/financial-ser...
Of course, mission critical business software still needs to be deployed long after fashions change.
It is a shame that Self and F-Script didn't gain a bit of attention.
F-Script was really awesome, even more with F-Script Anywhere
StrongTalk was also very influential, because many of its developers also ended up working on core parts of the JVM and V8 JITs both at Sun and Google.
I loved pascal syntax.
IMO Pascal actually has much better pointer syntax than C, even with the caret (which is basically a pointer/arrow up without the tail):
1. Pointer de-referencing is done via something that looks like a pointer and doesn't reuse some math operator (despite writing C for more than two decades i still pause for a bit whenever i want to multiply a pointer to a number)
2. Getting the address of something is via the "at" character @ (which kinda makes sense as a pointer to something is a way to refer "at" it), no way to mix the two
3. No special syntax for referencing a member of a record (struct) through a pointer, all record members are referenced using a dot ".", all pointers are de-referenced through the caret "^" so of course all members to a record pointer is done through using both of them: "^."
Also i think that the reason C has -> is because in C pointer de-reference is a prefix whereas in Pascal is a suffix so without it you'd either have to write (*ptr).foo. Pascal avoids that by having both be suffix operators.
In the case of Smalltalk, it was Digitalk's "Methods" (the precursor of Smalltalk/V) that adopted := for assignment in place of the left arrow (though the IBM PC character set actually had one, just not with the same code as ASCII 1963) since all of its tutorials were aimed at Pascal programmers.
That distinction comes from early C (pre-ANSI / K & R C), where structures were really definitions of offsets (and said offsets were global, you couldn't have different 'offsets' for a member with the same name in different structs).
Say you had a struct 'struct mystruct {int a;int b;int c};':
You could do '200.b' to access the value at the 'b' offset of the memory location '200' ( '*(int*)( 200 + offsetof(struct mystruct,b) )' or '((struct mystruct*)200)->b' )
And '200->b' would access offset 'b' of the memory pointed to by memory location '200' ( '*(int*)( *(intptr_t*)200 + offsetof(struct mystruct,b) )' or '(*(struct mystruct**)200)->b' )
I remember learning Pascal and C around the same time coming from BASIC and assembly and Pascal didn't seem particularly easier only more verbose. New concepts like types and pointers were much more of a hurdle than syntax.
Function: something that may have side effects but that also returns a value.
Pure Function: something that only returns a value, and has no side effects.
That's how I always understood the terms.
I guess you forgot about the semicolons which must or must not be there depending on location. And the final dot.
And, if we move a bit away from pure syntax, strings, even modern ones, which start at 1 for backward compatibility reasons with the original very limited Pascal strings, while dynamic array start at 0 (and original limited pascal arrays start wherever you like).
The final dot is clarifying the end of the program or unit. It also makes a more clear distinction between the main body or subroutine of the program versus other functions and procedures that are only component parts.
I don't think people are saying that Pascal is a perfect language, as none are, but rather it's an easier to understand and read language. Many teachers would agree with that assessment.
As for Smalltalk I can say that it's the only language that has never clicked with me. I have tried the usual languages including Lisp and some weird ones like Forth, Factor and the APL family and never had an issue with those but with Smalltalk I can't go beyond writing factorial. Actually if you sat me in front of a Smalltalk environment it would take me ages to figure out again where to click in order to even have the chance to create such method. The biggest problem with Smalltalk is that you have to give up your tools in exchange of some alien technology that isn't really that good to begin with.
Disclaimer: I have tried Pharo, Squeak and some commercial offerings such as Cincom.
Also I think most super popular languages have a niche. Whether its webapps, text processing, low level, etc. Smalltalk lived in its own world and it was hard to integrate with the OS and other capabilities of the system, (try writing a wrapper for like OpenGL or other libraries for example) and I think that somewhat prevented it from finding a great niche to thrive in.
Often those who learned C family languages first (C, C++, C#, Java, JavaScript...), can feel there is "something wrong" when viewing the syntax and structures of other languages. When everybody only wears dark blue or black, somebody wearing purple or beige could be considered "weird".
The creator of Ruby, Matsumoto, was a C++ programmer. It appears his intention was to make the language a bit more verbose, as is Pascal and BASIC, to increase ease of use and readability more than exists in C++. A person who used Ruby, would probably find reading and adjusting to Pascal or Lua much easier to do.
Smalltalk gets a bit "weird", because of how they do OO. The structure and syntax used reflects a different way of thinking. Saying that Smalltalk "died", seems to be part of this odd labeling for any language not in the top 10. Are Rust, Kotlin, or Go dead because they never were nor are in the top 10? Smalltalk was never that tremendously popular to begin with, but it certainly didn't die. Go check out Pharo (https://pharo.org/).
You can see the opposite end of that spectrum in Go, which can be readily read by anyone since there's generally only one good way to do anything.
I don't program in Ruby but every time I look at it I see something different depending on who wrote the code and what their personal style is
I'd say this is a plus, at least in certain dimensions. One of the best descriptions I've heard of Rails' ActiveSupport (a collection of extensions to built-in types included in Rails) is as a dialect of Ruby, one specific for developing web applications. Other domains could and should use different dialects.Python and Ruby occupy similar niches. Python has many more libraries, so people will say, that's why Python has been more successful, but IMO, even if Python and Ruby had as many useful libraries, most people would still choose Python, because it's a simpler and easier to learn language. That may have contributed to Python getting a headstart in the beginning.
I'm sure some people would say that's a case of "worse is better", but I think you also have to have some empathy for other programmers. When collaborating on a programming project, having a project that is easier for others to understand is easier. Having a language with less weirdness, less peculiarities and less complexity is a plus.
Looking at recent Python code, it also seems like the simplicity is being lost in the rush to add features.
My inkling is that this is because it is hard to have a general purpose language that lends itself to both novices and experts, and to small scale and large scale. Not impossible, but hard. Those categories desire/need different things, and will use the language in vastly different ways.
BASIC (for all its problems) was a better first language for many people than C or Fortran, but BASIC also didn't scale well to larger systems (structured variants scaling better). Then you start getting BASICs supporting OO and other features which complicate them and lose the appeal to novices, but more effectively meet the needs of intermediate-to-advanced users and maintainers of medium-to-large scale systems.
See also Pascal and its evolution. From a small language to a still small but not as small language in Delphi and others. For the novices, they can still use that small core. But it's now harder for them to onboard with a larger project or a project developed by more advanced users of the language. Especially as the standard library (for the language implementation, if not the language proper) grows.
Scheme, over the various revisions, has seen a similar conflict, culminating in the effective rejection of R6RS and the division into small and large for R7RS.
And those are languages that were largely meant to support novices and learners. Python started off similarly. Now it's older, and its users have aged, and they want to do more powerful things with it and maintain larger and larger systems with it.
On the small vs large side, look at the evolution of JavaScript. It was literally meant for in-page scripting, to do small things and not be the larger system, but a mere component. Like a shell script that copies two files contrasted with the actual OS or shell implementation. But now 25 years or so later it's a vastly different language, and people are building distributed systems on top of it.
Python is the same way, there is still (and always has been) a core that is particularly well-suited for novices and a slightly larger core that is sufficient for any small-to-medium system. But as you gain expertise and want to use the language to do other things, or to develop larger systems, you will likely exceed that core. Delphi with Object Pascal was similar. You could still use that core, and accomplish a lot (same with Python), but you cannot guarantee experts or larger systems won't exceed that core.
And that's the language dilemma: How to appeal to both novices and experts, small and large system maintainers? The features that each will want or need are different. Go stayed pretty stable for a decade and kept to its mostly small core, but even it has added generics now because people (maintaining larger libraries, in particular) are chafing under its present constraints. And with these changes, you'll see a (potential, may not happen I suppose but I'd be surprised) major shift in the way the language is used by more expert users that pull it away from its previous model.
You're probably right about the availability of documentation, but I also think Python is much easier to generally understand at a glance. Less features, less quirks means the language is easier to learn.
this has been true for a while, the 'zen of python' was lost a long time ago.
i think this is partially because of how much commercial popularity it has, and all the different types of things people need to do with python.