Programming language notation is a barrier to entry
blog.sigplan.org
blog.sigplan.org
> As names, car and cdr are great: short, and just the right visual distance apart. The only argument against them is that they're not mnemonic. But (a) more mnemonic names tend to be over-specific (not all cdrs are tails), and (b) after a week of using Lisp, car and cdr mean the two halves of a cons cell, and languages should be designed for people who've used them for more than a week.
In particular, the last sentence. I find arguments that go "but this looks confusing to a beginner!" unconvincing. Language syntax and notation should be aimed at actual practitioners, not beginners.
----
Look at Perl for an example of what happened to a language with syntax that appealed only to people who were already Perl-practitioners.
Now, you are right that the syntax is for practitioners. But if it's too bizarre to non-practitioners, it is in fact a barrier to entry.
Sure, being a beginner is a temporary situation, but let's not excuse absurdities and unnecessary convolutions as a result.
The argument actually says "don't optimize notation for beginners, optimize for practitioners". It doesn't imply making things difficult for beginners on purpose; just that they are not the priority.
I just don't even think it's a useful remark because you can just claim people you disagree with are beginners and people you agree with are practitioners and now anything can be argued.
I wouldn't call people who disagree with me on PL design issues beginners. I would call beginners beginners. It's not a matter of disagreeing, it's a matter of experience and time working on real projects with a language.
What about languages who support multiple ways to do something in a semi-inconsistent manner?
For example with Elixir, in a lot of cases you end up calling certain standard library functions like foo(hello: :world) while others are foo(:hello, :world). Now, there's special magic happening in the first case but stuff like this is super confusing and it gets easily compounded with more complex examples.
As a beginner I found Ecto's various ways to calls things super confusing and I had to reference examples almost all the time to get the syntax right (which often had multiple ways of being "right"), but I've also talked with more experienced Elixir developers and they get hung up with Ecto's syntax too.
So where do you draw the line between making the language / DSL better and keeping it catered to only folks who are basically at the same skill level of the language creator?
One of the super cool things about Elixir is how easy it is to extend the language with DSLs. However (and Jose et al are very up front about this), the more DSLs and macros you use, the more confusing it's going to be. Using macros is a tradeoff because it can be a very neat solution but it also obscures what the code is doing quite a bit.
Re. foo(hello: :world) vs foo(:hello, :world) - that one I got the hang of more quickly, though I'll admit it's confusing at first. In the first you're providing a keyword list (which is merely a list of 2-membered tuples where the first member of the tuple is an atom). And when a keyword list is the last (or only) argument to a function it can be provided without the brackets. So the same thing could be written as foo([{hello:, :world}]) or also foo([hello: :world]). In foo(:hello, :world) you're just providing arguments. When you're first learning the language this can definitely trip you up but I think once you become a "practitioner" it's not too bad and ends up being pretty nice.
God Elixir is so cool.
They seem like a good thing to me!
The aim is never to make things difficult for beginners on purpose. The point is that making things less confusing for beginners is not worthwhile if it comes at the cost of making things more cumbersome for practitioners.
I actually agree with the comment which states "But that's not really a beginner-specific problem. Even practitioners will likely find it to be a nuisance."
But that's not really a beginner-specific problem. Even practitioners will likely find it to be a nuisance.
Concepts are way more important. Spending time on syntax is like having a course in tar --flags.
It became for several years one of the most successful and widely used scripting languages on Unix?
The idea that today's languages and technologies are truly better than the past is laughable, particularly when they recycle so much of what is old and slap on a new name
That's really orthogonal to the discussion, though. This isn't about general superiority, it's about legibility and barriers to understanding/skill due to syntax and keywords. Perl is infamous for the "code golf is the default style" syntax, keywords, and tokens it employs in this context, and in my experience at least its most ardent fans believe it is a badge of honor.
It isn't, really, and it's the same mistake PG makes when discussing "car", "cons" and the like.
Indeed. And I disagree with your opinion. For all its terseness its still wonderfully expressive and it goes against the horrible trend set by Python with its "one true way" bullshit. That we all so willingly accept the commodification of our trade so the paymasters can eventually decrease our demand is shameful.
We should strive for a certain level of difficulty if anything to separate the wheat from the chaff. You wanna play with the big boys? Learn to code in a big boy language.
Instead we let the trainwreck that is modern JS take over.
The community exhausted the entire design space (along with some full-blown prototypes). Eventually, the core Julia team chose to use more forgiving scoping in the REPL (virtually always the first point of contact for beginners), while actual projects enforced stricter scoping rules.
My key take-away is to consider how the language interacts with its ecosystem, not just how it should ideally operate in isolation. I have found the Julia team to be consistent in this pursuit. If the first point of contact is intractable for beginners, the project is dead on arrival. A technical tool should be tailored for experts, but you don't want to kill adoption along the way. Engineering is tradeoffs.
This article goes more in depth along the same lines: https://pchiusano.github.io/2016-02-25/tech-adoption.html
[1] https://discourse.julialang.org/t/another-possible-solution-...
Please please, how do we convey this to the mathematicians that infest wikipedia?
And in so doing, they slowly bootstrap themselves so far away from any grounding context that it is incredibly difficult for anyone to _learn_ mathematics from reading wikipedia articles.
There are groups doing great things like Setosa [1], nLab [2] and for concepts, Simplicable [3]
For example, ruby has plenty of warts and annoyances that become evident with experience, but I still find plenty of ruby code that’s lovely to read (like it’s telling a story). Plenty of atrocities too, which is exacerbated by the language.
I see the value in succinct languages too, which require a knowledge investment to be productive. Just more efficient.
It’s kinda like QWERTY typists and stenographers, but with far more middle ground. No one can argue that stenographers are faster, but it’s also not a realistic expectation for everyone to learn it.
1. It's vaguely mnemonic of FirST and ReST, but compactified enough that it's still parseable as abstract names with no particular semantics
2. It's composable the same way that CAR and CDR are, e.g. FFRST == CAADR.
Another good pair of names is LHS and RHS for Left Hand Side and Right Hand Side. Same argument applies.
Lather, rinse, repeat.
For extra credit, write a macro that will save you having to type every possible combination out manually.
Notation in any field is a barrier to entry. See programmers takes on both math and music notation that are often posted here.
The thing about notation that people sometimes don't think about is that it's used by many different types of people. Some things that would be convenient when notated implicitly make thinking about other things difficult.
Anyone remember all the different ways music was represented as they grew up? What about guitar tabs? I remember simplified math as a kid.
I think this hits on the right approach, which already exists: different languages with different complexity levels for notation.
I think it’s a really interesting approach of stripping away a ton of information and relying on the musician having heard the music before.
I don't remember different notations for math and music as a kid...
Music was always sheet music, even when I learned classical guiter; Chord names were "procedures", at times spelled out, and sometimes only referenced. Guitar tabs exist, but were only for people who wanted to play without taking the time to actually learn the craft (at least when I was growing up) EDIT: This was the attitude of every teacher around me growing up, not mine.
As for math - other than having variables written down as empty squares, circles, triangles and stars which were later replaced by letters, I think the first real simplification of math notation that I met was Einstein summation; Everything else was just building on previous things, sometimes with new concepts (like complex numbers) and sometimes as shortcuts (limits and infinites). But they were all shorter than what they replaced, and compatible.
> I think this hits on the right approach, which already exists: different languages with different complexity levels for notation.
There was a guy who proposed something along the article's lines in 1958, which went even farther - instead of "single word descriptions", he chose a single symbol graphical description, because -- though ignored by the author of this article -- most of the world does not have English as their first language. I like the result of that 1958 suggestion a lot; most people place a SEP field over it. The guy is called Kenneth Iverson and the language is called APL with modern descendants J and K.
This attitude really rubs me the wrong way. Tablature is absolutely ancient. It's been around for hundreds of years, and this glib dismissal is just the absolutely epitome of an elitist mindset.
(Yes, I did grow up in a culture highly inspired by older/east european ideals, where kids were welcome to do whatever they want such as learn the guitar, after they've finished their homework, practiced their piano lessons. And as long they're on a path that -- if kept -- would lead them to be on a professional level) I played classical guitar for a few years, but stopped because I just wanted to play popular music, and there were no teachers around who actually taught that who were accomplished enough to be considered worth the money by my parents (who had a very limited budget).
Tablature (for guitar/bass guitar, anyway) reduces note and rhythm information to merely a fret# on a string. This is perfectly fine for people who just want to play a song. I look at tabs sometimes for a quick reference since it's just mentally easier than sheet music.
But if you want to start understanding the theory and relationships between those notes and chords to any meaningful degree, you really do need to start using sheet music. "Fret #3 on the E string to Fret #3 on the A string" doesn't communicate information like "G -> C" does (like the key center, the tonic, etc). It also doesn't limit you to any specific section of the fretboard.
Sheet music is really not that hard to learn. If you want to just use tabs, go ahead, but it's not elitist to point out that sheet music is basically essential for progression as a musician. Music theory is a fundamental part of that progression, unless you just want to play tablature (or sheet music) robotically.
1) First learning multiplication with an "x", as in "4 x 3 = 12".
2) Then in algebra switching to " * " (dot) to avoid ambiguity with x-as-a-variable.
3) Later learning that " * " wasn't multiplication but actually dot product - it just happened to do the same thing with two scalars.
1) When I grew up, multiplication and the letter X were styled very differently in books so that that was never ever a confusion.
2) The switch to dot, was NOT to avoid ambiguity, but to shorten stuff (IIRC we did that when we learned order of operations and started solving simple equations - writing "4 (dot) a" instaed of "4 (cross) a" -- and then just "4a" but multiplications still appear often among parentheses - they were not ALL replaced with dot.
3) That's simply wrong. It is a different operation; dot product reduces to multiplication among scalars -- but so does cross product. It's okay to denote scalar product with either.
Granted, this might be dependent on the exact curriculum one studies through - but mine did NOT have a simplified one that was replaced - instead, it started with the most explicit one, and got simplified (with new, compatible notation that did not replace older one, only extended it) through the years. And it seems this was almost the case for you too.
On 3, I think you misunderstood my point - in 2, we learned dot (the symbol) is multiplication, never touching on dot product because that was a good year or two later, then at that point were corrected to it not being what we first learned.
It's not perfect by any means, but it's better than what people used before (Basic, Pascal).
Anyway, if you want an actually good option, Logo is still around. But it does not let you create anything really interesting, so you better get through it fast or your courses will become boring.
But I'm not sure it's better than old basic. Python is definitely more readable - but it is less helpful than e.g. line numbers in helping form a mental model of a running program; and from my experience, "goto" and then "gosub + return" are easier to explain than funcs with local scope.
If I had to teach newcomers today, I would definitely teach Python -- because, once they do get it, it is actually useful for whatever they need to get done tomorrow morning. But I suspect e.g. BBC Basic (which has goto, gosub but also proc and local) would be easier for them to learn, and would get them farther in understanding programming in a shorter time.
Also like you mentioned, in even the medium-short term the thing that will get people learning the fastest is the thing that gives them a reason to code. Pretty much always right now that's going to be JS or sure maybe python.
Looking backwards, it is my experience that line-numbers and goto/gosub clicked more quickly and more deeply than function calls.
But, as you say, and as I alluded to in my first post, motivation to write code tomorrow morning is likely a better thing to optimize for.
I would refuse to teach JS, though; it has way too many warts one can't avoid, creating damage that would take a while to undo.
Explaining functions and recursion is tricky in general, but necessary for other things to make sense. There are ways to do so that might be more intuitive for a newbie to grasp. For Python specifically, take a look at https://thonny.org - it's particularly well suited to forming a mental model of a running program, including function calls, in a very visual way.
9000 REM F=FOREGROUND COLOR, X AND Y ARE COORDINATES
9010 IF POINT(X,Y)=F THEN RETURN
9020 PLOT(X,Y,F)
9030 X=X+1 : GOSUB 9000 : X=X-2 : GOSUB 9000 : X=X+1
9040 Y=Y+1 : GOSUB 9000 : Y=Y-2 : GOSUB 9000 : Y=Y+1
9050 RETURN
Which does recursion, visually, with only gosub/return, lets me touch on stack overflows (trying to fill too big an image would on most interpreters).I have never had the chance to teach BBC Basic, but if I did, even though it has PROCs and LOCALs, I would still go the GOTO -> GOSUB/RETURN route, and only then proceed to DEFPROC / LOCAL.
[0] This would also be introduction to multiple statements - I would show this with one statement per line and multiple statements per line and discuss pros/cons
There's like 50 years of notational history, and it can be wildly inconsistent. Some things are somewhat standardized, but there's no central location to look this stuff up. Probably one of the better resources is Benjamin Pierce's "Types and Programming Languages" (often referred to simply as TAPL), but I've seen even recent papers deviate from TAPL's treatment.
So when new people are trying to break into the PL research sphere, there's a huge barrier to entry. Invariably, they'll need to ask other researchers to decipher some of the notation in whatever papers they read because the authors just assumed the notation was prevalent enough to not warrant explanation. This works fine for people privileged enough to already be working in PL research as undergrads, but it's much tougher if you're coming from a different background. The PL research community is fairly active on Twitter, but explanations of decades-old notation do not conform to 280-character text messages very well.
It's a big issue that a lot of us younger PL researchers talk about often. I ask myself: how did I learn this stuff? Many meetings with my first PI as I had to bring up symbol after symbol from whatever papers I was trying to read. I took a graduate course in PL semantics which also helped, but then I've also seen deviations from that notation, so...??? And there's still notation I'm unfamiliar with. I was reading about substructural type systems recently (which are super cool, by the by) and encountered some symbols I didn't know and my current PI didn't know and none of my lab mates knew, so I resorted to a PL Discord server where some grad student at another university across the country was able to chime in with the answer. Such a headache.
I can tell it is super solid content, but, damn if I can crack it.
Why is that? If our ideas are good, don't we want the most number of people to know them?
What if PL research isn't getting the population of people and ideas it desires? Look to organizations that have been around longer with even more inscrutable forms of knowledge, how well do they age and evolve?
I think there is a lot more at play here when it comes to optimizing for beginners. Imagine if we had metalanguages for conveying knowledge and that we could distill common subjects from 6 years -> 6 weeks -> 6 hours -> 6 minutes. Or tooling that allowed us to write in both the Simple English Wikpedia, the standard and at the expert level at the same time? Or refactored our writing so it was in the Inverted Pyramid [1], automatically generating tl;drs and keeping our long winded posts on track?
Wouldn't that be amazing.
[1] https://en.wikipedia.org/wiki/Inverted_pyramid_(journalism)
We know that there's a lower limit to distillation that cannot be breached, "Kolmogorov Complexity", which extends Shannon entropy significantly. Furthermore, though this limit is universal, it does (in a well defined way) depend on your prior knowledge. It may seem paradoxical but isn't - it's just too long for this space....
It is not proof in the mathematical sense, mostly because your statement cannot be qualified in such a sense, but it strongly hints that "it would be amazing" because it is impossible.
Which is not to say we cannot do better than where we are today; but there is likely no solution that is, along all scenarios, better.
Are you saying that I am asking for a universal perfect compressor, like I am winning the Hutter Prize [1] ? I don't believe I said that.
To add to the tooling, what if the forum itself "learned you", and took that into account when I respond. It would know everything you wrote, potentially everything you read (on this site), and you in turn would know me. It might help mediate a discussion or prevent folks from talking past each other. What if our systems, could allow us to contextualize communication so that it was grounded in the same basis of knowledge as the receiver?
Might some of those signaling mechanisms be the norms and affectations of each branch of study?
You have given me some things to think about [2]
[2] http://users.ics.aalto.fi/pkaski/kca/
[3] Looks like they have some thought provoking books and interviews https://en.wikipedia.org/wiki/Gregory_Chaitin
Is Mathematics Invented or Discovered? https://www.youtube.com/watch?v=1RLdSvQ-OF0
[4] A physicist looks at linguistics through a complexity lens http://www.its.caltech.edu/~matilde/LinguisticsToronto7.pdf http://www.its.caltech.edu/~matilde/MOL2019MathLinguisticsSl...
> Are you saying that I am asking for a universal perfect compressor,
In many ways, I think you are, of the Kolmogorov/Chaitin kind, if you are talking about language (which I thought you were). If you are talking about tooling, that would expand the same text at different level to different readers, then -- I didn't understand that, and that wasn't my response; I'll have to ponder that for a while before I respond.
I do think that algorithmic complexity is the right context to consider notation, if that wasn't clear -- not in that we can (or should) distill knowledge to the smallest representation, but in the implications of what different "interpretation machines" (that is, persons reading ...) need in order to understand a notation.
We want that. But notation is for working with and on ideas, not teaching them. You want a tool that maps well to the problem domain, that lets you efficiently operate difficult concepts by proxy of simple symbols, whose relationship maps closely to the rules governing the problem domain.
The task of a beginner isn't to grasp a complex domain immediately (as that's impossible). Their task is to gradually ingest ideas (and increasingly complex notations) until they can operate at the level of the experienced, with the tools of the experienced.
Demanding that a notation used for work is also easy for beginners is just handicapping the experienced, and making advanced work near-impossible. It would be like demanding that every excavator be no larger than 1.5 x 1.5 x 1.5 meters and be made entirely of colorful plastic, because that's what little kids can work with.
The "barrier" that a notation creates isn't there to keep people from crossing it. Building ramps over that barrier is a good thing to do. But the barrier has a purpose - it's there to contain and control the forces of thought that are unleashed inside it.
Do you think the tools for broad communication and peer practitioner communication could be very different because they have different aims? Spreading findings in cutting-edge research as broadly as possible is perhaps in some cases not the primary goal.
That is not to say we should in any way that we should gatekeep knowledge! That is unconscionable. It's just that perhaps Watson and Crick's paper on x-ray crystallography of DNA is perhaps not the ideal time to stop and explain what electromagnetism is so that a reader can understand what an x-ray is.
We want beginner languages to make it as easy as possible to learn programming, and to be able to write a program to do simple things. Our PL ideas may be good, but no, we don't want beginners to need to know them. We want them to be able to learn to program at all, and the fewer ideas they need to learn to do so, the better (that is, the more of them will learn).
And then we want to encourage them to move on to more advanced ideas. To do that, we want to make the more advanced ideas as accessible as possible (which I think was your point). And to do that, we need to make ideas as orthogonal as possible; that is, we need to decouple them as much as possible. This lets someone learn as much as they want about one area without having to first learn all the other areas.
Agreed! I read in some essay (maybe one by pg, even) the following tidbit of wisdom: "I designed this for practitioners, i.e. people with more than a few weeks of experience with the language". This is what languages should strive for.
edit: almost! It was a comment by pg: https://news.ycombinator.com/item?id=21256727
This really comes out in math, where choice of notation can call attention to different aspects of a concept. The Wikipedia page for ordinary least squares, for example, freely intermixes (depending on how you count) 3 or so different styles of notation.
I think that this is why I've always been fascinated by Knuth's literate programming. If you want to make something more understandable, saying it twice in two different ways is a great technique.
[0] https://analyzethedatanotthedrivel.org/2018/03/31/numpy-anot...
K below:
,//
|/0(0|+)\
The first expression is usually called "flatten" (take a nested list and produce the list of all items in a single-level, in order). The second expression is a solution to the maximum-subarray-sum problem. Both of them leverage a hell of a lot more than "array" programming; in each case, every single character brings its own semantics, and the combination yields the semantic of the entire expression.Not disagreeing that NumPy is more popular. Am disagreeing that they are "extremely similar".
It clearly developed as a way to reason about music that already existed in the west at that point, rather than be a principled way to go about understanding the underpinnings of music.
And like democracy, it's the worst idea (except every other system that has been tried). I've looked at many notations over the years, and each seems to get one thing exceptionally right at the expense of nearly everything else (and unusable for some cases), whereas standard sheet music is universally bad for all uses, but not bad enough to make it unusable for any common use -- at least as long as you stay within the 12-notes-per-octave world)
Im 57, and have been playing the piano for 50 years. (I still practice 2 hours/day).
The problem with piano isn't reading the music, it's getting your hands in the right place, at the right time, with the least amount of effort, to be able to control the keys the way you want to.
Most of my practicing is getting from one place to another on the keyboard. I'm practicing the space "between" the notes (or positions on the keyboard; I start off by 'grouping' things).
That's why I'm amused when I see "learn keyboard" systems that somehow make you believe the problem is simply seeing a C# two octaves above middle C on the staff and figuring out where it is on the keyboard. You can learn all that in a week. The rest takes years.
Generally, but not always. I find the deep chromaticism of Messiaen or even Vierne (organ, not piano) very badly served by traditional musical notation - I have to stare at it for ages before it becomes apparent what I should be playing, whereas I can sight-read a typical SATB hymn.
I humbly submit that this is the perspective of someone who learned to read music 50 years ago and at an early age, and so who has, for most practical purposes, never not known how to read music.
That perspective will not necessarily be shared by a beginner who has been learning to read music for a few weeks. Sure, you can learn what the lines mean in each clef, and then you can learn to recognise lines above and below the standard five, and then with time you even learn to recognise lines far above and below. You can learn which lines mean something different by default if you have sharps or flats in your key signature, and eventually you actually remember that natural sign just after the start of the current bar before you incorrectly play B-flat for the middle line of a treble clef because your key signature says so. Oh, and you're playing the piano, so you also have to do all of this for a major chord. And a minor chord. And chords with added sevenths. And chords with suspended seconds or fourths instead of thirds. And inversions that almost look like chords with suspended something but that actually have four notes at odd spacing. And then a further note on the octave to make sure all five digits get a workout. With a random double-flat in the middle. Which lasts a strange amount of time, because you're playing that chord as part of a quintuplet. With markings showing suggested dynamics over the whole phrase.
I doubt anyone learns to read all of that fluently in a week, or a month, or even a year unless they are playing far more than any normal beginner. Yes, you can learn the main lines on a stave in a week, but the other 99% of reading music also takes years for most students.
This isn't to say that getting your hand to the right position efficiently and then playing with the action you want doesn't matter; of course it does. I'm just arguing that reading "the right position" from realistic, non-beginner scores is itself a significant challenge for most players, often for many years.
Terrible for what? My understanding is that it is quite good at accomplishing its intended purpose (ie transcribing western music in a dense manner that can be rapidly understood by a professional with only a glance).
I don't know much about music theory myself but the topic of alternative "improved" musical notations came up on HN regarding a music related Show HN that made the front page. Frustratingly I can't seem to dig it up right now. In that discussion many points were made regarding a number of tasks including simultaneously sight reading rhythms and chords, on the fly transposition (I think this was somehow related to symmetry?), and notational density.
The issue, in my opinion, is that notation and technical literacy are skills that students are "expected" to pick up through proxy. They do problem-solving problems, so they are getting experience coding. But the issue with this is that since novices are struggling with notation/syntax, they are juggling problem-solving issues (logic errors) and technical issues (syntax errors).
[1] https://dl.acm.org/doi/pdf/10.1145/3373165.3373177
[2] https://digitalcommons.usu.edu/cgi/viewcontent.cgi?article=8...
Every field develops jargon and writing shorthand. It doesn’t have to be an artificial barrier to entry, and it’s a bit weird to me that this seems to be the case in programming.
Similarly, martial arts often see a high attrition rate. This has a number a reasons, but one that can be addressed is forcing novices to spar with more advanced students way too early. I know the Gracie Jiu-Jitsu schools shifted away from letting new students spar until they've trained for at least 6 months specifically because they didn't want students to show up, get destroyed, and quit out of frustration.
Sort of my point is that the jargon and notation shouldn't simply be introduced and expected to be assimilated without explicit practice that is not confounded with additional problem-solving skills. Those can still be another practice activity, but students should be given the opportunity to just practice the skill without it.
Do you think that reducing the amount of syntax-errors via a powerful structured-editor would help reduce the syntax burden?
> Do you think that reducing the amount of syntax-errors via a powerful structured-editor would help reduce the syntax burden?
Truthfully, no. It would resolve syntax errors for sure, but I think there's something inherit to actually interacting and fixing mistakes that can't be replicated with technology. If you think about learning a second language, having an automatic translator will help you speak to other people, but you would not actually know the language. Not attempting to invoke Turing's Chinese Room, but rather there is a educational need to make mistakes, identify them, and resolve them that helps.
Yet, we think programming is somehow so different, that people should be able to make new creative works (ie: problem solving) at the same time as they're still trying to get a handle on the techniques. Being in the industry for 15 years has shown me there are a lot of really great problem solvers who are terrible at programming, despite having done it sometimes much longer than I have. There's one consistent distinction I generally see. The ones who can code fluently, for whom code just roles off their fingers, all started programming well before a CS program, maybe even back to middle school or earlier.
And I don't believe it's that programming education is that great for the younger age groups (often it's non existent formally), but they have one key advantage. For the most part, they learn to program without the distraction of having to learn CS theory. Not so different from how a musician would likely first learn some simple scales, chords, and songs without any theory. For the unfortunate ones entering a CS program having written little or no code, they're bombarded with very difficulty CS theory and having to learn programming at the same time. I think we'd have both better programmers and better computer scientists if the education system did a better job separating the two.
Thanks for linking your research, I'm going to give it a read soon. Is there a way I can get in touch? I'd love to discuss these ideas with you further and learn more about your research?.
My current research is on exploring different exercise "types" with the hope of building a practice recommendation system that can identify which complexity a student should practice next - a typing or coding exercise, or something in between (Parson's Puzzles or Output Prediction, etc.)
[1] https://research.csc.ncsu.edu/arglab/people/agaweda.html
Running [code] is a good thing because in Cartwright and Fellisen[1], [there is a] trivial typo (- instead of =). When you write things in TeX it's easy to make some mistakes, it's a good thing that it happens as a clear typo here but it could be a less clear typo and then it requires significant thinking [..] what was the correct thing. The advantage of using Haskell as a metalanguage is that if you replace = with - by mistake, almost certainly the compiler will not let you get away with this, you will catch it right away.
Oleg does exactly this in [2]. If the sequent calculus-style rules put you off, there's an OCaml implementation!
As an example I did this to the denotational semantics of R5RS Scheme, translating them to Haskell, the final result[3] is essentially a one-to-one mapping from the math to code.
However there are some aspects of PL theory where you cannot write things as code directly (e.g. operational semantics, typing rules), and you have to resort to notation, but it's used pretty consistently.
[0] https://youtu.be/GhERMBT7u4w?t=4395
[1] http://citeseerx.ist.psu.edu/viewdoc/download?doi=10.1.1.41.... (typo fixed)
The author quotes a tweet about someone complaining about the notation used for a basic type judgement for being cryptic. Any introductory text will immediately cover this and looking it up, you can learn the notation (which is really rather simple) in about 5 minutes. Yet, the tweet author claims that they saw this notation 12 years ago and still can't understand it. Well, to me the only reasonable conclusion is that they simply didn't bother spending a few minutes looking it up and instead expected that they would magically understand it after enough time has passed. I seriously doubt that gammas being used instead of the word 'context' is why they're struggling.
And this is my core issue with a lot of these "notation is a barrier" complaints. Instead of being valid complaints about esoteric notation (and especially using this notation without defining it!), it's people who believe that knowing how to program means you should be able to read any PL text without any preparation which is just unreasonable. If you want to seriously engage with a paper, you'll have to read up on the basics first in literally any field.
I recall attempting to read a PL paper for the first time. My background is in mathematics so I found it a fun puzzle trying to figure out what the notation meant. I think I maybe managed to look up a few things and it was reasonably easy to guess the rest. Some things are very hard to search for (how do you google the notation for inference rules where a horizontal line separates the antecedents from the consequents? I spent 20 minutes trying to find the name for the notation and failed).
The other difficulty with reading PL papers is that you need to know that often you should just read the introduction and definitions and skip the extremely long and tedious proofs by structural induction of (somewhat boring) properties like soundness. It tends to be the case that the contents of these proofs is not generally useful (though it is often necessary that such things are proven).
[1] https://en.m.wikipedia.org/wiki/Backus%E2%80%93Naur_form
In both cases, entry is the easy bit. Learning the notation is a somewhat orthogonal and arguably much smaller impediment than the inescapable difficulty of figuring out what you're trying to accomplish, and then how to accomplish it.
Personally I've always found math notation feels like it was optimised for space efficiency and writing speed on paper, akin to legal shorthand rather than the beauty I'm always told it has. Similarly for attempts to bring programming syntax closer to math syntax. In college I found myself translating equations of medium complexity such as the Fourier transform into pseudocode to understand.
Obviously however, my familiarity with programming notation colours my perspective as much as their experience with math notation colours theirs
>@> <$> <*> <|> <=< <+> <$> --> <&&> <||> ||| .|.
None of the above are made up, all are from real code.
If you know what those symbols mean (aka the interfaces they belong to.)
If you don't have that knowledge and are instead complaining about the fact that they have symbols period, I suggest you quit spouting answers when you should be asking questions.
data Person = Person
{ age :: Int
, name :: Text
, height :: Int
}
Person
<$> parseAge
<*> parseName
<*> parseHeight
Applicative style is one of the easiest constructs to read relative to the power of what it typically does.Most other times when you use <$>, you're using it instead of something like this:
fmap f $ x y z
Which is simplified by f <$> x y z
Google is a non-issue. Hoogle exists and works great (and can even index all symbols importable by your project) ($) :: (a → b) → a → b
(<$>) :: Functor f => (a → b) → f a → f bMathematic notation is horrible. We have to teach kids the forms for _years_ and most are still shaky at it once they're adults. You have to memorize everything because the syntax is basically shorthand, nothing is self explanatory. Then we wonder why kids struggle with math.
If a programmer decided to replace his function names with single characters from a different language he would be dope slapped by his colleges. If he always used single character variable names he would be kicked out.
To this day, when I'm reading computer science papers, I have to look really hard to make sense of the math gibbrish and get at the actual meaning, and often when I do I realize that it would have been much clearer in a few lines of JavaScript or Python or Go or Rust or just about anything else.
Well, perhaps anything except Haskell, OCaml, and similar. Even after learning the syntax, I still haven't internalized it as effortlessly as I did with other languages. Note that lisp is not very algol-like and I have a much easier time parsing it than I do Haskell despite relatively equal familiarity. I think Haskell has given us a lot of great ideas, and thankfully they've filtered down from the Ivory Tower of academic languages into pragmatic mainstream languages, such as Rust.
Haskell syntax is not more complex than an Algol descended language. It's only confusing if you expect an Algol descended language, for example because you're already familiar with them.
When I first learned Haskell, I only knew C. Haskell codebases were immediately so fun to dive into and learn from. I wouldn't hesitate to hit the "Source" button on Hackage. The fact that all code could be reasoned about locally with just simple symbolic substitution was a breath of fresh air.
To a person with 0 experience at all, Haskell or Algol-descended languages are going to be equally alien. Your experience is because you had asymmetric familiarity.
So I'm not surprised by your experience. But it says more about you than Haskell.
It might require a library or file of accompanying "assumptions" that enables a concise "display" notation that appears in the text, but it would allow anyone to examine and understand what was being specified by having a common, concrete base of reference and the ability to query and examine it from whatever angle works for the reader.
I can take some pseudo-code and turn it into a working C (or python) implementation to poke at without too much trouble but papers with pure algorithmic notation usually get put in the "I'll look at that later" folder which I never get around to looking at.
Granted, I'm not the intended audience for most of the papers I read since I have no formal CS (or math) education -- though enough information is out on the interwebs depending on how motivated I am to try out whatever shiny-new I'm learning about.
It goes farther than suggest in the article by dropping words entirely and using symbol, which removes the "need to use English" barrier which is invisible to people whose mother tongue is English but is most definitely there.
It gets treated with a SEP-field by most of the programming community.
The problem is not that objectively "X is a barrier to entry" (for whatever X is). The problem is that everyone has their level of comfort with something they consider natural and everything else needs to be simplified. But that something has a huge variance and unclear mean/median/mode.
From the linked text: > others shun clarity lest their work is considered trivial
This is a real problem in industry. Being told that the simple solution you spent 5 days finding for a very tricky problem proves that you were just goofing off when it was that easy is annoying, to say the least.
[1] https://www.cs.utexas.edu/users/EWD/transcriptions/EWD13xx/E...
We would expect that an aspiring architect needs to learn a little bit of notation, and a lot about the tradeoffs and constraints of making a building. We should expect that programmers learn a little bit of notation, and a lot about the tradeoffs and constraints of making software.
Knowing your tools well is important, but not sufficient. Focusing on "programming languages" makes people think that the important part is the notation. It is not.
There is also I think a difference between Architecture and software development in that an architect aims to design and plan with drawings and documents a building that in the end is distinct from them, whereas software development produces a set of documents that are synonymous with the software itself.
No. This has been tried a lot and has mostly failed. The notation can be a barrier but it isn't the hardest part.
Since the author is writing about mathematical notation, it is true that mathematical notations sometimes can be bothersome. `Σ` for example has a better (crude) representation in programming language `.forEach` or `.map`. Someone who doesn't use them very often can be really frustrated by it.
But, this also applies to non-programming languages. Logograms, like Chinese hanzi, Egyptian hieroglyphs, are all symbols and can describe things more succinctly, but has a really really steep learning curve.
Not that far, German has the word "Tschüß" which means "bye". An English-speaking person not actively reading German may imagine the sound of it as "Tshoob" while it actually is closer to "choose". The letter "ß", by intuition, is closer to "B" while it is actually an "ss".
I believe that the other half of the PL-notation-entry-barrier problem is which notations are required for which layer of programming language API. If learning is like going up a stair, it's not just about the angle of the stairs, but the height of each steps.
CSS for example (I would argue that CSS is a non-turing-complete declarative programming language), is rather good for this. You only need to use `.class` and `#` when you're styling something with class or ID. Only after that you will learn how to use `:pseudo-class` for a more complex behavior, after that @keyframe, etc.
For example, regular expressions and date strings can be written in a nice and compact form, but how nice would it be write `Not(DigitalCharacter)` instead of `\D`? Or `TwoDigitYear` and `FourDigitYear` instead of `y` and `Y`. The same convenience exists in short and long command line flags; and optional named parameters in function calls.
They might have preferred this "font"
https://mindsarentmagic.org/2020/02/19/a-picture-of-grahams-...
If the language is strongly typed the compiler should know whether the types are compatible and do it for you or refuse to compile.
I suppose my point is that casting was a poorly chosen example because the obvious solution isn't to have a better notation but to have no notation at all and follow Python's lead with duck typing or functional languages and strong type inference.
Moving between different C-style languages, once you know one, is not terribly hard however. It really is getting over that hump of learning one language from which you can then take that knowledge into another similar one.
Some time later I did learn Python and I really love it.
and I am not the only one saying this.
I later went back to bracked notation and never really had a problem anymore. Having each bracket combination be coloured uniquely helps a lot too.
It worked wonders for me. I always felt that brackets were too messy. And I bet it didn't help that, at the time, I usually worked with other students that generally didn't uphold any style guide. So brackets weren't indented properly and not much time was spent in making the code readable.
So when I got my internship I was thrown into Ruby with a style guide to adhere to. Which definately made my introduction a lot smoother compared to what I was used to.
I still prefer the do-end notation for methods, though.
If I may add a pet peeve of me, I would say that auto formatters and bad style guides are a barrier to entry. Some of the things out there seem to not have bothered to be actually adapted to the human visual system. Also, good formatting is a very difficult job to do automatically so the results are likely to be suboptimal. As a visually inclined person non-optimal formatting really disturbs me.
The problem is that you still need to worry about the semantics of the language. If that's standardized across languages, then you only have one language.
Consider a bunch of languages that all share a standard lambda syntax. If one of the languages has different capture semantics than another (e.g. move vs copy), then you only superficially understand the standard syntax. They mean different things. Having the same syntax doesn't buy you very much.
The assumption here is that this is a bad thing. But there's a reason we have doors, gates, locks and so on -- other "barriers to entry".
Yes, because we don't want people to get into our house.
Are you seriously suggesting a world where we want to make it harder for people to understand programming language literature?
I'd like the barrier for entry to be a bit higher if I'm honest. Maybe then I wouldn't have to work with code created by so many idiots.
More people able to understand the literature means more people who can contribute and contribute effectively. That means better programming languages, better tooling and better educated people.
I've worked with some brilliant booksmart people that produced god aweful code. Being able to understand literature/notation isn't going to make you a better developer by default. There's a lot more that goes into it than that.
In the post the author mentions an excellent talk by Guy Steele that goes over issue of notation and balkanization titled, "It's Time for a New Old Language" where he outlines "the history and current status of computer science metanotation". [2]
Notation, like syntax is important, but like a plethora of ill defined and specified Domain Specific Languages (DSL), notation also needs to be defined, regular and learnable. For notation takes words and turns them into pictures that compose into thought structures that cannot be expressed in human written language to the same specificity. Think Feynman [3] and Railroad [4] Diagrams, or musical notation [5]
The abstract explains it beautifully,
> The most popular programming language in computer science has no compiler or interpreter. Its definition is not written down in any one place. It has changed a lot over the decades, and those changes have introduced ambiguities and inconsistencies. Today, dozens of variations are in use, and its complexity has reached the point where it needs to be re-explained, at least in part, every time it is used. Much effort has been spent in hand-translating between this language and other languages that do have compilers. The language is quite amenable to parallel computation, but this fact has gone unexploited.
[1] https://simplicable.com/new/accidental-complexity-vs-essenti...
[2] https://www.youtube.com/watch?v=7HKbjYqqPPQ
slides, https://groups.csail.mit.edu/mac/users/gjs/6.945/readings/St...
paper, https://sci-hub.st/10.1145/3155284.3018773
[3] 6 minute video from Fermilab on Feynman Diagrams https://www.youtube.com/watch?v=hk1cOffTgdk
[4] https://sqlite.org/syntaxdiagrams.html
[5] https://www.mfiles.co.uk/music-notation-history.htm
[z] Wadler's Law
There's no reason to consider barriers a good thing, unless one has a motive to restrict use of programming languages. Why would one want that?
They skipped the related xkcd.
(judging by the picture of Susie at her desk, development environments' handwriting recognition was awesome in 1968, which explains Landin's 1966 comment in Next 700:
https://www.cs.cmu.edu/~crary/819-f09/Landin66.pdf p. 160
> "[The offside rule] is based on vertical alignment, not character width, and hence is equally appropriate in handwritten, typeset or typed texts.")
Syntax is where the rubber meets the road, but this is like blaming the stick and puck for making hockey hard.
Which possibly means - the more notation-capable-but-semantics-challenged people are working in the field, the more work for the semantics-capapble people to clean after them.
Just like how there is a class of programmers who don't understand simple CS concepts like algorithm complexity and memory allocation and write horribly inefficient code.
Of all the people I have tutored thus far, if they struggled with a language like Java or C#, trying to program with Scratch was just as hard for them. Taking syntax completely away and making it visual did not make a difference.
I’d understand this argument about a language like J, but we can’t dumb J down just so everyone can use it without putting in some effort, or it wouldn’t be J.
The most powerful tools will always have the highest learning curve.