The design side of programming language design
tomasp.net
tomasp.net
The Human Computer Interaction (HCI) community should care about this, but is beset by two problems: people in PL don't take people in HCI seriously because the people in HCI don't do proofs, and, the people in HCI generally don't have the background to start tackling the mathematical principles of programming language design.
This seems like a natural time for some kind of cross discipline collaboration, but the two forces above keep potential collaborators apart. PL people talk about judgements and sub-structural typing and HCI people's eyes glaze over, HCI people talk about human subjects, experimental methodology and statistics and PL people's eyes glaze over.
I've advocated for the creation of a new field of study, call it Programmer Computer Interaction. No one takes it that seriously though, and the intersection of researchers that care simultaneously about things like good experimental design, statistical significance, type theory and constructive logic seems to be just me.
The problem I have with that is that it seems to be limited to verifying existing structures and solutions. That kind of quantitative research has its use, but is not really helping with exploring novel ideas in human friendly interfaces.
For example, I stumbled across Céu a few years back and I have never seen concurrency and timing done in such an intuitive style before[0]. Other things that have blown my mind in the last few years include Halide's decoupling of algorithm and scheduling[1], Jonathan Edwards' experiments with schematic tables as an alternative to if-statements[2], vega-lite's grammar of interactive graphics to declaratively construct interactive plots[3], and aprt.us' hybrid graphics/text programming environment[4].
These are all novel ideas that you cannot find through quantitative measurements about which syntax is optimal. Not that the latter is without use, since it can make people shut up about pointless disagreements (although I think the better solution to ending a holy war on bracket style is to get rid of it altogether and use something one default style like elm-format[5]). And maybe we're also not looking outside of our own field enough.
For example, I recently read Steven Pinker's "The Stuff of Thought". There was a chapter discussing all kinds of (human) language paradoxes and hidden rules based on how humans have different ways of thinking about aggregates, and how we use those subconsciously in our daily language. As I was reading it, it made me think "this makes so much sense of how different languages use collections differently, they're just applying different styles of intuition described here!"
The funny thing about most language paradoxes is that they require very specific ways of framing a question, and that they disappear when framed differently. Which sounds a lot like how some problems are easier in one style of programming or the other. And that makes me wonder if we can't learn a lot from these branches of linguistics about how we might set up our computer languages in such a way that humans are less likely to make errors of thinking in them.
[2] https://vimeo.com/177767802
[3] https://vimeo.com/177767802
[4] https://www.youtube.com/watch?v=i3Xack9ufYk and http://aprt.us/
I don't know other research exists, but I wouldn't personally value any research on beginning users comprehension because you're typically a beginner in a new language for a few days to a few weeks but an experienced user for years or decades, so beginner experience is just not something I care about in the big picture. (Obviously there are reasons other people might care about it... it affects initial adoption of the language and so on, but I just don't care about them personally).
This might just be an issue of perception on my end, although it seems consistent with the HCI-conscious PL work I've seen.
Personally, I have two priorities in PL design: expressiveness and aesthetics. Both of these seem perfectly suited to a design-oriented approach. But are they a good fit with the sort of work regularly done in HCI? I genuinely don't know.
Looking in the other direction, I would like to think that even the most theory-obsessed PL people can appreciate that "intuitive" and "powerful" can both arise from clean design, but I've also been told, "your proofs are boring" as though that's a bad thing.
"intuitive" is very subjective, I think. I know lots of PL people that think that functional programming, for example, is intuitive. After many years of study and practice, I too find functional programming intuitive, but I can remember a time when I did not, but I could still program. From this I conclude that something being "intuitive" is very subjective.
Expressive power, on the other hand, I think we can put an objective definition on. My favorite take of this is Felleisen "On the Expressive Power of Programming Languages." Of course, what use is something being powerful?
(I'm not current on PL/HCI literature, so it wouldn't surprise me if they already have these concepts, just with different labels.)
[0] https://www.over-yonder.net/~fullermd/rants/userfriendly/2
I'm not convinced language expressiveness can really be quantified that way.
emphasis on reducing constant overhead at the expense of adding more linear overhead after a certain threshold has been reached
Who exactly is emphasizing this?
Well that's not true. If the HCI people actually had empirical data supporting their positions, you'd be a fool to ignore them. But what are PLT researchers to do with HCI claims that are completely unsupported, or whose supported positions are extremely limited in scope? We can't halt work on PLT to wait for them to catch up.
Some collaboration on empirical studies on programming ergonomics are needed. Given how huge much money is invested in this industry, I'm surprised more effort isn't put into this.
http://program-transformation.org/WGLD/
However, it seems to be drifting back to traditional SIGPLAN topics lately. The problem is the conferences: there was a concerted effort to make them more sciency and less designy in the last 15 years and....well, design no longer sells to acedmics.
The HCI community has similar problems actually. You can't really equate empirical evaluation of user behavior (a topic in itself) to design, where many aspects are quite difficult to validate empirically (things like Fitts' law can evaluate simple behaviors, but don't really scale to complex behaviors).
> I also believe that interesting programming language research is not about finding a keyword that will make writing for loops 3x easier
If we go out on a limb here a bit, "a keyword that will make writing for loops easier" seems to be a metaphor for everyday ergonomic issues in coding.
(I don't have a hypothesis about what the article is about, though).
A few principles:
- Acknowledging that the user is probably familiar with other languages, so stay within those conventions when possible. This is inspired by Jakob Nielsen's maxim that web designers should acknowledge that their users send 99% of their time on other websites.
- Making the most common activities the easiest. Larry Wall refers to this as applying Huffman coding to the syntax itself.
- Safe defaults. Making dangerous operations less convenient than the safe path. PHP has pretty much the opposite behavior, where many of the default design choices lead to security issues.
- Cognitive Load. Minimizing visual noise and the number of micro-decisions the user has to make. i.e. There should be one (good) way to do it.
- Preferring shorter, clearer terms over technical jargon. e.g. I use the terms "Flag" instead of "Boolean" and "List" instead of "Array". One can get pedantic over the exact meanings of symbols, but during the actual act of programming, simpler terms can help reduce friction.
The project is at https://tht.help if anyone is curious.
I'm not familiar with THT and this is a nitpick, but doesn't this contradict the first principle? Pretty much every language I've ever used has some notion of "boolean", and I don't know one that uses "flag". Don't want to bash, just curious how that decision was taken.
In actual use, you rarely interact with the names anyway -- you use `true` and `false` as values like most other languages.
It some cases "Flag" makes much more sense than "Boolean", but when using logic gates does "Flag" make any sense? Wouldn't it make people think of actual flags more than an on or off signal.
"List" and "Array" mean different things in CS. Lists cannot be accessed arbitrarily whereas arrays can be.
> One can get pedantic over the exact meanings of symbols, but during the actual act of programming, simpler terms can help reduce friction.
"Simpler" is a matter of perspective. Any experienced programmer should not have any problem with Arrays vs Lists vs Queues, etc, and they are all necessary concepts to be able to talk about and program with.
Most jargon exists because it is able to describe things that are not used in everyday life much more efficiently than without it.
etymonline doesn't have an entry for flag as a noun meaning boolean, but this might be related:
> flag (v.2) 1875, "place a flag on or over," from flag (n.1). Meaning "designate as someone who will not be served more liquor," by 1980s, probably from use of flags to signal trains, etc., to halt, which led to a verb meaning "inform by means of signal flags" (1856, American English). Meaning "to mark so as to be easily found" is from 1934 (originally by means of paper tabs on files). Related: Flagged; flagging.
When I think of what's interesting about Go and NodeJS, and about even more ambitious ideas for the next 10 years, my head is around languages built for communities of programmers. Even better, designing languages that improve over time by leveraging the combined activities of programmers and users, in code and out.
- It's worth optimizing for more people than fewer.
- There are a thousand times as many future programmers than there are existing ones.
- One can measure how easy it is for a clean-room human to read and create software in different languages and build something for them.
If there's a good reason for the difference, by all means, go for it. But if there isn't, please reconsider.
So that leaves little room for alternative language paradigms then. If you want dictionaries to look like JS or Python then what are languages like Lisp and APL to do, not use coherent syntax. Or what about just making different use of the limited number of symbols on the keyboard to make the use more consistent?
Id rather have a more consistent language or one that introduced new concepts than one familiar in syntax just for the sake of being familiar.
It's reminiscent of this: https://xkcd.com/927/
Edit: further to your point, I agree that my gripes are merely about an up-front cost, whereas consistency can be helpful throughout your use of a language.
Edit again: I essentially seek consistency within and among languages. I don't mean to state one should come at the cost of the other at all times.
Yes, but what about the benefits from the regularity and uniformity?
If you're a programmer now, you're a secondary audience to the masses of people who will program in the future. If the language is going to have the most impact, what they prefer supersedes what you (and I) prefer.
> If there's a good reason for the difference, by all means, go for it.
Of course and agreed. Nothing should exist without reason. Antoine de Saint Exupery etc.
I think it is oversold, but the book "Design of Everyday Things"[1] goes over this for many common items. There is a long section on doors with many interesting points to consider.
[1] https://www.amazon.com/Design-Everyday-Things-Revised-Expand...
You're thinking of syntax as chrome. It is more than that.
I was speculating that this belief is from folks that think design is just chrome.
Many times I've discovered that altering the syntax for something (not its semantics) can completely transform its use.
old:
function double(double a, double b) { return a + b; }
new: (a, b) { return a + b; }
Currently, the syntax for in/out contracts is being revised for usability:https://github.com/dlang/DIPs/blob/98052839441fdb8c6cc05afcc...
D has a different syntax for templates than C++, and is far more approachable and convenient:
C++:
template<class T> void foo(T t) { }
template<class T> class S { };
D: void foo(T)(T t) { }
class S(T) { }[1] https://smile.amazon.com/Selected-Papers-Computer-Languages-...
Note that I am specifically not arguing against immutability. I have, however, seen bugs caused both ways. When I was a TA, Java students were notorious for not understanding why calling "trim" on their string left the whitespace. To often disastrous consequences. I would be delighted to see empirical studies to weigh against all of the anecdotes in my head. :)
Unfortunately, I have no more data on the real-world results than anyone else. :-)
http://dl.acm.org/citation.cfm?id=2884798&CFID=985095807&CFT...
It's a small example, maybe it's a start.
It is just a start, so I am hesitant to pick at it. I am excited to see this getting explored, I was hoping for a dive into bug reports against software, though.
That is, I share the same bias most of software developers share that mutable code is ultimately dangerous and should be avoided. The flipside, only when working with established codebases does this really bother me. Greenfield projects that are not close to shipping are usually quite clean and not a concern. Battle worn codebases, though...
To that end, the CWE attempt (https://cwe.mitre.org/) was an interesting start. Especially if it was backed by numbers in the CVE database. So my question is, how many of the CVE reports could have fully been laid to this class of bugs? (I'll see if I can put together a notebook going over this idea. Still hoping someone else has already given this a better treatment than I am likely to.)
In Scala, for example, immutable variables are declared with "val" while mutable ones are declared with "var". Here's there's no great cost in typing or screen space (and to the uninitiated, it may not even jump out at you). But Scala developers will almost universally use val; in the language idiom var is an unusual thing that feels wrong. And I say above that it won't jump out to beginners; but to experienced Scala developers that l vs r is a huge glaring difference -- your brain just gets trained that way.
I think the key takeaway is that some languages are designed with first-class support for immutable variables, and in those languages, you tend to use immutable variables. Even in Java, modern style tends to use "final" for local variable definitions, though the support in the language for immutable types is much weaker, so the tendency isn't as strong. In a language like C, immutability is a practically non-existent practice, and "const" does not serve the same function.
One of the more well known pieces that is worth a read is Cognitive Dimensions of Notations [1]. I've even used them in my research on the usability of debugging tools.
It is composed of 14 dimensions to evaluate your design (of a PL or UI).
[1] https://en.wikipedia.org/wiki/Cognitive_dimensions_of_notati...
It drives me crazy that in my main editors/IDE's of choice I can't just hide all traces of comments when I'm writing code until I'm ready to comment it up or that I can't layer comments, since half the languages now have annotations for everything from documentation to configuration either built in or as add-on's.
It's a lot of noise when you trying to grok just code and gets in the way.
There is a key analogy here for programming language design!
The analogy also extends to specific languages. There are some bikes which are keenly optimized for short distance performance. There are other bikes optimized for long distance performance. There are some bikes which are optimized for comfort. Yet other bikes which are optimized for folding compactly.
All of the above are valid designs, just for different contexts.
Perl 6 started with its design documents[0] -- the apocalypses led to the exegeses, which finally settled into a synopsis. The Perl 6 language itself is not any of these: The test suite defines the language, but the reasons they are the way they are came first. It has design features like hyperoperators to make parallel code easier to write -- which of course should encourage people to write parallel code.
Perl 5 is explicitly a postmodern language[1], and Perl 6 is the same but more. From my point of view, Python looks like it has a strong modernist philosophy. So I'm not really getting where OP is coming from.
[0]: https://design.perl6.org [1]: http://www.wall.org/~larry/pm.html
Mathematics solves this by appealing to a shared underlying "truth", that allows not only programmers with different backgrounds, but also computers to understand and process a programming language.
And I don't think math really "solves" any of the issues that are on debate here. You still need language and notation to express your precious maths, and that falls to the very same pitfalls as what is seen in PL design.
Many do, especially in academia, but quite a few seem to be standing on a combination of "this looks like it'll do what I want" and "this interpreter is the spec."
* https://en.wikipedia.org/wiki/Where_Mathematics_Comes_From
* https://en.wikipedia.org/wiki/Proofs_and_Refutations
As for mathematics being superior for underpinning programming languages - that's certainly how some people make it look, but nobody really tries to explain what exactly is the link between mathematical foundations of programming and actual programming. (I did this for my PhD and I'm still not sure how it's supposed to work.) Everyone just hand-wavy assumes it's somehow there. The best account for how this might looks is perhaps this: https://plato.stanford.edu/entries/computer-science/
the entries are 0 or 1.
If language L can do thing T, then there is a 1 in in cell L,T. Otherwise there is a 0.
For example, a thing could be, "you can feed it a math text and it will give a list of defined terms and expand the definition of any term in the list in the signature {<-}.