To the brain, reading computer code is not the same as reading language
news.mit.edu
news.mit.edu
My university went so far as to allow students with a bachelor's degree in music to directly enter the computer science master's program.
I have worked in both computer science and classical music, and found many crossover talents. Two previous Director of Music at my school (UBC in Canada) in the last few decades are both musicians and engineers/physicists.
(though it's been.... a while... since I graduated.)
I'm sure if we tried hard enough we could come up with a pretty good analogy between data structures / algorithms and scales / arpeggios / etc... but I've never felt quite the same when recognizing a familiarish algorithm in code as I would when sight-reading some music and having my fingers perform the right kind of movements out of almost unconscious muscle memory.
It took me a while to realize that I was doing this; I only figured it out from when it doesn't work. I can't do it when I'm looking at code that's dark-themed if I normally work with that code light-themed. In that case I have to read the code line by line. It's actually a strange feeling.
Some of my environments are dark themed for certain languages/environments so it's not the dark-theme itself that throws me off but something about the unexpected visual difference.
[0] The Dream Machine by Waldrop
A lot of the time editing, I'm not really paying attention to the meaning of what's written, but whether the meaning is conveyed properly and clearly in any given sentence or paragraph.
Reading through the research methods of this study, it says the participants read a single line of English text or Python/Scratch code at a time. (Example English sentence of "NOBODY COULD HAVE PREDICTED THE EARTHQUAKE IN THIS PART OF THE COUNTRY"). So this study was looking only at the very granular level of "reading".
I've also found when I'm deep in a problem solving mode like coding I tend to be far less empathetic. So it seems to mess with my people skills too until I come up for air.
I'm skeptical of this because he went on to be a successful, respected computer scientist and business leader, as well as starting a family. While of course many people who've done each of these things may be bad at empathy (or worse!), I'm not aware that Carmack has somehow continued along the course he was on at 14.
https://science.sciencemag.org/content/342/6156/377.abstract
My experience as well.
This is a side effect of a forced WFH. My kids are regularly at the receiving end of my outbursts when they interrupt me when I'm deep in work, more so when I'm debugging. In office on the other hand such 1-1 interruptions are rare.
But seriously, by recalling and introspecting the thought process occurred just before an interruption, I think coding is moke akin to navigating a landscape -- it has a substantial spatial component.
Simply put, it's like thinking about consequences and dependencies at every given line/statement. Something that we don't do (or do rarely) while speaking or reading.
[0] https://cacm.acm.org/careers/243179-forget-math-language-ski...
Sadly, this is underappreciated in the general psychology and linguistics community. Some researchers in Second Language Acquisition research community are more concerned about this kind of mis-interpretation and confusion of different meanings of "language skills/aptitude".
Its all correlated with g.
>[0]
They contrast "language learning ability", ie intrinsic intelligence, with "math knowledge", ie amount of schooling. For reference, you can take (a version of) the math knowledge test used here: https://www.idrlabs.com/numeracy/test.php
If the article wanted to give a more honest interpretation of the paper, they would have focused on the leading predictor of programming ability, "fluid reasoning and working-memory capacity". One supposes this to be quite useful in math and engineering.
I saw a paper that argues that there are two factors: g and math. But that's tangential.
But limitedness of the methodology means the conclusion can't apply to software development at large, something that the title misleads us about.
[1]: https://twitter.com/wcrichton/status/1339235494102753280
One thing I've noticed is that I can type out code at the same time I'm singing a song aloud. If I tried to do that while writing English, I'd start typing whatever I'm singing.
Yet, like the OP here, I can also sing along with music and code at the same time.
(Which is btw not really possible in the office, even with noise cancelling headphones. Thank God for remote work.)
Also the experiment was "The researchers showed them snippets of code and asked them to predict what action the code would produce.".
This may not be the best way to prove the thesis conclusively. Because when I read code I check the syntax fast and then I try to figure out logic.
"The researchers saw little to no response to code in the language regions of the brain"
This may mean that the participants quickly scanned the sytntax(language part of brain is used) and then they tried to figure out the logic(multiple demand network of brain is used)
> This may not be the best way to prove the thesis conclusively.
It certainly isn't. For starters, it does not provide a comparison to normal language.
> The researchers saw little to no response to code in the language regions of the brain
And I say: screw that logic. It's faulty beyond comprehension. That a brain region doesn't light up doesn't mean anything. It could even mean their scanner was broken.
Oh well. Time to go code golfing!
A program can also be seen as an awkward natural language, conveying programmer's intention to other programmers, of course. But it's a different modality of the code.
Exactly. That's why laws or contracts written by lawyers are so verbose. Think of a rent contract or a software license like GPL.
Contrast with WTFPL and similar...
My argument would be that legalese with its strict "May" "Should"s and all that could be sufficiently similar to a (verbose) programming language.
Anecdotally, I certainly have entered "debugging" mode when trying to check if a specific clause of a contract was applicable to my situation before, rather than trying to interpret it liberally.
Of course, anecdote, and only demonstrates that the (my) brain _can_ operate "programmingly" when reading legalese, not that it does so "naturally".
It's interesting to think about, and makes me suspect that this is why I tend to prefer writing my own recipes out as tersely as possible, particularly if I know the general procedure. At this point I don't need to be told to preheat an oven or to scrape down the sides of the stand mixer, so I'll just write down things like "Creaming method, 350° in 8x8 cake pan." I might even start writing them in pseudocode now.
My partner takes it a step further with her cranberry sauce recipe; the instructions just say "C'mon."
Learning a second language, likewise, is best done the same way as the first: through immersion, if you can swing it. The high school student who spent a summer in Mexico came back far more fluent than I was even after three classes in high school and one in college. The plasticity, though, of this area of the brain will decrease over time. It is easiest to pick up a new language before adulthood.
Most people don't appreciate this or don't believe it. For more information, read The Language Instinct, by Steven Pinker.
Programming? Yeah, I guess it spins the CPU more than any specialized chip. I hate how I can't be interrupted while doing it, and for a few moments trying to come out of it, I am in some kind of daze.
I've noticed that people become more logical in all aspects of their thinking as their programming skills improve. Also, I've noticed that a lot of religious people become atheist or agnostic as their programming skills improve.
It trains you to become good at holding a lot of information in your head and being aware of how they relate to each other and how they conflict.
When it comes to coding, I prefer dynamically typed languages because they don't have any constraints and that makes them more effective at training your brain because you can't outsource or put a boundary on any part of your thinking.
[0]: http://web.mit.edu/9.s915/www/classes/duncan.pdf [1]: https://i.imgur.com/xdYzbev.png
In human language, you can play with idioms, metaphors, double entendre, etc.
In writing code, you miss one symbol and it breaks.
[1]: For example it’s occurred to me a few times that early Node interfaces were an opportunity to introduce monad types to idiomatic JS development. Sure, it was callback hell, but basically everything that touched IO had an Option type, just in the form `type Callback<T> = (error: Error, value?: never) => void | (error: null, value: T) => void`
Edit: I’m sincerely disappointed the sibling reply got downvoted and dead, it was a valuable contribution. To resurface, the sibling commenter distinguished the cognitive experience of code versus human prose from the perspective of someone who’s autistic. It’s a distinction I share (diagnosed ADHD, undiagnosed but strongly suspect ASD spectrum).
Edit 2: no longer dead so I'll defer to the sibling comment.
YMMV, but the way I reason through while I read/write/troubleshoot code are mentally visualized and mapped in very different ways. When I compare what I visualize with the junior engineers I support, their visualizations tend to stop at a much higher granularity level. I strongly suspect that I engage in techniques that would be extremely familiar with and be on far more elementary level to those who practiced building "memory palaces" [1].
I have tried to teach this to the junior engineers, but the vast majority are simply wired differently it seems, and the few who "get it" do not sufficiently make it a habit to internalize it, much of which at first involves tedious pencil and paper.
I remember as a kid while blundering about a PDP-8 drawing on paper what I thought was going on at first with labeled boxes and lines (most of which was laughably inept and naive), then gradually transferring that to memory. Over many years of practicing that, it eventually became a reflexive habit. For more difficult problems I encounter today, I still sit down with pencil and paper and diagram it out to talk myself through it and build up a mental model.
For me personally, engaging spatial, motor, visual, verbal, and auditory senses together makes up my "visualization" in an iterative process that gradually internalizes into a pure mental model. That's why I prefer to look like a mental patient in the confines of my private office than at a co-working space or coffee shop.
https://en.wikipedia.org/wiki/Structuralism_(philosophy_of_m...
Have you told anyone to grep, cat, mount, fsck, or sudo?
I'd say the more experience one has with a programming language, the more reading and writing it feels like a natural language.
This was my reaction as well: what, they needed to do a study to figure this out?
The confident students who aren't concerned with getting all the specifics of grammar exactly perfect are the ones that are making the most progress toward being conversational, they're able to express themselves and actually have something resembling a conversation, despite using the wrong prepositions here and there or conjugating verbs wrong sometimes. The more "academically" minded students who obsess over getting the grammar correct like a math equation are still struggling to form even basic sentences outside of solving homework problems on a piece of paper. I've noticed that I've been making much better progress now that I've been focusing on just trying to express myself even if I'm making mistakes.
Which is obviously the complete opposite of using a programming language.
In the early phases of learning a new coding language, I learn how to write lines of code. Sometimes it compiles, most of the time it would cause errors.
At some point it works, but a more seasoned engineer could easily find places to refactor it in order to make it more extensible/flexible, easier to read, less brittle, and most importantly to this discussion on coding language, more idiomatic.
Now, you could say surely I could've seen these by reading the docs and learning about certain library features or patterns specific to the language. But in my experience, I tend to learn faster by first exercising my brain and trying it out on my own take before having someone (be it a mentor or a tutorial/book) correct me so that I have a bit more context about why an alternative approach is better.
In both cases, when you get it that wrong that your meaning doesn’t come across correctly, it’ll usually get pointed out to you, either by the compiler or the person you’re talking to, so you can go back and try again.
Eventually, you’ve potentially gotten it right when the compiler stops throwing errors or the person you’re speaking to understands. Unfortunately, in both cases you may still not end up with the desired outcome as your compiler is “stupid” and cannot understand intent, only what you’ve written. Similarly, but different, the human you’re speaking to may bore of your broken language eventually and just move on.
There’s no perfect system for everyone, but it never hurts to remember the saying: perfect is the enemy of good enough. It doesn’t matter if your language skills are perfect, just that they’re good enough to convey your meaning effectively for the environment you’re in.
Perhaps having an automatic compiler changes our perception much as an auto-tickling device doesn't work, but surely early levels of math and languages have similar learning paths.
> Eventually, you’ve potentially gotten it right when ... the person you’re speaking to understands
I can see that I wrote poorly here myself, and here you (and child) are correctly correcting me, while simultaneously arguing against that happening, which is kinda ironic!
What I failed to adequately express was that human languages are a lot more forgiving of errors, and as long as the message is understood, you’ve gotten it “right”
Tldr: Compilers you need to be fairly precise, natural languages you barely need to be in the ball park.
And if the context fails, things can go just as badly as with computers, in the short run. In the longer run, people notice that something went wrong and are likely to go back to an earlier point and try to clarify, whereas computer systems are not great at it.
If you write in C 'foo(a;', that's equivalent to saying 'ball foo with a' - the compiler really can't understand what you mean.
The equivalent of 'I go store' is something more like writing 'int p; printf("%d", p);' - the compiler understands what you mean, but it's bad style, and possibly wrong (you're supposed to initialize a variable before you read it).
"I go store" is reasonably close to correct: you have a subject, predicate, and object and there is one meaning that fits far better than all others. However switch it up to "I store food" and now you have some ambiguity: are you saying you are storing food for later or are you saying you want to go to a grocery store or are you trying to talk about the food you purchased at the store? Again, a human being incredibly intelligent might be able to tell from context, analogous to how someone reviewing a block of code might recognize what a block of code is supposed to be doing, but in general you could expect the interpreter to raise an error, either of the form "bad syntax" or "sorry, I'm not sure what you mean."
If we consider "human languages" like legalese or diplomatic language, where the goal as with programming languages is to limit ambiguity as much as possible while maintaining the speaker's ability to potentially say as many things as possible, syntax becomes very important, and typically it's preferred that the interpreter raise an error rather than try to guess at the speaker's intentions.
Humans are "fuzzy compiling" to such a degree (and giving no errors or warnings to the writer) that it is hard to compare the two.
> If we consider "human languages" like legalese or diplomatic language, where the goal as with programming languages is to limit ambiguity as much as possible while maintaining the speaker's ability to potentially say as many things as possible, syntax becomes very important, and typically it's preferred that the interpreter raise an error rather than try to guess at the speaker's intentions.
Such a concept is not even possible in natural languages. Natural languages are so full of ambiguity that even native speakers from similar regions often confuse one another. Because frankly many don't realize there's: what you intend to say, what you say, and what is heard (times number of listeners). When listening to a non-native speaker we focus a lot of trying to understand intent from them. When listening to native speakers we focus a lot on what is said and what we hear (often ignoring intent). There is just so much fuzziness to natural language that it is a bit ludicrous to compare them. The fuzziness of natural language and strictness of programming languages are necessary features too.
> Natural languages are so full of ambiguity that even native speakers from similar regions often confuse one another.
These two statements are contradictory. There's a difference between not caring that an error is present and being robust to error. Natural language represents the former. Yes I might not shout "Error: finger shoes not recognized" but that error nevertheless occurred, I don't actually know what they meant. I'm going to assume gloves, but if their life depended on me getting them mittens they'ed probably wish I had said I don't understand. Luckily, we live in a world where getting things wrong is typically harmless, but history is full of disastrous misunderstandings. If you are doing something important like drawing up a legal contract, your lawyer should have no problem telling you "that wording is unacceptable." which is conceptually no different from a compiler's syntax error.
Natural language and programming languages are different in many ways in terms of degree. Natural languages need to be used in vastly more circumstances, they have been in use for vastly longer, and they generally represent an ad hoc development cycle involving millions of independent contributors. They thus have much more deprecated functionality and conflicting user conventions. But while different in degree, the two aren't really different in kind - every fundamental trait of one is shared by the other, which is why the fuzzy compiler for natural language will accept the terms for a programming language as a substitute without raising any error.
So let's look at the two parts you quotes from me and give some examples. We've conversing quite well, wouldn't you say? You understand me and I understand you? So #1 must be true, right? But have you ever played the game telephone[0]. With the general outcome of every game that must mean that #1 cannot be true, yet here we are. Like I said before, when communicating there's: what is intended, what is said, and what is heard. Usually these align pretty well. But small fuzzing adds up quite quickly. But it should also illustrate to you that what you hear (or read) isn't what the speaker (or writer) intended. Their job is to convey as best as they can to you while your job is to interpret as best as you can. Language being of this form allows us to express wildly abstract concepts, especially ones beyond our capacity to even understand. This is the usefulness of natural languages. But it is also why we created mathematical and programming languages, but why they will never be able to have the same expressiveness. You cannot program what you do not understand. But you must be able to talk about what you do not understand, if you are ever to hope to one day understand it.
I'd encourage you to read both more about linguistics and compilers. If you're feeling lazy, Tom Scott has some decent linguistics videos and are aimed at STEM audiences. For more technical maybe see NativLang. For a book, checkout The Unfolding of Language.
[0] https://en.wikipedia.org/wiki/Telephone_(game)
Edit: btw, you did get the meaning of finger shoes (German) even with the extreme limitations of only text based communication.
I don't know if by "we've conversing" you meant we are or we have been conversing. Those two possibilities are close enough to one another that substituting the wrong one will probably not cause a catastrophic failure, but perhaps you meant to use the pluperfect to imply that the conversation was done in some "drop the mic" moment that was completely lost on me.
Telephone is not a game of ambiguity, but of noisy communication. If you start with cat you might get bat at the end, but you wouldn't get feline. You could easily win at telephone by speaking the same meaningless statement repeatedly and you can lose at telephone with a grammatically correct statement which is nevertheless phonetically similar to some other statement. "Blue blue blue blue" is going to make it, "blue crew flew you" probably won't. A fault in transmission is fundamentally different from a syntax error, more analogous to a letter getting lost in the mail or a mistyped letter on the keyboard.
There is a fundamental difference between getting your point across and getting a point across that's similar enough to accomplish your goals.
Consider the following statements:
"I want you to feed my cat
"I want my cat fed"
"I want feed cat"
"You food cat"
"Cat food"
"Weird Dog Rice"
"seltsamer Hundereis"
I am certain you know what I mean in the first sentence. I would be amazed if you could deduce the meaning of the last without further clarification. Between these extremes there is some point where your compiler's heuristics fail, and likely long before that you start losing nuance and details. Natural language error tolerance has its limits.
Now consider the following javascript expressions:
for (let cat in catstofeed) {feedCat(cat)}
var i=1; while (i) { try { feedCat(catstofeed[i-1]); i = i+1; } catch (e) { i = 0;};
var i=0; while (i) { try { feedCat(catstofeed[i]); i = i+1; } catch (e) { i = 0;};
while (i) { try { feedCat(catstofeed[i]); i = i+1; } catch (e) { i = 0;};
If you substituted the second for the first, you'd accomplish what you were setting out to do. The third would lead to a wildly different result but is still valid. The final case is ambiguous, a smart compiler could guess the programmer intended for statement 2's functionality, but asking for clarification is the safer option.
You are right that trying to have a human conversation with only a few dozen words and a very limited sentence structure would be very difficult - a vast vocabulary and complicated grammar is a hard requirement for a natural language, and ambiguity is thus unavoidable. Dealing with such a high level of ambiguity is thus also a requirement of a natural language compiler, and such a compiler is beyond the capability of any existing silicon computer. But just because a language has more words and more valid expressions does not fundamentally change what it is, any more than adding more processing power to a computer eventually makes it stop being a computer.
If you wanted to make a programming language with 50000 keywords and an extremely forgiving syntax, you could do so. It would be extremely impractical for programming a current computer, which is only capable of a few distinct operations and thus has no need for nuance, but it would still work, and damn wouldn't it be great at self documenting.
That's because that's not what I said nor what I meant.
Of course, you also need to learn large amounts of library functions and their semantics, but that's more like learning who the people at court are and what their roles are than like the language. You could speak perfect French, but you still have to learn that you must call the Comptesse de Barry before calling the King if you don't want her to be upset. Even patterns are often more in this area of cultural/local context rather than the actual langauge - more like learning the difference between overly-descriptive prose and good prose is (which is why the knowledge about it also tends to translate better between different languages).
If your English grammar is wrong, I will catch you.
Make progress.
One thing I think I've noticed among a lot of people who have lived in English-speaking countries for a very long time is that, after their accent reaches general intelligibility and their grammar reaches a point where they're always quickly understood, they stop progressing and leave it at that. I knew an Iranian-American guy in his 50s who had been in the US since his 20s who would drop every article except when it was important.
One of my Danish teachers said that the move towards making people conversational is a fairly recent invention, and approximately a century ago (my memory is spotty on when the change happened) teaching was more geared towards making people literate in the new language.
One issue in the study is what a reasonable comparison is. Our ancestors would have talked about "there is a lion over there" and such mundane things all the time, so we shouldn't be surprised if we're just "nexting" those kinds of things with a specialized circuit.
Obviously that doesn't exist for coding, but the question is whether there is natural language that can't be nexted. I'm thinking of James Joyce's Ulysees or Shakespeare type stuff. It's weird even for the native reader. Does that activate the math network?
Also, isn't that a better analogy for how you program machines? You're telling the computer what to do in a language that isn't assembler, and doesn't have the same machine model as your coding language, with a heavy dependence on shorthands established only in the context you created (by writing functions, classes, etc). I thought the reason why a lot of people find it hard to read code is because you're inherently doing something awkward, translation.
I've been learning a new language as well, with the wrinkle that it's quite closely related to one I already speak. Mandarin is lovely, it makes me wonder why there are languages where the verbs and nouns are changed. It also means you can jump into saying useful things pretty fast, you are just restricted by vocabulary. WRT coding, this is similar. Coding languages tend to have very few things you must know (except C++ lol) to get started, but with a lot of reading libs necessary to do anything in a given domain.
Now that said, the idea of comparing programming languages with natural languages or even constructed languages seems very strange to me. Natural languages are tools for thought, expression and communication. Programming languages are tools for, well, programming. Even the relationship between the grammar of a PL and the grammar of a human language is distant (we often intuitively think of identifiers as similar to nouns and verbs in human languages, but in fact all identifiers are equivalent to proper nouns in human languages, and the actual equivalents of nouns and verbs are stuff like block openers, significant white space, function application, maybe the keywords).
It is a well-known fact already that humans can only naturally learn human languages, which have a specific kind of grammar. Even some clumsy attempts at standardizing human language end up introducing un-natural grammar that children can't pick up until later in school when it is drilled into them. It would have been truly surprising to find out that we can use our natural language capacity for something as alien as a PL.
If you skip over the hard parts, then you're not exactly learning a language. You might as well just learn from one of those "Hello, my name is Bob, where's the toilet ? " cartoon books.
Its a bit like those people who learn programming informally. Most of them end up writing spaghetti code for the rest of their lives.
Its the same with languages. Skip the hard bits and your attempt fluency will hit a wall at some point relatively early on. Grammar is tough, but stick at it, it will eventually become second nature. Don't be lazy.
Language isn't just some word for word translations and a new sentence structure, it's much much deeper than that.
The only thing that is required to learn a language is tens of thousands of hours of practice.
That's not true. Most native speaks have little knowledge of the grammar and syntax beyond intuitive. They have forgotten most of what they've been taught at school, plus, they could speak quite well before going to school and being taught grammar in the first place. We acquire language by osmosis, not by studying, and even less so by studying the syntax.
The fact that they don't know the scientific names of the grammar rules doesn't mean that they don't know grammar. As somebody how's native language is very dissimilar to English, my intuitive language rules are no help speaking English. The grammar you are taught at school is descriptive of English and not a prescription. The distinction s very similar to "laws in physics". Nature doesn't really care about the rules we impose on it.
Learning grammar alone is not enough. Learning vocabulary alone is not enough. Immersing yourself to a language without a guide (like your parents guided you into language) is completely ineffective.
Grammar, vocabulary and pronunciation are equally important if you want to get fluent in both spoken and written communication in a different language.
Similarly the fact that they don't study grammar and skip the "hard parts" doesn't mean that they don't pick up/know grammar, so the original argument is moot.
Gender in German, for instance. It seems to be completely arbitrary save for one detail: whichever “sounds” best. I compare this to using “a/an” in English (albeit the rules in English are somewhat more “strict” in this regard.)
His students were objectively less skilled than the other spanish teacher that focused on vocabulary and conversational skills.
At one point, I wondered if this same strategy would make for a good (human) language learning app. Kind of like you described... rather than something like Duolingo where you have to memorize answers, it might help to have a conversational app where you see the results of your attempts in real time.
This is an incredibly harmful attitude towards learning, but it feels nice because it's a lot less effort than actually trying to read or listen to something to learn. It's just laziness.
Learning C++ this way is how someone would end up with a buffer overflow every 30 lines of code they write. It's the reason some self-taught developers can't give you the fuzziest definition of the difference between O(n) and O(n^2).
The closest approximation of this is how kids learn to speak, but this is incredibly inefficient, and they receive many years of formal education anyway.
What works for you isn't what works for everybody. And I'm absolutely not denying the importance of learning theory.
I think learning programming by learning all the rules upfront and then trying to do sth isn't the best way to do it. Much easier to try to do something and iterate quickly learning what's needed to do the job. That's how people learnt programming on 8-bit computers - type a sample program and change stuff to see what it does. Then find something you want to achieve and try to change the sample program to do that. Eventually you understand how basics work and can learn everything else.
The "language" part in programming is really just a vernacular term for describing differences in syntax, and has no parallels with human spoken languages.
However, most of the people I interact with find my exposition methods curious in sense of structure. So, I shared this study with some of my friends, and they agreed with the study too. When I pointed out my surprise to them, they told me that I speak like a computer program in the sense of structure of exposition.
So, I guess I agree with this study in general, and actually would love to know what's going on in my brain.
For what it's worth, I am fluent in 4 natural languages from two different families and I'm learning a fifth language. Also, I'm fluent in multiple programming languages.
West Germanic: English
Indo-Aryan: Hindi, Bengali
Dravidian: Kannada
And currently, I'm learning German.
It wouldn’t surprise me if a lot of it feels like brute forcing syntax and keywords - there’s no way of reasoning through it, you just have to knuckle down and learn it.
It’s also part of the driving force behind my app (link in bio) - it can take a long time (i.e. years) to get comfortable carrying on a conversation, so I want to help people practise in a less intimidating environment.
Others here, and presumably your friends, have talked about how they think of programming "visually", but they don't think visually about conversing in natural language.
I thought of a concrete way to describe it: it "feels obvious" to me that any programming language could be "easily" replaced by a systematic diagramming method, with programs translated systematically into equivalent diagrams. If anything, a diagram would be clearer and more readable to me than the same program in plaintext. (Of course, it would be much more cumbersome to input such a diagram into a computer than plaintext.)
Whereas the idea of systematically translating arbitrary natural language sentences into diagrams is...nonsensical to me. I mean I wouldn't even know where to begin.
I'm curious if for you, do you feel like you could easily systematically translate a natural language sentence into a diagram? Or, because you don't think as visually as people like me, translating a program into a diagram is not "obvious" to you at all? Or maybe this distinction just doesn't feel significant to you, one language being diagram-translatable and another being diagram-untranslatable doesn't cause them to feel different, they feel equally language-y to you?
For what it's worth, I agree with OP, and have a very "language-oriented" thinking style. I certainly don't visualize anything while programming or doing math (except for geometry and the like). My thinking feels like it's more based on constraint solving and seeing analogies between domains.
I clearly completely failed to communicate what I meant about "diagramming" a programming language, as evidenced by all the people pointing out that syntax tree diagrams work fine on natural languages, and indeed were originally created for natural languages, which is totally missing the point.
What I meant was that the semantics of programming languages are so limited and constrained that they could easily be translated into an "executable diagram", such as a control-flow graph for imperative code, or a dataflow graph for functional code. The syntax of natural languages is indeed more-or-less similarly constrained as programming languages, but the semantics of natural languages seems completely nebulous and ill-defined to me.
To use the board game analogy from earlier, you could conceivably learn to play chess entirely in terms of chess notation ("1. e4 e5 2. Nf3 Nc6 3. Bb5 a6"), without ever learning about the 8x8 chessboard or the 16 pieces. Chess notation shares no syntactic structure whatsoever with the "syntax"/diagram drawing rules of a chessboard and pieces, yet their semantics are exactly equivalent.
In the same way, the "syntax"/diagram drawing rules of control-flow graphs has no syntactic structure in common with imperative code, yet exactly equivalent semantics. Could you imagine a diagram system that has no syntactic structure in common with natural language, yet completely captures the semantics?
I cannot begin to imagine that. The semantics of natural language defy description.
- Re-derive it by completing the square on the general quadratic
- Solve a specific quadratic you're interested in by factoring or graphing
- Try to find zeros with numerical methods
The point is, everything's connected to everything else, so it doesn't matter if you forget any part of it as long as you remember some critical mass.
To be fully fluent in a programming language, you have to memorize two or three dozen keywords and a handful of operators (mostly the same as standard math). To be fluent in a natural language, you have to memorize two or three dozen words a week for years. Not to mention understanding all the insane grammatical stuff.
With human languages, sure it's easy enough to remember that "nec-" means "dead" (especially if you consume fantasy novels / games where necromancy is a school of magic), but you're still totally screwed if you can't remember whether the word you need is "necare" or "necire," "necobitis" or "necabitis" or "necibetas". There's no rhyme or reason to any of it, your knowledge won't help you figure out whether that letter's supposed to be an "o" or an "i", or the difference between past tense and future tense and perfect tense and imperfect tense and tensile strength and gaaah why is this all so complicated!?
It's like if a programming language had 20 different versions of each keyword, and they were all spelled with one or two letters different, usually only the vowels are different. And they all compile, getting the wrong one just creates subtle semantic bugs at runtime. And there's a maze of rules telling you which version of the keyword you use if you're inside a loop body or a function body or a conditional statement body. Oh, and there are six different versions of the "if" keyword, and they're all two letters long, good luck getting the right one of those. And if your code's going to be executing two or more times, all the names of your variables will suddenly be different. And when you refer to a struct outside of the function that created it, all the field names are transformed. And whoever was writing the compiler put ten different special cases in the code for transforming your variable names, just because they could.
To me it feels a little deeper than just there being a "maze of rules". Your earlier mention of "everything's connected to everything else" resonated with me more.
In a sense, the expressive power of programming languages, and mathematical notation, feels very small. I liked the board game metaphor brought up in the top comment: programming languages and mathematical notation merely feel like arrangements of board game pieces (from an infinite box, and the rules of how they can be arranged are "context-free"). They don't have meaning except what we impose (ideally, assisted by comments).
A metaphor I brought up in another comment is that it feels "easy" to systematically translate an arbitrary program into an equivalent diagram, such that the diagram would contain 100% of the operational information in the program, and could be runnable as-is. (Would be cumbersome to input into the computer, ofc.) Whereas the idea of systematically translating an arbitrary natural sentence into a diagram just seems...nonsensical to me. I wouldn't even know where to begin.
Does that resonate with you?
we actually literally did this in high school. It was very informative for me.
Yes, if you try to apply it to some parts of "Moby Dick" it will be rather painful. :)
Basing a computer language on a subset of English was done for several query languages, ranging from SQL to Attempto Controlled English (https://en.wikipedia.org/wiki/Attempto_Controlled_English).
As others have said in this thread: programming "languages" are not like natural languages for a reason, and conciseness is a big part of that. The two fill a completely different role, and it should be obvious that the brain structures involved are different too.
To illustrate that: the only reason people are surprised that programming languages and natural languages are dissimilar is because our natural language is ambiguous: it calls them both "languages" even though they serve entirely different functions.
Ruby’s syntax does seem to be informed by Japanese word order, so that’s nice.
You take ideas in one form and express them in terms idiomatic to the other.
Sort of like computer viruses and biological viruses.
Legos with the all the rules of AD&D 2nd Ed.
However it's always good to have intuitions like this confirmed or refuted by actual investigation, and putting them on a more solid footing within the broader context.
My personal experience of reading code has always felt more like looking at a mechanical mechanism, or perhaps solving a paper maze where you see the whole thing at once but can focus on parts as well. It isn't something I look at one line at a time, I read it in structural blocks, almost a breadth-first reading. I see the if/else statement and structure before I see any of the details of each branch. It's a deeply non-linear process of going up and down levels of abstraction where necessary, and ignoring details at lower levels that are irrelevant to the task at hand.
What did surprise me, though, was the lack of similarity to mathematical thinking. Was, this, though, because they used fairly simple snippets of Python, a programming language which is not very mathematical in its notation? Would a different result have been obtained had they asked the subjects to reason about a multithreaded C++ program which copiously used autoincrementing pointers and polymorphism?
And, can personal preferences in programming languages be partly down to which parts of their brains people like to use?
And, as the problems were simple, the study doesn't say anything about expert programmers. Is the often-quoted 10× difference between the best and worst programmers associated with the best ones also involving the mathematical and linguistic parts of their brains?
It's probably not practical to measure programming at a high level (or indeed violin playing at a high level) inside an MRI machine, given the cramped and noisy conditions, so we may not have answers to the interesting questions raised any time soon.
I've noticed that many programmers hit a wall when dealing with Haskell or mathematics because notation that I parse effortlessly they have to puzzle over symbol by symbol, so I suspect it is a different system.
I'd hazard a guess that the results would be the same. It all boils down to (multiple) sequential processes, which can be analyzed by spatio-temporal reasoning.
Functional programming, on the other hand, may be different. When "execution" of your program is driven by beta-reductions and control flow can be abstracted away, it becomes less certain that spatio-temporal reasoning is the right tool for the task.
And, yes, I hit the wall madhadron mentioned in the sibling comment.
I find that I'm able to start writing a language and learn its grammar fairly quickly, but once I have to listen and converse in it it's like a whole different ballpark.
Furthermore, GPT-3 is only a language model because it is trained on textual data. Transformer architectures simply map sequences to other sequences. It doesn't particularly matter what those sequences represent. GPT-2 has been used to complete images, for example: https://openai.com/blog/image-gpt/
Transformer architectures do map sequences to sequences. What is not known is that the task of programming is a sequence problem. This experiment seems to suggest that maybe its not a sequence problem.
Unlike human languages: English (and probably most others) has hundreds of words that can be used to nuance, imply, suggest, refine, gradate, distinguish, obscure, conceal, intimate, impute.
If we don't know what they mean we can often skip right past them. (Computers don't do 'gist'.) Illogic is not only permitted, it's often essential.
Are they? Do you say this as someone who is familiar with actual laws? Genuinely asking. Because I know that engineers (my past self included) like to imagine formalizing laws and turning them into code, which is extremely naive, once you see actual laws and how much human common sense judgment is needed to decide on them, including inferring intent, judging potential consequences, interpreting imprecise terms, etc.
Perhaps a part of the tax code is like code. But a lot of laws are far from it. It's more formal than novels but it's firmly natural language that you need to understand as a human.
I'd also compare these two to mathematical proofs written in natural language in textbooks or academic papers. They have recurring structures and patterns, but are not formalized proofs, leave things explicitly or implicitly as an exercise to the reader or assume cases as trivial. (The gap to a formal representation is often nontrivial effort and has never been totally done for most of math)
But then again, we have lots of similar things. Household equipment user manuals, revenue reports, soccer match descriptions in newspapers, stock market updates, phone calls to make an appointment at the dentist, etc. Lots of things on various points of the structured/formal/closed -- unstructured/informal/open spectrum.
The French tax code is peculiar (though I'm not sure its unique) in literally being defined by the rules encoded in an official program. A recent paper on a new language for compiling those rules, "A Modern Compiler for the French Tax Code" (https://arxiv.org/pdf/2011.07966.pdf, via HN a couple of weeks ago), contains a blurb at the very end,
> Closer to the topic of this paper, the logical structure of the US tax law has been extensively studied by Lawsky [18, 19], pointing out the legal ambiguities in the text of the law that need to be resolved using legal reasoning. She also claims that the tax law drafting style follows default logic [24], a non-monotonic logic that is hard to encode in languages with first-order logic (FOL). This could explain, as M is also based on FOL, the complexity of the DGFiP codebase.
See https://en.wikipedia.org/wiki/Default_logic and https://en.wikipedia.org/wiki/Non-monotonic_logic But note that it's written in that "style", not that it's done consistently or correctly. Unfortunately I couldn't easily find a copy of the cited papers, but I made a note for later reference.
EDIT: I once took a class in law school taught by a Professor of Systems Engineering from the parent university. I forget why, but long ago he took an interest in the rules of evidence in the common law. While the rules evolved organically over more than a thousand years, they actually strongly reflect a formal system of abductive reasoning: https://en.wikipedia.org/wiki/Abductive_reasoning. And there's significant, hidden rigor. In one project he had everyone in class create a tree consisting of evidence (e.g. size 7 shoe) at the leaves, with the branches as inferences all the way up to the top, which was the core legal claim--1) Jim did 2) break and 3) enter into 4) said building.... Anyhow, being a programmer I decided to create my tree using the DOT language and render it using Graphviz. While I was doing that it occurred to me that some of the more esoteric, technical rules of evidence, such as that any piece of evidence can only be used to prove a single element, actually ensures that the tree is an acyclic graph. I'm not sure precisely why (perhaps the professor had more insight), but I believe it's because the rules are in part crafted to ensure that the finder of fact is presented with a tractable problem to solve. The process would get too confusing (and possibly logically inconsistent?) if you didn't keep the tree of evidence and inference simple. No judge or lawmaker likely ever had in mind the literal shape of the evidence tree. In fact, AFAIU nobody thought to represent it as a tree until the late 19th century, when the famous scholar of common law evidence, John Wigmore (https://en.wikipedia.org/wiki/John_Henry_Wigmore), developed a system for analyzing and preparing trial evidence in that manner.
Anyhow, a proper and infinitely more rigorous treatment of this is presented in the book, "Analysis of Evidence, 2nd Ed." by Terence Anderson, David Schum, and William Twining. (Schum taught the class.) A large part of the book explores and explains the law of evidence as a system of Bayesian reasoning. (Oh, that's why! I think Prof. Schum first stumbled into the legal realm after he got the idea of using Bayesian statistics to resolve historically famous and contentious legal trials.)
Math is not concerned with implementation details.
!falseThe language is often a reflection of the problem it was intended to solve. Would you really argue that Brainfuck can solve your problems just as well as Python?
I don't see meaningful differences between languages whether they are computer languages, human languages, math notation, except how the human brain processes them.
>linguistics research at the time which defined language as a set of strings.
Fundamentally there is nothing wrong with that. You can have trivial languages that follow simple rules. In practice humans tend to use more complicated languages but that is just their personal and cultural preference.
People here are too hung up the message that is being communicated rather than the general idea of language as a communication method. Programming languages usually express algorithms and trigger logical sections of the brain but nothing prevents you from using a human language to express the same algorithm and trigger the same sections of the brain. The reason why people associate programming languages with logical thinking is that they were exclusively designed to communicate logic and nothing else.
They name Python and ScratchJr, but no hints on what spoken languages were examined. I've been told Chinese and Japanese are much closer than Romance languages to code.
But a language is not its script.
By approximation all modern languages that have been written have been written in the Latin script. That doesn't make them "similar languages" in any real sense. And conversely, many languages use several different scripts for the same language in different regions/groups depending on politics.
That said, I like to understand a program (especially my own) by explaining it in English. There, I'm using generative, rather than recognition, processing. There might be some fruitful research in asking people to explain how programs work, rather than what they do, under fMRI.
Or programming system in analogue of writing system. Becomes progsys for short.
I know - why not just call it code? That's what we already do. We can just make it the preferred term.