On the foolishness of “natural language programming”
cs.utexas.edu
cs.utexas.edu
Most people who have never written a computer program, have probably never even been through the experience of having to express exactly what they want someone or something else to do for them, in a specific and non-ambiguous manner. It really is a different way of thinking.
A really great formal language actually frees you from "every detail, no matter how obvious". That's the heart of abstraction.
Good mathematics is heavily abstracted because it forms a superhighway between mathematical concepts that allows you to lift and transport intuition from one context to another without forsaking formality.
To some degree these abstractions must always be built from nitty gritty details. In mathematics it's a matter of didactics to find the most efficient bootstrapping process. In computer science it's a matter of library and language design.
Didactic context is personal, the context required for or used by a particular person in order to conceptualize and build confidence with a set of concepts (and their associated syntax, though it's close to meaningless so long as it doesn't get in your way too much).
Mandatory context is more like a compact notion of why and how some mechanism is applied. For instance, you might bootstrap homology using point-set topology (educational context) and then enhance it using algebraic topology. Eventually, perhaps, the mandatory context is that homology is a measure of non-exactness of chain complexes and this concept can be lifted from its base context and presented wherever exactness is an interesting measure.
Point being—language hardly exists in a vacuum. Good language invokes powerful concepts in an efficient manner and allows you to construct arguments using them. Good arguments can link seemingly disparate concepts with ease.
I.e. it can be very productive to do the reasonable thing, until it's proven unreasonable.
Why else does every coding language provide plain-language comments, but legal documents do not?
In technology, the machine is a stand-in for a human relationship - the cell phone app keeps track of restaurant ratings so people don't have to keep asking one another. The code is a part of the system of distributing the power in the form of information.
In the legal world, the intently obscure language is the means of hoarding and securing the power.
W/r/t to legal systems collapsing, they're protected by guns and myths. The baillifs and police believe Columbus was a hero and they kill to uphold his dream. The incentive to power is enough to entice enough new scholars each year to sustain the illusion.
So to respond to the poster above, the pains of the programmer and legal system are no more similar than the programmer and the chef. Each attempts to communicate a procedure, but the legal system has an entirely different and less noble set of requirements.
This is true enough in many parts of the world, but a careful reading of the history of English Common Law reveals a fascinatingly iterative legal process that formed a bulwark against monarchal oppression for hundreds of years.
For a good introduction, check out this book:
http://www.amazon.com/Inventing-Freedom-English-Speaking-Peo...
Constitutional law is under constant iteration, the Supreme Court often makes decisions about which cases to take based on whether it thinks it will be able to advance the state of that area of law.
If you subscribe to blogs like Popehat, you can get a feel for this iteration/evolution process and how it works in modern times.
The history of English common law is far more akin to natural evolution than any intelligent design.
It's the same with governments and people -- history proves that relationship red in tooth and claw as well. The surviving elk outrun the wolves, and the surviving people stay one step ahead of predatory governments.
When Stalin signed the non-aggression pact with Hitler, he began a purge of his military. Anyone who had spoken against Hitler beforehand was purged -- shot or sent to Siberia. Then, when Hitler broke the pact and invaded Russia, all those officers who had spoken in Hitler's favor were purged. The survivors were those few who didn't have an opinion, or who didn't dare express it.
The Cultural Revolution in China purged all those bourgeois elitists who had a college degree or who had acquired skills like science and technology, or any significant academic achievement. Now everything has changed and individual Chinese are allowed -- nay, encouraged -- to educate themselves for success, acquire wealth, and grow the economy -- exactly the opposite of the Cultural Revolution outlook. It is a very wise person who avoided any problems during the Cultural Revolution, and who avoid any problems now.
My point? Elk who survive do so by avoiding wolves. People who survive do so by avoiding governments. It's true that wolves improve the genetic stock of elk, just as governments improve the genetic stock of people, and by the same method -- by tearing the weak and sick to pieces.
Even if government is the best way to organize society at the moment, that doesn't mean it's unreasonable to look for better alternatives. But again, none of this has to do with the analogy to intelligent design I was making.
So are many of the software projects I've been called on to maintain.
"The computer is a machine, but a codebase is an organism."
http://www.meltingasphalt.com/technical-debt-of-the-west/To be fair, though, I think you see technology and code in too positive a light, because code is also often a means for accruing power.
Your own example of a cell phone app that keeps track of restaurant ratings could serve to illustrate this quite well: such services have a tendency to become highly centralized with at most a handful of alternatives. [0]
Whoever controls the rating service then has considerable power, because they get to subtly influence the way that results are displayed, which can directly influence restaurants' bottom line.
This type of thing is even more obvious with services like Google and Facebook. Just something to keep in mind before patting oneself too much on the back...
[0] This is due to network effects: people use the service with the highest pervasiveness (basically, the most ratings and comments), and the service with the most users tends to get the most ratings and comments.
That's only surprising if you consider the language of laws to be of paramount importance. But generally, that's not the case. Most legal outcomes just come down to the discretion of police, prosecutors, judges, and juries, in descending order of prominence.
A lot of people, and probably especially programmers and engineers, think of the legal system as a formal system. It's quite common, when discussing some court case, or perhaps some unscrupulous act from a government official, to hear people say "but isn't that illegal?" The surprise implied by that question only makes sense if you assume either that a legal document has actual power to prevent an act (which we know it does not), or that the people who decide legal outcomes are somehow bound by the meaning of a legal document rather than the other way around.
It's not an impossible task, just a difficult one.
Case in point: the debasement of the term "literally".
* It's possible since you can always simulate a human brain with software, but there are of course more practical artifical general intelligence systems like the AIXI-mc.
Sorry, but this is a completely shocking article to me. First in it's immediate dismissal of any formalism inherent in natural language, but also with the ease at which he dismisses the proposition without any real consideration.
If we learned anything from Chomsky, it's that the underlying grammar we are born with is both instinctual, and follows formal rules. To say differently is literally, demonstrably false. Irregular verbs, for example, aren't learned in the traditional sense, one must actually unlearn the formal rules. Any child that tells you she "swimmed" all afternoon is using a more formal version of English than you do.
The idea of a natural language programming language is flawed, but not by formalism. It's flawed by the evolutionary nature of natural languages. That is, the very people that he states "are no longer able to use their native tongue effectively," are probably using a new dialect that shares a common ancestor with his more "traditional" usage.
Many people in this thread are talking about the inability of plebes to express what they actually want. This is a fair point, but not a problem with language specifically. Communication tools would need to be employed by a computer in the same way as humans use them. E.g., a simple ambiguity checker could work wonders here, as it does between humans when someone you are talking to simply says, "What did you mean when you said you 'realized you forgot your phone at the train station'? Did you forget at the train station, or realize at the train station?".
What IS a problem, however, is Quine's indeterminacy of translation. That could pose serious hurdles that may be insurmountable, however, we still have effective communication between humans, so it's simple to see how this may only be a theoretical problem, rather than a formal one.
This subject should be under the purview of analytic philosophy and linguistics, not mathematics or computer science.
It seems an appalling number of people agree with Dijkstra here - "we not only don't need natural language, we don't even want it."
I'm sorry, but if I can express to a human being a set of directions to fill out a form in a minute or two and expressing that to a computer takes much longer and is more error-prone, that is an inefficiency in software development which it is extremely desirable to address.
There is nothing magical about human brains that would make them theoretically impossible to express in software in such a way that we can give a program natural language directions, and the fact that so many people want to dismiss this endeavor out of hand is ridiculous to me.
This should be our holy grail, something to strive toward, not something to ignore.
In fact I would suggest that it will most likely be the only way out of the mess of such a wide variety of software standards (ever have fun moving your things to a new system and having to re-learn many things just because you changed PC or phone operating systems?), whereas natural language is a standard we already have and works fine.
This way you essentially wouldn't need to learn a new set of incantations - just tell the damn thing what you want the damn thing to do, dammit.
Ever moved to a new job and had fun figuring out the local terminology, and learning the in-jokes (and figuring out when you were the butt of the joke but didn't quite get it), and when your boss told you to "just do your damn job and quit pestering me", did you immediately know he was just having a bad day and it wasn't your fault at all?
A mistake in understanding these "natural language incantations" will lead to you doing the wrong thing, or feeling embarrassed or depressed. Not so different from a system crash, when you think about it.
Thankfully machines never feel embarrassed or confused about the "damn thing" you want them to do. They only do the things we tell them, in a painfully literal way.
After all, when people want to be really precise, they use mathematical notation :)
Certainly not perfectly of course, but go and ask people who use some sort of CRUD system that used to be manual what they like and what they dislike about the new system.
Likely very few of them will say it is easier to use than just asking Sharon in the next cubicle to verify something.
They will have some likes, sure - but they will be things like the almost infinitely increased speed or the fact that they can access the system 24/7 while Sharon needs to sleep and take a break to eat the occasional Ding-Dong, but the ease of use that comes with dealing with another human being is a sacrifice that they make for these benefits.
(or to put it the reverse way, imagine the same office and a new employee named Eliza comes in and behaves exactly like a computer, "only doing the things you tell her in a painfully literal way." How quick would you want to give her the boot?)
> After all, when people want to be really precise, they use mathematical notation :)
Again, this is absolutely true. The problem is when you're dealing with a simple CRUD app for your insurance house or just copying files over to your iPod, you're not interested in being really precise - you're interested in the shortest path to get a relatively simple thing done.
If that path is blocked by the fact that you don't know the particular menu item, keyboard shortcut, or command switch for something that you can express in English without even thinking about it, then I regard that as a huge opportunity for technology in general and our industry in particular.
I also agree that a lot of people tend to shy away from formalism. I think that's what Dijkstra was lamenting.
By the way, what is a CRUD? John Doe doesn't understand the word. What do you mean, Create-Read-Update-Delete? He understands some of those words, but I'm unsure they mean what he thinks they mean. Are you sure that, without some training on the formalisms of the system (which Sharon obviously had!), you want John Doe to delete something from the system? He might try to unplug the harddrive, maybe that's what he thinks "deleting" means.
Yes, it's easier to teach John Doe to use a limited UI instead of, say, teaching him SQL. But he'll be able to do less complex stuff with just the UI. (And the fact SQL has some English-sounding keywords is helpful, but SQL is an extremely formal system with few parallels to natural language).
Highly unlikely that someone familiar with the system (even in an informal way) would do this, for the same reason that you have no trouble understanding me when I say "fruit flies like a banana". Is there technically a chance that I mean "all pieces of fruit fly through the air in the same manner that a banana flies through the air"? Sure. But it's so ridiculously low that you simply ignore it and are willing to accept the infinitesimal risk of misunderstanding.
Should all programming or user interface work be done this way? Of course not. I'm just saying it would be a very effective level of abstraction for a great many use cases, in the same way that I don't need formal mathematical notation to write a Python statement to print "hello, world".
Of course, he formalism has to be there once you get down far enough, in the same way that showing Sharon how to do an account credit in the CRUD system means that you are altering neural structures with electric and chemical signals.
However, the person who trains Sharon doesn't need to know exactly what neurons to stimulate in Sharon's brain with exactly what voltage in order to teach her that system - the relatively lofty level of abstraction provided by English works just fine.
Allow me to add some random thoughts:
- Python's print "hello, world" IS a formal notation. It's just that this particular notation and this particular task are so simple that we can delude ourselves into thinking it's English. But when you move to actual Python scripts, the only ones who believe "it sounds like English" are programmers :) I wouldn't trust my mom to write a Python script, after all.
- Let's go back to our CRUD/office situation example, and allow me to make it a bit more realistic (but still funny):
"Sharon, please print the report."
"Which report?".
"The one I asked you about yesterday."
"Uh, you asked about two reports yesterday. Do you mean the one about fruit flies or about bananas? Or do you want both?"
"Yes."
"Sorry, yes what? I asked you multiple questions!"
"Yes, both reports. I forgot about the other one, but I want it too."
(...)
"Ok, even though I have a terrible headache, I printed your reports. Here they are."
"Oops, sorry Sharon. I didn't mean you had to print them now. Tomorrow would have been fine. Also, please don't get mad, but I didn't want them on my desk. They are actually for Jane on the fifth floor... Didn't I mention that? Also, why did you print them using the expensive printer?"
----
My point is that just doing CRUDs with English is probably fine, but as the complexity of the task approaches that of a general purpose programming language, the level of precision you must use with your language approaches that of a formal system. Which is what programming languages are...
There is actually a theory of how two entities reach agreement through conversation: http://en.wikipedia.org/wiki/Conversation_theory
We can explain to a human how to fill out a form because the human has probably seen a lot of forms before. Humans make a lot of assumptions that often end up being wrong: but enough of them are correct that they're still useful.
We trust computers to be unimpeachably accurate because we as humans are not. If computers need to make the types of assumptions that humans do, then they lose a good deal of their accuracy (and their usefulness).
Human language is also visually difficult to read. The biggest improvement that symbolic languages (specifically, modern programming languages) make is the use of spacing and symbols to break apart complex processes into sub-sections, loops, etc.
Furthermore, I disagree with your statement that "natural language is a standard we already have and works fine." Language is not static, nor is it standard. Sure, we may have "standard" grammar rules, but even those can vary from region to region and many people don't follow the rules on a day to day basis. It's not a static target, so developing something that could interpret natural language means developing an artificial intelligence capable of taking nuance, context and the like into account.
EWD was simply claiming that such a system applied to general purpose computing would be so complicated as to be wildly impractical.
I'm going to guess that you've never given a set of requirements to a programmer before. Programmers ARE the human interface to computers, as they are often writing software to specifications created by someone else. Many of the problems arise in the ambiguity of the human-human communication. Another part comes from the lack of specificity combined with different ideas about how to handle unspecified cases. Your idea of what is obvious is not the only one. Some people lack domain knowledge that is assumed in requirements and leads to poor choices where specifications are not complete. In the end, natural language assumes a broad swath of "common sense" that computers do not have yet.
> In the end, natural language assumes a broad swath of "common sense" that computers do not have yet.
Absolutely. I'm not at all saying that we'd have had this last Tuesday if we'd just take our heads out of our asses - I'm saying it's something we should strive toward and not ignore.
> Many of the problems arise in the ambiguity of the human-human communication.
Human-Human communication works, and works well - once again, we do it every single day, all the time.
Do we encounter problems with ambiguity? Sure. But they are by far the exception and not the norm. After all, forms get filled out, Driver's licenses get renewed, complicated Starbucks orders get filled - these common use cases work.
By contrast, have you ever had this fun experience with a terminal program?
> quit
Unknown command: "quit"
> exit
Unknown command: "exit"
> shutdown
Unknown command: "shutdown"
or my favorite:
> quit
Unknown command "quit." If you want to close the program, type "exit."
Simple English statements like these work extremely well in human-human communication and do not work hardly at all in human-computer communication. I'd just like us to get from A to B, that's all.
I think we are still a long way from understanding what really happens inside our brains. About telling a computer what to do: They don't 'do' anything. They run programs. (Free after Weizenbaum)
Agreed 100% - I'm in no way saying that it's easy or will be done in our lifetime or the next 10 lifetimes - just that it's not impossible because we're not magic and we do it every damn day.
> About telling a computer what to do: They don't 'do' anything. They run programs. (Free after Weizenbaum)
What I want is (to reference an example in another comment) to go to a command line and type "Copy the report to the share" and (if I'm on a linux box) have my computer translate that to "cp /path/to/report.pdf /path/to/share/" without my ever having to know what it's doing behind the scenes.
Whatever categorical bucket someone wants to put that in doesn't matter to me whatsoever - all I'm saying is that's what I want to see, and that's what I think we should work toward.
When will we get an OS that lets me persist my objects, uniquely identify them with a uuid plus arbitrary attributes (print date??? give me a break), migrate and cache them anywhere and sign them for authenticity? That would be a real OS feature.
Sure all that can be cobbled together on one machine with different libraries. But to be an OS feature, I need servers that understand and respect all that. Object browsers that let me create a relation to view pertinent objects. Security managers that limit access to apps with digital authority etc. All on the network.
It's not a new idea, nobody ever was able to make that kind of storage work.
As for making it work, there's no obstacle. Implementation is straightforward. And since any current file system API is trivially implementable on top of it (create a relation using parentDir, filename, {dates}) there should be little integration issue.
Error: line 2: 'damnit': no denotation
Just kidding. I agree: pace Prof. Dr. Dijkstra, automatic programming is a grand challenge problem and deserves more attention.Gathering requirements is generally done in natural language, but it's a very slow error-prone process. Even after sign-off on requirements it's pretty standard for them to be wrong in critical ways. Frankly this is the part of many software development projects that dooms them to failure.
Ignoring the difficulty of actually getting precise natural language, you'll still get to the point where no one can understand the language. If you don't believe me, go read some Kant.
I'd appoligize for the rambling comment with grammatical and spelling mistakes, but they further my point :)
Indeterminacy of translation is indeed part of the problem. But you're not saying anything new by bringing it up - it's just a jargon term from linguistics to describe a problem a programmer might illustrate with the (buffalo)+ sentence or "(time|fruit) flies like (an arrow|a banana)" example. They already know the core of the problem without needing the whole weight of a linguistics education.
The obvious problem of indeterminacy of translation is why, when computer scientists talk about natural language programming, they do not normally mean natural language processing - even more so in Dijkstra's time, when computers were slower and our NLP algorithms worse.
The core that Dijkstra is getting at here is symbolic reasoning. He's pointing out that natural language is a poor fit for symbolic reasoning, that there's been a history of movement from rhetorical reasoning to symbolic reasoning in mathematics - in fact that mathematics stagnated where rhetorical reasoning persisted.
Even if we solved the translation indeterminacy problem, we would need Strong AI to convert such a high level description into something concrete enough for a computer to do. We propel computing machines using levers made out of abstractions - the higher the tower of abstraction, the longer the lever, and the greater the power. But the problem of programming is not in pushing the lever, it's in building the lever. In a word, it's engineering, not philosophy - it's about how, and not what.
Firstly, the indeterminacy of translation is a much, much deeper problem than the buffalo sentence and the time/fruit flies sentences.
The buffalo sentence is interesting in that it is grammatically perfect, and would be easy for anyone to understand following formal rules, but so bizarre that it's confusing at first.
The time/arrow sentence is also not an issue for humans as it's simply the result of two homonyms that happen to be verbs for one definition. They are clever, but not deeply ambiguous, and i seriously doubt they would pose a serious problem.
Second, natural language IS symbolic reasoning. I don't understand how you can justify a claim that it isn't. The term rhetorical reasoning presumes that symbolic reasoning is already happening.
4 + 4 = 8.
Let there be a function f taking one argument, returning a result that is the argument multiplied by two.
f x = x * 2
Do you see the difference between rhetoric and symbol? The point is, perhaps, more literal than you suspected?
Of course you can argue that the 4 number, is more abstract than the word, and that the variable x is more abstract, than the 'concrete' number 4 ... even so, you can build arbitrary complex structures with just 1 level of abstraction/indirection.
In this sence, I agree with scoofy: '(usage of) natural language IS symbolic reasoning' indeed.
It's trivial to define a mapping from the symbolic to the rhetorical. The reverse isn't trivial, but even if it were trivial, we'd still prefer the symbolic.
It's easier to read.
If that were generally true, we wouldn't have bothered to invent mathematical symbolism and syntax. Natural language has any number of pitfalls -- "are you going to sleep or watch TV?" The "or" in that sentence differs in meaning from the formal logical "or" which would suggest doing both at once.
If I want to express the difference between time and space in relativity theory, I might say:
t' = t √(v^2/c^2)
With that equation I have said a lot, referring to the Pythagorean Theorem, orthogonal dimensions, and the constancy of the speed of light. To translate the above into natural language would require many more symbols, as well as the acceptance of much more ambiguity of meaning.
It's uncontroversial to say that, as we get closer to describing nature accurately, we use more equations and fewer words, and not because of an irrational preference for equations.
The rejection of this idea is not only not "demonstrably false", it's actually pretty commonplace among linguists. It would be more accurate to say that Chomsky's theory of a universal grammar is demonstrably false.
"the evolutionary nature of natural languages"
One of Chomsky's more widely criticized ideas (and a pretty bizarre one), is that the language instinct could not have arisen by evolution through natural selection.
By "the evolutionary nature of natural languages," i do not mean natural language as a product of natural selection. Instead, i mean the constant flux of natural language dialects. The loss of some words, phrases, etc. over time with the adoption of new ones. "Reprise" has been replaced with "remix," "how do you do" becomes "howdy" over time. This isn't exactly controversial.
Skinner and Chomsky both came up with some ideas about language acquisition that were useful and stimulated subsequent research. They are both quite wrong, in the sense that neither theory withstands the scrutiny of research over the past 50 years in anthropology, psychology, and [computational] linguistics. Pinker is one of the more prominent critics of Chomsky, who also points out the value of some of his ideas.
In stark contrast, Darwin's theory remains largely intact, with some modifications and additional insights. It's been magnificently confirmed by discoveries in genetics, of which Darwin could have known nothing. Darwin was right and created a successful theory that will live forever, even as it's added to and enhanced. We have nothing like that for a theory of language acquisition, neither from Chomsky, Skinner, nor anyone else, yet. On the whole Chomsky was not more right than Skinner in any ultimately significant way.
Sorry for misapprehending what you were getting at with evolution of language. I think the idea that human language changes over time is not unfamiliar.
Could we generally agree that language is instinctual and follows formal or formalalizable rules? If so, then i'd say my point still stands.
No, not without some evidence.
----
> First in it's immediate dismissal of any formalism inherent in natural language
That's not what Dijkstra's article is saying.
> the ease at which he dismisses the proposition without any real consideration
Dijkstra provides several paragraphs of clarification. IMHO, these do provide useful context and show the reader why he has the opinion that he does (of course, we may just disagree here -- but that's okay). What do you feel is missing?
> This subject should be under the purview of analytic philosophy and linguistics, not mathematics or computer science.
Why?
> To say differently is literally, demonstrably false.
Why? Can you demonstrate that, or point to a source that does?
> Many people in this thread are talking about the inability of plebes to express what they actually want. This is a fair point, but not a problem with language specifically.
Why not?
> Graduate of analytic philosophy
Possible appeal to authority.
> This is not up for debate
Why not?
> it's easily provable.
Then please prove it, or point to something else that does so.
> I could site dozens of linguists, Pinker immediately comes to mind offhand.
Then please do so, and please be more specific than just giving us "Pinker". The more specific you can be, the more useful the citation is. I (and presumably others here) are not familiar with his work.
Nouns themselves can take a number of different meanings (specifically proper nouns and names) depending on context. Communicative languages often don't distinguish between equality and identity; a critical distinction in computer science (i.e. I can truthfully say "I've eaten at same restaurant 3 days in a row" if I ate at 3 different McDonalds locations on consecutive days). These assumptions about identity and equality are not the same from language to language (or even generation to generation, as you discussed). Context is also an issue; we often discuss things in ambiguous contexts which often require clarifying questions from a human. We may have effective communication between humans, but misunderstandings are common. IMO, language is an effective communication tool almost specifically because it is imprecise: we fill in the gaps with our own experiences and it mostly works out.
Overall, I don't think EWD was saying that natural language programming would be impossible; just that the effort required to program in it would likely be more than learning a symbolic programming language. The computer would need to ask so many clarifying questions to reach the level of specificity required for computer science that it would be a very arduous task. Rather than making computer programming more accessible, natural language programming would make it substantially more difficult.
Equality and identity are essentially the same with being verbs. Your sentence is an error of clarity, you'd fool a human in the same way as you'd fool a computer. I believe the ambiguity you are going for is that of "123 Main St is McDonalds" (identity) vs "123 Main St is a McDonalds" (subsumption, also called class-inclusion). Here however, we have the "a" which clarifies that the statement means subsumption. Why? because it would fool humans otherwise, and so we created a rule for clarification.
I certainly think a natural language programming language would need to ask clarity questions via a parser before running a program, but i honestly don't think this is a serious problem.
It would really only be a problem with intentionally ambiguous words and phrases, such as the verb "to hook up" which is intentionally fuzzy enough to allow the speaker to mean one thing, while allowing the speaker to imply something else.
Certainly we could come up with linguistic rules to differentiate them, but for the sake of convenience (or maybe because each of us possesses an incomplete knowledge of linguistic rules) we may refer to them as being the same. But this problem becomes much less difficult in a symbolic language, where many languages have defined an identity operator (usually ===) and an equality operator (usually ==).
You're right that we have language rules; but many of these rules are arbitrary and mutable in their own right (and do change over time as your previous comment pointed out). The problem starts to come in that language rules are a lot more complex and subject to multiple different (and potentially correct) interpretations.
I side with EWD on this one: the interpretation of natural language is going to differ from person to person and as such, building a system that can perfectly discern intent is going to require so much context as to be impractical. It's just easier to program a computer using a symbolic language, while building in well understood formal rules to the symbolic languages that cover the vast majority of cases.
It's not that it would be impossible to develop a general purpose compiler capable of interpreting natural language to machine code; merely impractical. Symbolic languages are much better adapted to these tasks; just as natural languages are much better adapted to conveying uncertainty.
Now, this is sort of a problem, but it's not as though we are walking around confused all day. We make it work, we formalize, we come up with a perscriptivist framework, even if we live in a decriptivist world.
We can do this with any issue brought up with being verb "problems." Our Julius Caesar issue is fixed by a simple convention (yes we have conventions in English as we do in formal languages). Getting people to follow the rules is the only problem, and it will eventually fail, but that doesn't mean we won't have a few hundred years of it working, possibly more.
The issue of practicality is an entirely different question. Currently, it's an absurd proposition, that isn't to say however, that it isn't doable. Especially with greater computational power in the future. I think saying it's impractical is a cop out. The amount of energy that goes into learning languages is immense. If we were able to learn one syntax, nearly identical to English, that could convert instructions into readable code, we'd enter a renaissance of programming akin to art after the photograph was invented, or the current renaissance of music we've created with the invention of the synthesizer and sampler. Skill would evaporate, ideas would reign. People would complain that proper rules aren't being followed like they always do, but the amount of programs produced would expand so far that the cream would rise to the top and the world would be a better place.
If, for security reasons, you wanted to shackle an AI from making and acting on its own choices, you would need to shackle its ability to interpret language as well because they are the same thing. You have to make choices about implied intent when interpreting language; and those choices. You can't just restrict a machine to making choices on linguistic interpretations if those interpretations then lead directly to action (in the case of natural language programming).
We can (and have) created sets of natural language interpreters for specific situations: Siri is a good example of that. But by and large these are hacks that flag specific situations (such as creating a reminder or opening an app) and pick out the relevant phrases and plug data into fields.
I might ask my friend to move the report he's working on to a shared network location so I can load it into my computer and read it: "Hey Joe, can you move the report to the share?"
Joe might ask the computer to do the same thing: "cp /home/joe/reports/cool_report.pdf /network/share/reports/cool_report.pdf"
The actual ideas that are communicated are very similar, but not the same. English is good for communicating one idea while bash/GNU is good for communicating the other.
Just because English has some established formalism doesn't mean it's good at communicating the ideas we want to communicate to computers.
BTW, I don't care which field you put the issue under; it's the same issue and anyone who cares about it might contribute to the discussion.
"Hey Joe, can you move the report to the share?"
is a great way to communicate something you want done. it doesnt matter if its to a computer, a person, or a dog. If it can't operate on those bounds then its not sophisticated enough to actually meet the needs of the user. One day computers will get there, they haven't done so not because its a 'bad way to talk to a computer' but because computers have not yet become that sophisticated.
It might as well be a different language in a lot of ways.
> Remark. As a result of the educational trend away from intellectual discipline, the last decades have shown in the Western world a sharp decline of people's mastery of their own language: many people that by the standards of a previous generation should know better, are no longer able to use their native tongue effectively, even for purposes for which it is pretty adequate.
Because though they are an expert in their small field, but no more informed or unbiased than your friend at the pub, they end up sounding dumb.
In this and some of other writings, he would do well to think and gather data about social issues, or talk mostly about algorithms.
When someone who is famous talks about all sorts of things, it's fashionable to diss and mock them for "speaking outside their authority", yet everyone else is free to ramble about whatever they like.
You're doing an inverse arugment by authority here; "whatever Dijkstra says about [not-programming] is wrong and should not be listened to, because of who the speaker is".
The man in the pub may have a good point. If you read Djikstra's other essays where he wanders off subject, you'll motice he rarely does. He sounds like an angry old man, lost in dreams of a better past that never was.
While I can't speak to total trend (measurements of literacy have changed, as have the prevalence of testing), it's fairly clear in a qualitative fashion that total reading capabilities have declined over the last 100 years. Examine pulps (cheap entertainment books) from the late 1800s, along with childrens' books of the time... significantly more complex paragraphs and much larger vocabulary.
I'll also note that there's a distinction to be kept in mind between literacy (mastery of the artificial skill of using the written word) and fluency speaking one's native tongue (a skill all mentally healthy human beings pick up naturally).
First, http://nces.ed.gov/naal/ is the organization I was doing my reading with. Their surveys span the last 20 years. There's a 200-page PDF describing the surveys.
Second, "literacy" is, as you point out, a nuanced term. NAAL has broken it out into 3 categories with 4 rankings possible from Below Basic to Proficient. Based on my reading of the survey questions, I drew the line of "literate" as "proficient".
So, http://nces.ed.gov/naal/kf_demographics.asp is a summary page, and " Percentage of adults in each prose, document, and quantitative literacy level: 1992 and 2003" is the graph.
You can see that the percentages are not collected into a single number. I made the ENTIRELY GROSS assumption that proficiency in one area probably leaks into the other areas. I would like to calculate the actual "Total proficiency" score based on the data though.
At any rate, the percentages are abominable.
When you're dealing with variables, "morning star" "evening star" problems are essentially irrelevant because you are defining things, rather than merely naming them.
'Morning'/'Evening' star won't matter until it does - and if it were something that wouldn't ever matter, you probably wouldn't have input it.
With naming, we can be wrong, because of the inherent negative knowledge of our framework. That is, in science, we can only prove with certainty that things aren't the case, not that things are. This is Karl Popper thesis.
With defining, we cannot be wrong about things, we have the ability to have perfect knowledge, so the things we are referring to are defined to be such, we simply cannot be wrong about their identity and structure.
People comparing natural language programming to EULAs are missing the point entirely: natural language input guides a search process for formal programs, it isn't literally "the program" itself.
In other words, you'd need formal language to be able to specify what you truly want, then would "translate" it to natural language to execute the search, the processor would search for the proper formal language expression, and then you would verify.
Or you could just skip all the intermediate steps.
Anyway I don't think we're disagreeing with each other: there is always going to be a place for formal languages, they won't ever go away and be "replaced" with natural language interfaces, it's just that Jane Random will be more than happy to speak some gibberish to her computer and let the machine figure things out.
If the computer changes state it's because of the execution of a program, and if no programmer actually wrote that specific program then the act of creation can be considered "programming".
One alternative I've seen is what happens with artificial intelligence: if we can do it, it's not AI, just "algorithms" or whatever. In reality the definition of "artificial intelligence" is broad enough to cover the things we consider mundane today, like movie recommendations and search engine results.
If MS Paint was souped up with things like repetition and conditionals, would you be so opposed to calling that "programming"?
Natural Language can be seen as a very efficient compression algorithm. The phrase the speaker chooses is the shortest message, given the context of the receiver.
Programming a computer with Natural Language is incredibly difficult, because Natural Language alone, without the context it is built upon, really lacks much of the information the computer needs to operate the program.
Legal documents contain the most precise language we humans can create that can still be considered to be "natural", and they're still all but unreadable without the proper degree and legal context.
Not precisely what I would call a "win" for the ease of programming
Another time a contract had a weird compounded interest increase in percentage ownership that, the way it was written, kept increasing the percentage every year so that by the time the contract was over the other party would have something like twenty million dollars more equity than I did!!! I didn't trust the lawyer that wrote it (again one of the top lawyers in town) so I walked away from that deal, ignoring the claims that it was just a mistake.
Reading 'round trip' translations via Ithkuil is interesting.
Our father in heaven
hallowed be your name
May the [metaphorical] environment which fosters/sustains your rule eventually [metaphorically] permeate us
As for your aspirations, may they be made real
on Earth, not just in heaven
please be one who enables us to eat and drink our critical sustenance
And forgive us in regard to our moral transgressions in the way we grant and receive forgiveness amongst each other
may we successfully avoid having desires [that are] against our better judgement)
and be one who enables us to successfully avoid ideas associated with the Devil).
Without poetry, things are more clearly weird.
For natural language programming to work, computers simply need to be capable of making prudent guesses - so long as we show the information clearly, we can always fix those guesses. We can allow for limited instances of constructed languages, rewiring connections between words, etc.
You mention that such references are limited in nature, but laws are rarely as limited in their references as EULAs are. Plus, there's a whole layer of context which exists outside the written document - the exact legal definition of words.
I just don't think that a "prudent" guess will be enough, especially as you start distributing such programs to machines or situations where the context is different.
Suppose you want a robot you can give instructions like "clean the kitchen". Programming the robot to understand this as performing the action 'clean' on the object 'kitchen' is something we can already do. The problem is that the robot doesn't know how to perform that task. It doesn't even know what what state of affairs constitutes the desirable end result of a clean kitchen (as opposed to e.g. an empty kitchen because it threw out all your food and cutlery along with the trash). That knowledge is the meat of the problem, and it's just as hard in any language.
"Therefore, although changing to communication between machine and man conducted in the latter's native tongue would greatly increase the machine's burden, we have to challenge the assumption that this would simplify man's life."
I had clients once where both of these limitations prevented us from being able to produce a product for them. They where incredibly successful Wall Street types. They wanted a software system written that would automate some of what they did. They were traders, but not interested in high velocity trading. After a weeks of meetings about high level goals, we had a meeting to finally get down to specifics about the procedures and functionality that they wanted. They simply couldn't describe in precise language what they did! They had been doing it for years and years very successfully, but we couldn't help them because they couldn't describe what they did. It was very odd.
They worked with sophisticated mathematical models and strange rules of thumb, the morning news, and the perceived level of activity on the exchange floor. Some of them used fancy interactive graphs while others relied on a simple printout of a spreadsheet full of numbers.
The universe they worked in was very complex and involved decision making that, apparently, was hard to describe in words. Imaging a world class boxer trying to put into words the algorithm that he used to win a match. It was a bit like that.
A programmers wife asks him, "Please go to the store and buy a carton of milk, and if they have eggs, buy a dozen."
The programmer returns home, and his wife is very angry with him. "Why did you buy twelve cartons of milk?!"
English SE question detailing the linguistic mechanisms: http://english.stackexchange.com/questions/40234/bring-6-egg...
>> At some point I hope to have computer systems I can program by voice in English, as in "House? Could you wake me up at 7?"
> Yeah, well, I fear the answer will be yes (it could), but it won't do so since you haven't asked it to wake you up, only if it could.
From an exchange on python-list, https://mail.python.org/pipermail/python-list/2003-October/1...
I feel like the halting problem and Godel's theorem imply programmers will never be out of a job, and the best we can do is build out more and more solutions for specific use-cases, like Wordpress and Shopify. I don't think there will ever be a general-purpose AI that can write computer programs. At least not in my lifetime or my children's.
And btw, has anyone else felt that the Halting Problem and Godel's theorem are two sides of the same coin? Is there any formal connection between them? I feel like they are not "independent" (in the sense that Euclid's 5th postulate is independent).
The theorem of the unsolvability of the halting problem is used in modern proofs of Gödels incompletness teorem(s) [1]. Hofstadter writes about this relation in Gödel Escher Bach as well.
[1] http://www.amazon.com/Lectures-Logic-Set-Theory-Mathematical...
UTexas + "Foolishness" => Dijkstra.
Given this- natural language programming can hope to achieve reasonable levels of precision in very few languages. English, for example, has over a million words, while French and Italian are both under 100K. That 10x discrepancy in words means that there is necessarily far more imprecision in romance languages than English. Graamatical constructs can add precision and clarity but cannot make up for the inherent gap in vocabulary precision.
Someone with more knowledge can comment on the feasibility of precise NLP in non-English, non-romance languages.
If, however, acceptance testing is defined through the precision of a low-level programming language, but the problem is defined in a less-precise natural language, there will always be a precision mismatch between the languages used to define the problem. A solution may be implemented that addresses the natural language issue but cannot meet the constraints of the programming language.
I think of source code as being more analogous to architectural blueprints or formal logic than someone's verbal description of an object. It's just a coincidence that source code is governed by structures and concepts that exist in human languages.
Usually when people say they want to be able to create software with natural language, what they want are better tools that allow them to more efficiently create something, with less time spent worrying about syntax or technical minutiae. This outcome doesn't require natural language, it just requires better tools.
Why do proofs have commentary, and why do we have physics text books?
Python is nothing like English. It is an extremely precise formalism, and thank god for that.
Beyond this, natural language programming doesn't preclude explicitness. You could still have portions of limited explicit language or even portions of actual code if that is insufficient.
Natural language programming gets us two advantages:
1. Easy entry (presumably what the author cares about)
2. A language which is composed of an infinite number of DSLs - but without the pains of limited language scope and in which DSLs are actually easy to write. This actually fixes the problem author was taking about - using a language for a purpose it clearly wasn't meant to.
And if you think natural language programming is going to get rid of symbolism in language, I've got a U+1F309 to sell you.
"The black truck driver ran through the red light."
Is the truck black or is the truck driver black?
Is the driver driving, or is he running on foot?
Is the red light a traffic indicator, or is it a just a beam of red light?
Teaching a computer to figure out the intended interpretation here is a monumental task, and teaching the people to understand how their use of language is actually very easy to misunderstand is even harder.
In Western countries you'd assume it's a black driver. In a country where almost everyone is black you'd assume a black truck. So the code could have bugs depending on the location/culture (and thus context).
(Along with a user defined value that let's the computer take its best guess)
The goal is not simply to replace curly braces and other symbols with English words; unlike what COBOL would have us think, that road doesn't lead to natural language at all.
As someone who has interacted with Google, Microsoft and Apple's natural language systems I am glad this was ignored. Happy someone stayed hungry.
What Google, MS and Apple's NLP systems are doing is mapping your words to a selection of pre-prepared levers. It's simplistic in the extreme compared to the real problem of being able to create new levers.
IMO, in dismissing Dijkstra, you completely missed his point.
I definitely think the aspiration for natural language programming has only made things better and something we all look forward to
We often make this mistake with computer systems. I.e. the need to specify everything up front. Imagine a workflow where the process of design is a conversation with the system. To some extent, REPLs capture this, but they're still driven wholly by the programmers.
though, to be fair, eliminating useless jargon is helpful, but you've got to keep the useful jargon.