Do we really need another programming language?
paulbutler.org
paulbutler.org
> ... until programming languages reach the
> expressiveness of written English, I’ll keep
> welcoming them.
English is so inherently ambiguous that you are really asking about legally expanded, lawyerese, and not English. And that's going to be several times longer than the source code.I agree there's space for improving languages, I believe Google's new "Go" probably isn't that big a step, but I think asking for the "expressiveness of English" is dangerous nonsense.
One point that I may not have properly gotten across is that I'm not suggesting that programming languages should approach English in form. I was just trying to establish English as a lower bound for the amount of information required to communicate an idea. (It's not even a lower bound, since English is fairly redundant, but it's an upper bound on the lower bound if you know what I mean)
If I said "until programming languages reach the information density of written English", do you think that would make more sense?
Real Information Theory, as opposed to Data Transmission Theory, has to take into account the receiver. The transmitter has to have a model of the receiver, and then devise the data that will transform the receiver according to the intent of the transmitter.
For example, if you're a mathematician and I talk about the topological space obtained by gluing the edge of a mobious strip to the edge of a disk being the same topological space obtained by identifying antipodal points of S^2, you'll know what I mean. If not, I have to explain further.
Similarly, a computer program can start with a top level description of your intent, and a competent programmer can then complete the steps.
But a computer is not a programmer. When you describe your algorithm in sufficient detail for a programmer, you haven't given enough information for someone who isn't a programmer. A non-programmer will look at what you've written and then say "How do I do that?" Thus you need further instructions.
So you might say "We sort this vector of numbers by picking one at random, scanning down the vector to separate into those that are smaller, equal and larger, then recurse." and a sufficiently skilled programmer who's never met the Quicksort can now implement it.
But there are details missing.
You programs should (perhaps) be arranged so that the main routine is sufficient for a very skilled programmer to complete. Each routine called then is sufficient for a slightly less skilled programmer to complete, and recurse.
The problem is that the terminating case is the computer, and not a programmer, or even a person. Thus the terminating case is a long way down.
Perhaps there is a language design that will bring the terminating case higher.
The nice thing about the terminating case being a computer, though, is that it's a moving target (Moore's law, etc.). At the extreme case, a compiler could model every molecule in a human brain and use the simulated intelligence to convert the description into a language that it can compile.
I would be crazy to expect that any time soon, but I think it shows that in theory we can do better than existing programming languages do.
By the way, I re-worded the last sentence to "as long as programming languages are less expressive than written English, I'll keep welcoming them.", because I don't think it's inevitable that programming languages can become as expressive as natural languages. I do think that as long as they aren't, we can do better.
The error rates we get on modern computers are scary for most "algorithmic" problems.
It seems that in general, for any two entities the level of the abstraction used for communication between them, should be the highest common level of abstraction that both of the two entities can operate in.
Take your example of the mathematician. If he is using an application that understands what a topological space or a mobius strip is, it seems then he can use those terms in the language that he uses to communicate to the computer.
Some computers understand what a Vector is. One can refer to vectors in the language used to communicate with that machine. For example a command like "multiply vector a and vector b" would seem reasonable. However if the machine is just a bare x86 box without any operating system, it will not "understand" what a vector is so one has to "talk" to it in terms of registers, memory and IO ports and so on.
This idea of course leads into DSLs. A machine can rise very high on the abstraction level scale today but the abstraction will be confined to a very narrow domain usually.
In comparison with a 3 year old as the subject (instead of a computer) the difference seems to be all in the ability of the recipient to absorb and apply contextual information.
Programming languages are formal specifications of instructions to a machine. They have to be incredibly deterministic/unambiguous because the machine does no inference of it's own. Even with a so called 'high level language' that handles a lot of under-the-hood stuff like memory allocation and garbage collection for you, there are no intuitive leaps. If you wrote the algorithm in English, but made sure to be rigidly unambiguous, you'd probably find the English version to be not so short.
To make the argument more convincing, don't use English. The grammar is ambiguous. Lojban, a created language, is not. It is being used as a spoken language (though it's not very popular yet). Still, it is being used very differently from a computer program in a fundamental way.
One of the major issues is the concept of object definitions. Object-oriented languages have a very specific meaning for what an "object" is, while spoken languages do not. Rather than creating a specification for what data and functions an object contains, spoken languages group objects by properties. Ambiguity aside, it's a problem of whether you're defining it top-down, or bottom-up. Prototypes are closer to the way we use language for object definitions, but it's still not quite the same.
Another issue is how the language is used. A random portion of the source code is completely non-sensical to a computer, while a random paragraph of a book can still be meaningful. The portion of the source code could make sense to a programmer, maybe enough to reconstruct enough of it, but no computer can process it.
So, to go back to the original point, programming languages and spoken languages are both supposed to convey ideas. Spoken languages do a much better job, hands down. OP used length as a rough measure, but the idea is valid. The question is how can we make computers understand spoken languages in a meaningful way. Answer that, and you'll have a better programming language.
Despite the progress that's been made, computing is still a very young field and there's a whole lot of room to grow. We need better languages, better compiler technology, better kernels, better paradigms for handling tasks (processes and threads don't cut it), and above all else, we need new ideas.
If people aren't experimenting with new ideas, we won't move forward. While most of the implementations (literal implementations or specific language designs) will fail, the good ideas will carry on and support the next generation. We should all welcome innovation in this space, even if we don't like what's being created.
For example, imagine describing a sorting algorithm to someone and you decide that it should be done in an imperative and sequential way (not functionally or without any parallelism involved). Describing it in English is too ambiguous and verbose, x86 assembly is probably too presice and too hardware specific, a turing machine program is too theoritical and too rigurous. So then you make something up, and I think eventually it looks similar to Python, Ruby or Pascal.
However, in practice the language itself is not worth much outside academia unless it comes with a solid library. At the end of the day, the user of the language will have to open network connections, render web pages, access databases and draw GUIs. If the language doesn't let them do that, no matter how elegant it is it will remain just a toy language.
As for libraries, this is becoming less and less of an issue. With .NET and the JVM, you have a single set of libraries that works with any compatible language and allows language designers to focus on what they're good at: the language. The effects of that can't be understated; it's allowing for a huge amount of innovation with fairly little time investment and ease of adoption.
Yes, the language does't have to come with its own libraries. It can piggy-back on some other libraries. If there was just one operating system. with a good and stable API the language might not need any libraries, just the ability to make system calls. So for .NET and java, Sun and Microsoft already did the hard work as far as library code is concerned so any language on those platforms is already miles ahead of a new lanuage without any "batteries".
Interestingly enough, if an API is stable and well designed, a language could be created to take a better advantage of it. I am currently looking at Vala, it basically started as a better language than C that would take advantage of glib.
So you have to become popular enough to get a large enough community to broadly implement lots of libraries. This means the language can't be too innovative. It has to be Blub++. A lot of the old baggage of the older languages will be carried forward due to cultural expectations.
Isn't this just how human beings work, though? Isn't this exactly what we should expect given how people work?
I tried this with a short Lisp program a few years ago and found it to be false. I may have been too detailed in my descriptions with the English version, but I don't think so. Evolved[0] languages for human communication are notoriously imprecise.
The "plain English" translation of a legal document will be much shorter than the original as well. Such language isn't used in legal documents because lawyers can and will argue over the interpretation. To make programming like everyday English, the runtime would have to sort out the ambiguities. I think that would require human-like AI.
Given human-like AI, much of what we do as programmers could be done by the AI instead. A non-programmer could simply ask the AI to create the program or provide certain kinds of output. Even in a situation like that, I suspect there would still be a demand for human programmers who could more precisely communicate their intentions to computers, if only for the purpose of creating better AIs[1].
[0]That means pretty much any language people actually speak. I'm not sure if it's true for designed languages like Esperanto.
[1]It might be a good idea to restrict AIs from certain kinds of self-modification or from creating new AIs. I think we've all seen that movie.
Different languages have different strengths and weaknesses. Some are particularly good at number crunching, others make working with text easy. Some provide rich, powerful type systems, others go for minimalism. Some languages are designed to enforce OO, others build on list/stack/tree data structures. Some go for functional purity, others are based on logic proving, others still are based on the flow of data instead of the application of functions. Some languages promote stacked hierarchial components, others promote flat hierarchies instead.
Until we find that sweet spot, where all these different features are combined in a powerful, expressive, yet simple way, we will need to continue experimenting with new programming languages. As we gain experience in writing different kinds of software and solving different kinds of problems, we can model programming languages around these tasks to simplify them and build abstractions upon them.
Personally, I don't think our search for the ultimate language will end until we can find a nice balance between imperative and dataflow programming languages, since they seem to be the two most fundamentally different paradigms, yet some problems are better expressed in one and others in the other. (Eg, imperative is very good at expressing purely mathematical concepts as well as imposing order on computation, while dataflow is great at highly concurrent processing - until we can merge the two, I believe we will always have problems)
So yes, we do need more programming languages. Hopefully some day someone will create a multi-paradigm langauge which manages to gracefully balance the various pros and cons. Unfortunately, I think it could be a long while yet, before an ultimate language is created.
So I guess my idea of the limit programming language/environment would be that it's automatic (or at least trivially easy) to make 90% of every new program into a gem, and then automatic/trivial to find the gem that does exactly what you need, as well as to debug it. (maybe rubygems are already this good, i haven't looked into them).
* No - We have languages that are turing-complete so new languages have no benefit.
* Yes - New languages can take trends that make their way into other languages in the form of libraries and built them directly into the language, making the source code cleaner.
* Yes - New hardware (cf. multi-processors) enables new features which can't be exploited by older languages.
* No - New languages segment the population, meaning we keep re-inventing the wheel by writing the same libraries over and over again.
A man who knows how to write a new programming language
But then doesn't!
Definitions applying to the female gender left to other posters.You need your engineering staff to use a language like Python instead of C or C++ for multiple reasons. C's too dangerous and C++ is too slow and full of traps. But Python doesn't compile efficiently into native code (forget Swallow for now).
My hunch is that Google will move their internal staff to coding in this language for newer projects, and not really care who else out there is using it.
Certainly.
Just not another 1970s abomination which firmly traps programmers into the compile-pray-debug cycle.