Programming languages have to break free from the tyranny of ASCII
queue.acm.org
queue.acm.org
How do I type it?
How do I type it?
That's the fundamental question worth repeating three times. Numerous languages support Unicode in the syntax and it doesn't matter because nobody can use them because who wants to use a language where the 'or' operator is a textual copy and paste operation for a new user?
You can incrementally stuff more symbols onto a keyboard but I'd want to see some science that shows that people are able and willing to deal with a keyboard that contains a their local language, a full suite of mathematical symbols including Greek, Hebrew, and the various mangled Latin characters that mathematics uses, plus logic, set notation, plus the useful not-necessarily-mathematical operators that language designers either want or in some cases already use/permit such as arrows, boxes, and all the other things. And we still need our full editor keyboard shortcuts and if you're a power user (and we're talking programmers...) our keyboard shortcuts for the window manager.
So, you want to predicate learning your language on me learning all that? I'm the kind of fruitcake that learned Haskell but I still laugh at the idea of learning something like that. No sale. (And I don't use the Unicode Haskell permits because while I'm actually very comfortable remapping my own keyboard, for instance I actually have interrobang bound to a key because I found myself using it so much, why would I do that to anybody else who ever tries to read or use my code‽)
I did that on a Mac using the option key and a few custom key bindings.
The real question is how do you READ it. The problem is that unless you are very selective about your fonts and which characters you use, it can become impossible to read because many unicode characters are visually indistinguishable.
The exceptions to this policy are usually writing systems that didn't have standardized encodings at the time they were incorporated into Unicode. Also, the rather notorious CJK unification project resulted in several new overlapping layouts as Unicode policy changed over time with respect to Chinese characters.
Added - they did; here is an older version (3.0) of the standard online http://www.unicode.org/book/u2.html
But that's not the actual point. Learning a new language is a challenge. Learning a new language when you haven't already learned 10 is an even bigger challenge. Learning your first language is very difficult, which most of us have probably forgotten. Now, to the list of things you have to learn to do so, we want to add learning several dozen brand new keystrokes that nobody knows offhand? It may not seem like much compared to everything else one has to learn to program, but having to learn a lot of stuff is annoying; having to fight your input devices is like being rendered mute. Try switching your keyboard map for a few days to see how you like it. This is not a level of frustration worth pursuing.
Capslock::F15I'm prepared to be corrected but ..
Greek capital sigma is the mathematical "sum" symbol just as capital pi is the product symbol.
Another reason this makes no sense is, that programming languages are actually little about syntax and much about paradigms (proc, func, oop, etc.) with technical details, like garbage collection and runtime features like interpretation. That's the important part, not the character set.
Yes, I'm one of the geeks that use Dvorak, vim and just got me a mechanical keyboard with tactile feedback. I care a lot about my typing, and I sure as hell don't want to type in unicode.
Really, I actually live in Germany, and have the extra 3 umlauts on my secondary layout, and that's already a royal pain in the ass. I also did some CJK text input (Japanese is a hobby of mine). well, don't get me started.
No, more than the ASCII character set is defiantly the very last thing you might want, eclipsed only by syntax highlighted critical path display. He could at least have taken the time to imagine what the result of that would have been, or make a mockup, then he would not post this crap!
http://projectfortress.sun.com/Projects/Community/wiki/MathS... http://projectfortress.sun.com/Projects/Community/wiki/Fortr...
Yes, if only developers could devise some way of remapping keyboard keys, affixing new labels or even creating entirely new alternative layouts and devices. But that sort of rocket science is at least, what, three decades out?
Second, what "productivity benefits"? Are you such an astonishing programmer than your primary bottleneck is typing your symbols? It's too slow to type -> instead of Alt-Gr->, or whatever? Or maybe the agony of having to read -> instead of →? I find it hard to believe that's much more than a rounding error.
I guess that mathematical notation is often easier to read and comprehend. I'd rather see something like this http://upload.wikimedia.org/math/9/6/9/9695855b7c5869aad99fa... than something like
sum [ for x in s -> f(x) ]
on the other hand, without an IDE that's capable of translating between both notations, editing would be rather painfulPractically, having to regularly deal with Japanese encoding, a big issue is interoperability. Granted, a lot of issues are caused by not using unicode (unix vs windows encodings which sadly are still pervasive), and those could be mitigated by enforcing unicode.
I am not so much concerned about readability: mathematics go away with tons and tons of symbols, and it is not really ambiguous - certainly less so than most programming language syntaxes I know.
Meaning?
But that's not very interesting, because that's not really how we program, at least not anymore: you write something to be read by other people more than something to be understood by computer, and two equivalent programs are often not equivalent at all in readability. So maybe my choice of words was not optimal, and I should have chosen readability instead of ambiguity.
I think the focus on being executable by the computer is the wrong one. That's the trivial part.
Soft keyboards and composing-entry are the future.
Right now, I'm learning APL using it; it's pretty comfortable for that. An important point, though, is that APL has a finite number of extra keys. Every symbol you insert in a source file is a symbol other developers have to have a way of typing—so, while this system can work for APL or arbitrary-symbol-lang-with-a-keyboard-conf-file-in-the-installer, it won't let you just define whatever symbols you please on a file-by-file basis.[1]
[1] ...for now. I'm thinking about writing an IDE hook that scans your project for symbols your keyboard can't type given your board-set, then dumps them into a temporary "project keyboard palette", that you get to by swiping to the left/right. This would let you go beyond Unicode to use arbitrary pictorial symbols (i.e. SVG images) as characters. If you could drag keys off your palette and onto your personal boards, keyboard symbology would then become a memetic process, spreading from project to project when the symbol is useful to the reader, much like words spread in conversation.
Let's say you want to express hints to the compiler about what paths are common. And you like the idea of expressing that with tinted blocks. Here's what you have to do to make that work:
- rejigger your editor to input and save such tint blocks in the code;
- convince everyone else using your language to use your editor, or make similar modifications to their editor;
- ignore the cases where you don't have a particular editor available, such as SSH'ing into a minimally configured machine;
- exclude color-blind and blind programmers;
- abandon the convenience of being able to view your source code in pagers like less, diff tools, revision control tools, web browsers, IRC, instant messaging, and web services that allow you to paste snippets of code.
Or, you could:
- add a new keyword to your language.
- add a new syntax coloring rule / code folding rule to your editor, and get almost exactly the same effect.
We have a basic information science problem here. Every feature of the programming language is going to have some concrete representation as a stream of bytes anyway. As long as we use multiple tools and environments to manipulate and view that stream of bytes, it makes sense for the representation and the serialization to be identical -- that is: plain text.
However, there's no reason why we can't have the odd unicode character in source code, as long as there were some well-known shortcuts for typing. A lambda symbol would be wonderful for many languages. Similarly, a Chinese programmer ought to be able to define variable names and strings directly in Chinese.
Not only that, but we actually debated using the unicode characters for !=, >=, and so on instead of the standard multi-character ASCII equivalents. In the end it was decided that it would be a gratuitous change for little benefit.
So it's a shame that he levelled his criticism at Rob Pike, a guy who clearly knows a lot about unicode (he co-authored UTF-8 with Ken), built the first Operating System with unicode support in its foundations (Plan 9), and has taken great care to support unicode in Go.
I've never understood the arguments against operator overloading. Yeah, you can create unreadable programs that way. But it's equally easy to create a function named "sort" that erases your hard drive. If you apply the same disciple to naming your operators that you do to naming your functions, it's not hard to understand. And it's a huge boon if you happen to be (for instance) defining a new numeric type.
Big_int.( 23 567 mod 45 + "123456789123456789123456789")
where only the operators in the parentheses are overloaded.
Moving to this sort of setup will result in a) conflicting representations, b) significantly slower code, or c) epic frameworks on the order of all-of-CPAN with tens of thousands of things you must know to be even remotely efficient.
Barring a Sufficiently Advanced Compiler, of course, but until one exists (and puts all of us out of a job) optimization is still frequently important.
The main hurdle is actually a UI problem: How do you make a structure editor with the ease of use of a text editor?
Chuck Moore, of course, has done this with Colorforth. I read him somewhere saying that by encoding some of the program in color, a visual part of the brain not normally involved in programming was engaged, freeing up other cognitive resources to think about the problem at hand. He didn't put it that dogmatically; it was more of a speculation about he found this use of color so helpful.
Of course, the solution to this would be to instead provide a UI (presumably in your editor) to change which colors are associated with which syntax bits... which is exactly what we have in any reasonable editor. So... what's the problem?
Or take a language like Scala, which I happen to like a lot, but has one hellaciously complex type system: how much would it benefit from an additional syntax abstraction? This question is worth asking!
I like the columns of code idea a lot, and it's something I've thought about for a long time. But my conclusion so far is that this is actually better done with IDE functionality. Just establish a common annotation for metadata which IDEs interpret as a display configuration. Or have the IDE open a .src file and a .doc file in the same directory, displayed side by side. The doc file has the same line count as the .src file, and there are conventions / macros / etc. that make use of the second column, etc. (Of course the .doc could be compressed into some kind of metadata file when stored.)
To bring it back to the topic, aren't we all REALLY interested in something that makes coding... just, better?
$ perl -e 'use utf8; $水=1; print $水,"\n"'
1http://search.cpan.org/~ewilhelm/lambda-v0.0.1/lib/lambda.pm
It's sort of telling that the manual page has a lengthy explanation about what to do if you can't see the symbol, and a vim macro that you need to install so you can type the symbol. And it's just one symbol saving 2 characters (and no keypresses) :-(
Just as you mentioned of perl, very few people use it.
Python code is written by many people in the world who are not familiar with the English language, or even well-acquainted with the Latin writing system. Such developers often desire to define classes and functions with names in their native languages, rather than having to come up with an (often incorrect) English translation of the concept they want to name. By using identifiers in their native language, code clarity and maintainability of the code among speakers of that language improves. </quote> http://www.python.org/dev/peps/pep-3131/
$ python3 -c '水=1; print(水)'
1
But I also don't know anyone who is actually using them.If the syntax is changed to allow '打水' instead of 'print(水)', where the language parser knows the single CJK token 打 has a standalone semantic, then needless spaces can be removed.
I used to mock LabView (arguing speed and complexity), but now that I work with a colleague that exclusively uses it--we are a medical research engineering group--I really start to respect its utility.
So the question becomes: Would an opensource cross platform scripting language, modeled after LabView, fit this bill?
However, they tend to fall down for many general programming paradigms. LabView is probably the furthest along that I know of in attacking more general problems, but ultimately they are designed around building a pipeline: datain->do something->data out.
LabView also builds in interesting levels of instrumentation and interactive components that can make it seem more general purpose like.
In that sense, the old UNIX tools (simple, single function tools you can pipe together to create complex systems) are the same thing.
But there are other graphical programming environments that aren't focused around pure dataflow.
Is actually a very interesting experiment in using graphical components to build up a program. But ultimately it's just a nice wrapper around a regular old programming language, and in playing with it a bit, I found that it was simply faster to just type things in a regular old language.
One of the more interesting things about graphical programming systems, is that you can often coax people into learning to use them who otherwise would have no interest or inclination towards writing code. I think there's probably a huge untapped market for a fully graphical system that can compile and perform like native code...
Isn't that what computing is all about?
while not (user clicked 'close') do
something
done1 - Display something (perhaps put some data out) 2 - Receive input 3 - do something (perhaps combine input with data in) 4 - goto 1
This is true of everything from video games to operating systems.
Dataflow models are simply a subset of the models of computation we deal with these days...they are extraordinarily useful models, but still limited in what they do.
So far I've been wondering if a graphical interface a la labview could work if it translated down into say, erlang code. It all depends highly on who the target audience is...a test engineer, or someone wanting an easier way to write scalable distributed applications. If the latter, I'd expect the interface to be a very well thought out interface that lets you effortlessly flow between drilling down into functions and seeing the greater view of all the code in your application.
In the following program, what public methods exist?
package main
import "fmt"
func ॐ(ॐ string) {
fmt.Printf("%s\n", ॐ)
}
func main() {
ॐ("x")
}As it is, parsing is ugly because we'll write 1D text visualized as a 2D projection; then we use parse rules to blow it out into a tree so that your (parens and {braces}) can work properly; and then (if writing a compiler rather than an AST-walking interpreter) we collapse it again into a 1D executable format after one or more rounds of semantic analysis. It's a lot of steps, and there's plenty of room for error. It offends my minimalistic sensibilities.
What would the alternative be, though? It's nice to be able to just type out code, and to copy+paste the data from other formats like HTML documents. So in the end you still want to have an editor that can deal with plain text, even if it is "natively" aware of the AST rules of the language.
My other alternative is to write more assembler or Forth-style languages that don't have to be parsed as a tree. Not what most people would want...
Because it is simple and everyone can read it. I don't want to see code that is "greek to me."
I also think it's a terrible idea to use color. Not only are a lot of people color blind, but what happens if you need to print your source code on a black and white laser printer? Those are still heavily used.
I don't think the author of the article thought these issues through very carefully.
Isn't this obvious? What keyboard has 'Ω' or '₀' easily accessible? I'd be pretty upset refactoring or reusing any code that forced me to copy-pasta or type some crazed key combination just to render a character. It might not look perfect, but the lowest common denominator (ASCII char set) for programming works because every keyboard layout on the planet can enter those characters in (mostly) one keystroke.
Am I missing the joke?
I think Go actually gets this right: its syntax doesn't require any wacky characters, but it allows identifiers to consist of any Unicode characters classified as letters or numbers. The syntax caters to the lowest common denominator, but if I really want to name my variables in Greek, I can.
Footnote from John McCarthy's Recursive Functions of Symbolic Expressions and Their Computation by Machine, Part I
For example, I work now on a Java code base which was originally developed on French keyboards, with accents and other diacritical marks typical of the French language in comments and sometimes in constant values.
This code base, developed on Eclipse on Windows boxes which defaulted to CP-1252 encoding, was kept in CVS on Linux boxes, then migrated to Subversion on Linux boxes, and now is edited on Eclipse on Linux boxes which default to ISO-8859-1 encoding (when the developer has not fat-fingered the Eclipse configuration). None of the original diacritical marks has been preserved and the comments in particular are a mess. You can however choose to make an effort to interpret the meaning of the comment and to replace any funny looking sign with a similar-looking ASCII character.
None of this would have happened if the editor had accepted only ASCII characters: it would have been simpler and because of that much more robust.
#!/usr/pkg/bin/tclsh8.6
proc გამარჯობა {} {
# Georgian "hello", via Emacs (C-h h).
puts "Oh hi!"
};# call above proc
გამარჯობა
;# I can eat glass, Georgian, via http://www.columbia.edu/kermit/utf8.html
set მ "მინას ვჭამ და არა მტკივა."
puts ${მ}
lset newrhs $position ${sym}\{Ø}
which lets one's code read like the printed algorithms in a book.
Mr. Kamp slags syntax designers as putting compatibility with the ASR-33 above expressiveness, but I dont' see a way around it with today's standard hardware.
Who's gonna do that? And how is it going to compete against C++, Java, ObjC and the rest out there? When you come up with an idea, that radically changes the programming landscape, you need to have a vector to introduce it to the world.
The only way I can see that happen is, when someone like Apple declares this its new iPhone language, or Microsoft's new .NET language, but yeah .. the odds for that happening.. exactly.
http://hackage.haskell.org/trac/haskell-prime/wiki/UnicodeIn...
The first program to be converted to UTF was the C Compiler. There are two levels of conversion. On the syntactic level, input to the C compiler is UTF; on the semantic level, the C language needs to define how compiled programs manipulate the UTF set.
http://plan9.bell-labs.com/sys/doc/utf.html
Plan9, what Unix did next.
Ω is a valid variable name in C#. (But Ω₀ isn't).