Non-English-based programming languages
en.wikipedia.org
en.wikipedia.org
I`m a native portuguese speaker, for example. In portuguese, we "break" the "to be" verb in two forms: The "ser" verb to denominate an immutable state like "The sun is hot" and a "estar" verb to denominate a transitory state "It is hot today".
My guess is that, if programming languages were made since the beginning with a language like that, the coding world would be a more pleasant place today. Things like differenciate mutable/immutable state would be much more natural.
So, it`s possible even to create a Non-English based programming language IN English. Russian, japanese, etc. languages have their peculiarities, that could shape coding very different than what it is today.
Ruby it`s a nice example of it. Just google about the subject, and found this material (http://blog.new-bamboo.co.uk/2010/12/17/learning-japanese-th...)
Any rubyist-japanese-speaker who would like to give some thoughts on the subject?
http://blog.new-bamboo.co.uk/2010/12/17/learning-japanese-th...
Empty chess file :: set comments (character * s, integer n)
{
If (n> = maximum number of comments)
For (; maximum number of annotations <= n; maximum number of comments + +)
Comment [Maximum number of annotations] = NONE;
If (s == NULL or the string length (s) == 0)
Returns;
If (annotation [n]! = NONE)
Delete Comment [n];
Comment [n] = new character [string length (s) +1];
String Copy (annotation [n], s);
}
¹ http://zh.wikipedia.org/wiki/%E4%B8%99%E6%AD%A3%E6%AD%A3² empty where one might expect void, character, for, string, &c.
I have a translation of "Design of the Unix Operating System" in Chinese, and take great comfort in the fact that I can still get the gist of all the source code listings--even despite the comments in Chinese.
I admire the simplicity of the grammar in Chinese (from the year I took of it in college), but honestly I find logographic languages are kind of gross.
EDIT:
Fine, fine, I admit it: a language with millenia of cruft is totally reasonable to use as the way of persisting the cruftiest programming language in the world.
Indeed, the same years of hard study that are required to write and read Chinese literately should be added onto the same years of study required to write and read C++ reliably.
This is such a comically bad and obtuse idea I think we should propose it as the next draft standard.
EDIT2:
Look, explain your downvotes (in the language of your choice!). My opinion is simply that alphabetic languages (here exemplified by English) are superior to logographic languages--mostly because they require knowing fewer characters.
I may be grossly misunderstanding Chinese here; as claimed, my schooling in it is limited.
Finally, with an appropriate method, it doesn't have to take years to learn enough characters to write and read literately, I can read Japanese at a high school level, and I've been at it for a year and a half.
As I understand it, there are three alphabets: kanji (lots of Chinese characters), katakana, and hiragana. The latter two are used to spell out the syllables of words, in some sense acting like an alphabetic language. Kanji seems to have a thousand or two characters in use, whereas katakana and hiragana have around fifty.
Beyond looking up stroke numbers and radicals, I found dictionary usage for Chinese characters somewhat hard--English lets you basically do a very easy binary search on a word (start at most significant character, find section, move to next most significant character, etc.).
You can do the same thing with dictionaries in Japanese or Chinese, though it works best if you use a dictionary that lets you handwrite in the characters (it helps a lot if you learn your radicals and stroke orders well, so that you can easily write characters you don't know).
"if you use a dictionary that lets you handwrite in the character"
I'm unfamiliar with any paper dictionary with that capability.
You end up with much lower population literacy rate than countries with alphabetic systems. ex) China at 92% vs Korea at 99%.
Hangul was created 600 years ago to counter the difficulties faced by ordinary citizens attempting to memorize the Chinese alphabets. Hangul also by far one of the more exotic alphabet system. This video sums it up nicely:
체스::리플달기 (캐릭터 * ㅅ, 숫자 ㄴ)
{
면 (ㄴ >= 최고리플번수)
용 (; 최고리플번수 <= ㄴ;)
면 (ㅅ == 무 아니면 문자 길이 (ㅅ) == 0)
도라오기;
면 (리플 [ㄴ]! = 없음)
리플 지우기 [ㄴ];
문자 복사 (리플 [ㄴ], ㅅ);
}
}
What's interesting about Korean alphabet is that you could read and write the code above within a few days of memorizing the Korean alphabet system. Even a non-native Korean speaker can read the above example (you could read about 70% of it as most words are phonetic spellings of english words like Chess = 체스. and write it without speaking a word of Korean if you knew the consonants and alphabets. There's very little Korean word in that example above.Good luck with writing it in Chinese (memorize all 4000 characters) or Japanese (Kanji, Katakana, Hirakana a clusterfck). If you're gonna build an Asian programming language, Korean alphabet's flexibility makes it easier if not more efficient to express developer intention.
Majority of the english words have been phonetically typed in Korean, there's very little semantic Korean meaning.
clavis hashus nominamentum da.
Is equivalent to [1]:
@keys = keys %hash;
And in Klingon Perl [2]:
De'pu'wI' bIH yInob!
Is equivalent to:
%data = @_;
[1] https://metacpan.org/module/DCONWAY/Lingua-Romana-Perligata-...
* Alphabet has 33+ letters, so there is no space left for programmer's favorite symbols on layout: @#$^&{}[]|~`<>. Yep.
* In Math variables' names traditionally are Latin/Greek letters, Cyrillic is used only when teaching kids. Again: switching layout all the time? No, thanks.
* Words are simply much longer. And I mean much. In English even short words tend to become shorter (variable -> var). In Russian there is simply no such pattern exists. `var-set` -> `установить-переменную`. `var-get` - `получить-переменную`.
* Words are stable in English morphologically. e.g. `last-msg-delivered` - `last-msgS-delivered` -> `последнЕЕ-сообщениЕ-доставленО` - `последнИЕ-сообщениЯ-доставленЫ` - if you change gender or from singular to plural etc you should change few words in a row.
Seriously, it seems to me like trying to make things more complicated.
One does wonder how soviets ever managed to program their computers.
Yep, that's one way. But it make things harder. If you set up Alt to use as English layout modifier to type {} you need to press two modifiers now: Alt + Shift. However I agree it's not the main issue.
> The English alphabet isn't the Latin/Greek alphabet either.
At least it's Latin.
> And I am absolutely astounded that Russians apparently have never invented the concept of an abbreviation or contraction.
Of course such things do exist.. they just don't work very well. English's short words is a relatively distinctive feature which it has developed being a mix of Latin/French/Germanic. For Russian words usually are complex. Var - переменная - prefix пере- + root мен. Yes, you can contract it to `перем.` But it also sounds like contraction of `перемещенная` which is `moved` or `перемещать` which is `move` or `переменный’ which is `variable(meaning: alternating)` and so forth.
Министерство здравоохранения => Минздрав (Ministry of Health => Minheal)
Министерство юстиции => Минюст (Ministry of Justice => Minjust)
etc, etc. As a foreigner, I found that quite peculiar though it doesn't seem to be as common now as it was before but it's still noticeable in everyday life.
Symbol; English; Russian
, ; , ; Shift + .
[ ; [ ; AltGr + х
{ ; Shift + [ ; AltGr + Shift + х
See pattern? Though I agree with you - it's not a dealbreaker.I could set up a lot of symbols though altgt but they won't mostly much English layout. But I don't need to because languages I program in are English based.
If I knew German, I would use exclusively German layout as apparently you do. But I need actually 4: En, Eo(that not much but still), Uk, Ru - those two every day. So I contracted it to 2: Latin/Cyrillic. :)
I`m a native portuguese speaker, and think that romantic languages can be more "emotional" than english. But its amazing how a new idea can be expressed so easilly in English.
German and Latin can create new words or ideas much easier than English can. English creates them 'easier' by wholesale importing them. Take the concept of 'karma', there is no word for this in English. In German this word is schicksal. This concept doesn't exist in English at all other than the Hindi import.
That said, your example is not a very good one. "Schicksal" can be translated as "fate", "fortune", "destiny" and "lot". Not just "karma".
Both languages have their expressive poets, authors, etc. It's easier to rhyme in Portuguese (it's almost like cheating, since the verbs all end in the same syllables), it's easier to modify the grammar classes of words in English (turning nouns into verbs and vice-versa), they both have rich vocabulary (although English seems to give more multiple meanings to individual words, and have more synonyms too). But in the end, I'd never call either of them more expressive than the other...
Portuguese, for example, is much more expressive in the sense that you can communicate the same idea in many different ways, which may have different meanings.
There is also the contractions issue and the omission of the subject of an expression ("eu estava" == "eu tava" == "tava" == "I was"), which complicates it even more. And all of this is actually valid, depending on the linguist you "follow". ;)
Well, I did not make myself very clear, but this is what I wanted mean when I talked about expressiveness.
PS: I wanted to write something about linguistic relativity too, but I can't remember, lol.
For example I don't even think about what it would mean in English if I write
while(true){ i++; if(i>100) break; }
I'm curious if native English speakers look at code as real English text sometimes? It should be funny, because when I translate my code to my native language(word-by-word) it's just meaningless and funny. if not callable(something): make_callable(something)
Which can (sortof) be read off in English.For things like "for", "wend" (while-end), "class", "switch", "main", they're divorced enough from any real English meaning that I just think of them as arbitrary coding words.
So, it's kind of both, at least for me.
Every statement has a meaning, and every block is a story. If it wasn't readable we could just as well use BANCstar [1].
(Btw, I absolutely hate that Microsoft translates Excel formulas (and so on) in localized versions. I simply cannot program in my native language: I basically translate back to English mentally.)
Although I suspect that if this did get written and translated around a lot, many non-English speakers would still pick the English surface syntax due to network effects...
(Edited for clarity)
* Formatting isn't preserved and the file is retabbed each time you save it.
* Earlier versions had French and Japanese compilers/decompilers
* Each application can add AppleScript syntax (like VBA/OLE or whatever). If a script uses commands for an application you don't have installed, you see the raw FourCC codes in the code instead.
It's actually a quite interesting language. http://en.wikipedia.org/wiki/AppleScript
It also uses the same file type for tokenized and untokenized programs: untokenized programs are tokenized when run, and tokenized programs are untokenized when opened with the built-in text editor. Short of opening the file with some external program like a hex editor, there's no way to tell whether the program is tokenized or not. Which, of course, can cause problems if you send an untokenized program to a calculator set to a different language.
Frankly, as a non-native speaker, I absolutely dread localized source.
The turtle academy is a programming environment with lessons for teaching kids logo.
For professional programming, a non-english programming language is just a curiosity, non of them has ever gained global traction nor will probably gain.
For kids, for learning, supporting multiple languages is a must.
English: http://turtleacademy.com/lang/en
Russian: http://turtleacademy.com/lang/ru
Hebrew : http://turtleacademy.com/lang/he
Spanish: http://turtleacademy.com/lang/es
Chinese: http://turtleacademy.com/lang/zh
And we are looking for volunteers to translate to further languages.
http://speedata.github.io/publisher/manual/index.html - Switch to the other language in the footer
I don't know if it is an urban legend: I read that in English the length of a words is shorter than in Japanese; so in WW2 the English language was better for shouting out orders - less time for communication means more time for action;
But the summum is reached with perligata.
Some ideas (warning - long and useless read):
- everything is an expression
- expressions have Cases. Default case (without postfixes) is Source Case
- cases are like roles for parts of expression in given complex expression context (allow us to change order of arguments in expressions without named parameters for example, and to overload functions/macros depending on Cases of the supplied arguments)
- you can define pre-, -post, and -in- fixes that work like macros/functions
- you can use pre/post fixes to tag expression with cases in the context of an expression
- you can use 1-arg functions/macros as pre and postfixes, or as regular functions
- you can use 2-arg functions/macros as infixes or as regular functions
- more than 2arg functions can only be used as regular functions
- you can overload functions/macros depending on case of args
- you can define your own cases
Example:
% %1 %2 etc are unamed arguments in
macro/function definitions,
like in lambdas in clojure
"-" is used in pre/in/post fixed macros
definitions to mark where the base
identifier goes
%-! is 1arg macro that defines variable of name %.
Example: a!
foobar!
%-a is postfix macro that tags variable %
with Target case
%-i tags with Helper case
%-p tags with PositiveConditional case
%-n tags with NegativeConditional case
%1a=%2 is infix 2arg macro that assigns to
variable %1 (must be in Target case)
value %2 (in Source case).
Can be used as regular macro: = %1a %2
or equivalently
= %2 %1a
We can also overload = to be equality operator
if both args are in Source case %1=%2
works with regular macro too:
= %1 %2
or
= %2 %1
We can overload binary operators with 3rd and 4th args in PositiveConditional and NegativeConditional cases and we don't need IFs :)This also allows us to write if with only else clause in the form:
> x 0 (OUTPUTa="X SMALLER THAN 0!!!")n
%1-;-%2 is infix macro that evaluates in sequence.
It's overloaded for all cases, especially
for %1a-;-%2a
# is python-style comment
Example code to calculate real roots of quadratic equation:
a! b! c! delta! # definitions of variables
= INPUT aa ba ca # overloaded = macro with
# many targets (a b c in Target case)
# and one source
# to easily fetch many items from input
deltaa=((b^2)-4*a*c) # precedence will be a problem
# to define generally
= delta 0 # overloaded = equality test with
# 3rd arg being positive conditional case
# and 4th being negative conditional case
(
= OUTPUTa "x0=" (-b/(2*a)) # overloaded = macro with one target
# and many sources - will join the sources
)p
(> delta 0
(
= OUTPUTa "x0=" (-b-sqrt(delta))/(2*a);
= OUTPUTa "x1=" (-b+sqrt(delta))/(2*a)
)p
)nI know a bit of Finnish, and I can imagine that a programming language built on the principles of Finnish grammar would be beautifully expressive as well as extremely terse.
For example, maybe '-ss' denotes 'class'. Functions and methods with '-ef' with an optional return type '-int', '-ing' (string), '-oat' (float). Start a loop iterator by adding '-oop' to an iterable? Modify conditionals change their behaviour, e.g. 'ifret' (if/return) etc
Fooss:
barefint:
thingoop(i)
ifret i == someValue
Looks strange, but I if you're used to the concept I wonder if it would be more efficient.• You want to learn how languages work.
• You have a problem which is difficult to solve in other languages but easy in yours.
• You want to combine features that have never been combined before, to see if their interaction creates something of value.