I'd hope there is a Unicode mapping into other characters - is there any language that supports this? Even for libraries would be nice.
I'd hope there is a Unicode mapping into other characters - is there any language that supports this? Even for libraries would be nice.
In schools, kids learn to code using Algol-like language with identifiers in Cyrillic (essentially, Russian) abbreviations, e.g. "нц для i от 1 до n" rather than "for i := 1 until n do", but even here you see Latin letter "i" ;) However, it's not uncommon to just go with e.g. BASIC or Pascal.
There is also de-facto standard accounting & ERP software producs from 1C which use Visual Basic-like language with Russian identifiers like "Константа" for "Const", etc. It's widespread in its domain, but not used anywhere else. 1C programmers are sometimes (half-jokingly) considered to be of a different caste. Sort of like COBOL programmers, maybe, or - probably closer - SAP developers.
I don't think anyone has any serious issues with basic English language for identifiers. Literature, documentation, discussion - here, language barriers do matter (a lot!), but not for programming languages themselves. That is, unless someone wants to show off as a true hardcore Russophile, of course.
Code comments, identifier names and e.g. log messages are different matter. I don't think there is any common practice for this - more like a matter of personal preferences. Some try to stick with English, some use Russian or transliterated Russian (mixing it with some English) liberally so it's not uncommon to see something like `log.Error("Дом не найден %s, %s: %s", ulica, nomerDoma, responseText) /* TODO: Добавить поиск лучшего совпадения */` in their code.
This gives me nightmares of dealing with an old patched Joomla install, where the developers had managed to bring in plugins etc. that used multiple inconsistent encodings in comments, so you could get all of the code to render correctly with any single terminal setting.
Been there, saw that. The quoted part alone is way more than enough to give nightmares. I bet, encoding issues is just the icing on the cake. ;)
The number of keywords in a modern programming language is not very large, and most are short, so even a beginner with no prior knowledge of English would not have very hard time using something like Python (or Java, or Pascal, or whatever they use at their school).
There was (is still?) a project to translate all the German comments in Open/Libre Office into English, a legacy of the German company that originally wrote it.
e.g.
// 教科リストを取得する。
String[] kyoukaList = getKyoukaList();
(Japanese university backend systems have two concepts of "a subject", with kyouka being a subject like math is a subject and kamoku being a subject like Algebra II is a subject.)
class Pessoa(models.Model):
idade = models.IntegerField()
nome = models.CharField()
Then some things wind up having mixed-language variable names: class PessoaForm(forms.ModelForm):Off course there are the accented characters but those are just transliterated to ASCII. Even when the language accepts unicode identifiers (like Python or Javascript) they are not transliterated to ASCII so "função" is not the same as "funcao". Everybody just sticks with ASCII for identifiers (occasionally people throw the Greek letter for pi or lambda in a formula but not on my watch).
The only thing more awful than these "Portunglish" codebases are Brazilian projects trying to stick with English names only. These are often full of misleading translations - for example using "Budget" instead of "Quote" because in Brazilian Portuguese the word "Orçamento" means both. Very confusing.
Pretty much nails it: https://temochka.com/blog/posts/2017/06/28/the-language-of-p...
https://docs.perl6.org/language/quoting
"Quoting" is all defined inside of a sub-language called "Q"
%<x>
%"x"
%{x}
%!x!
Unfortunately the rules are loose enough that even space is a valid quote character. So this too is equivalent to the above:
% x
("% x ")
A question mark followed by any other single character is treated as a string literal containing that character.
Perl 6's string quoting sub language has various feature flags that you can turn on or off.
Here are a few of the basic ones shown in valid code (comments and newlines included)
# start quoting with no features enabled
# (not even backslashing the delimiter is enabled)
Q
:scalar # enable $foo
:array # enable @foo[]
:hash # enable %foo{}
:closure # enable { 12/3 }
:function # enable &foo('bar')
:single # turns on backslashing the delimiter
:backslash # enable \n
:!exec # turn off executing it (redundant)
:exec(0) # another syntax for turning off a feature
{…} # various delimiters are allowed (generally punctuation)
There are shorter variants of each of the feature flags. :c => :closureI would like to point out that the parsing of the :closure and :function part of the sub language is reentrant. (Actually most of them add some form of reentrancy, it is just harder to show it for the others)
"a b &foo( "c d &bar( "baz" )" )"
Note that this reuses the base Perl 6 parser, and that is why it is reentrant.There are shortcuts for regularly used forms
「」 =:= Q[]
'' =:= q[] =:= Q:q[] =:= Q:single[]
"" =:= qq[] =:= Q:qq[] =:= Q:double[]
=:= Q:b:s:a:h:c:f[]
=:= Q:backslash:scalar:array:hash:closure:function[]
The ability to turn off features can be useful qq :!c [a {\n b\n} c] # { and } don't form a closure here
I would like to point out that the :foo syntax is used everywhere in the language for named parameters to routines and operators (most operators are implemented as subroutines) :foo =:= :foo( True ) =:= foo => True
:!bar =:= :bar( False ) =:= bar => False
:$baz =:= :baz( $baz ) =:= baz => $baz
If the delimiters are paired () <> {} 「」 «» “” ‘’ you can double up on them. q<FOO> =:= q<<FOO>> =:= q<<<<<FOO>>>>>
This is useful to avoid having to backslash a delimiter within the string, or trying to find a delimiter that isn't in the string. q<< <a> >> =:= q' <a> 'Personally I prefer it when they use English obviously (I don't speak German), but it also means you don't get ridiculously long identifiers.
Variables in English, strings in Chinese when they have to (it manages a list of city names). I could work on that code and I don't know Chinese.
This is more or less what I see in code here in Italy. Sometimes somebody slips a variable name in Italian but among pro developers it's limited at cases when there is really no obvious English counterpart. Example: legal terms, they tend to be very specific and you'd need a lawyer to translate them correctly. Accounting too.
Other languages, like Russian or Japanese, can be transliterated pretty unequivocally, so transliterated identifiers are often found in source code, especially in business-related identifiers where a precise translation may not even exist.
Does anyone know how governments in China, Japan or Korea handle this? Do they try to use non-english programming languages?
You can test this by changing your input method to Japanese or Chinese. At least in Japanese, the sound of the character is typed out using english letters (romanji) and various options are presented for autocompletion.
You are right that reading characters is probably faster, though.
Reading characters is a bit faster, but only because you had to memorize 10,000 of them already. People reading English will similarly chunk many of the words they are familiar with so that they are not so much read as they are recognized.
Technically speaking, Japanese have kana-based keyboard layout, where a single mora is a single keypress rather than two. However, there is still an extra (second) keystroke for dakuten/handakuten (か -> が), which is not required with romaji input (where it's just "ka" vs "ga"), so it's not a 100% speed increase. So, typing is somewhat more efficient than with romaji, but I think I've heard no one uses that, except for, possibly, professional typists.
In any case, even though mostly everyone prefers the IME on their computers, I haven’t seen any Japanese people who prefer romaji input on their smartphones.
Though, I wonder if the varying size of English words makes it easier to distinguish patterns just from the shape of the code.
Back in 2000s, text editors often did not default to UTF-8, so if you are from Taiwan for example, your file might be in big5 encoding, and would require a change if open by an editor in utf-8. Now this is a rare practice.
I am lead to believe this usage of English for code and Chinese for comments is quite common given the code I've read from public repositories and other companies we worked with.
On osx, capslock is used to quickly switch between languages.
The default way to switch is control + space.