2. Because Windows has case-insensitive filenames, and Linux, etc., have case sensitive filenames, we recommend that path/filenames be in lower case for portability. This annoys some people.
3. There are command line switches to map from module names to filenames for special purposes. They're very rarely needed, but invaluable when they are.
Overall, it's been very successful.
I'd say "valid X language identifier characters" should always be ASCII.
I never understood the BS fad for unicode identifiers.
Wanna allow some math symbols? Maybe. The full unicode gamut, so that you can have a variable named shit emoji? Yeah, no.
Or why not full unicode text support? Is there any real reason besides "some people might want emoji and I don't like that"?
But no, the totality of the argument always reduces to: "I'm not used to this and would find it inconvenient"
This means the code can be read by anyone anywhere in the world on any operating system and that string payloads can similarly be read by anyone anywhere in the world.
Uh, no? You are not supposed to be able to read this valid Unicode string literal in Korean: `"그뤼고 이 문좌열은 일부려 기ㅖ버역을 어럽게하러고 오타비문이 산개해 있구먼요."`
Also a significant portion (and possibly the majority) of codes would be ever read and written by a small group of people, often sharing a common language other than English, so non-English code is just fine for them. If you are saying that a public library should be written in English, I almost agree---there would be some exceptions though.
public class 그뤼고
{
private 이 이 {get;set;}
public 그뤼고()
{
var 문좌열은 = 이.문좌열은;
var 기ㅖ버역을 = Enumerable.Range(0, 10);
var 오타비문이 = 기ㅖ버역을.Select(요 => 요);
}
}
public class 이
{
public int 문좌열은 {get;set;}
}I have seen numerous instances of pseudo-English when it comes to naming. It is hard to name things in non-native tongues. When reasonable, reducing that overhead can be indeed beneficial.
Plus it is pain to alt+shift between languages all the time.
But in cases where the program will be dealing with some concept that doesn't exist in English, being able to refer to things by their actual name, in the native language (assuming that's also the native language of the customer and development team), is much better than inventing a confusing and unnatural English translation.
A programming language's identifiers is not the place to express one's national identity. They should be utilitarian, and easily understood by programmers across countries.
Since you're already supposed to understand the syntax of every major programming language (which is based on english) you can make do with english keywords too. Nothing worse than opening some code to find bizarro foreign language identifiers.
(And I'm no native english speaker, so I'm not speaking as someone who's ok with this because english is their language or ASCII fits their default keyboard layout: it just makes sense).
> Nothing worse than opening some code to find bizarro foreign language identifiers.
I'm not a native English speaker, and I have only to loose in this scenario.
Just like people know they have to write in English on HN to communicate, they also write their code in English when they plan to open-source it and share with the rest of the world. As for closed-source projects ... if your company doesn't conduct its business in English, why force the code to be in English? The only people whom that'd benefit are never going to see it.
We live in a multicultural world, one for which English as a default doesn't make sense in every context. Yes, it may be the case that Chinese, Spanish, Greek, German, Arabic, or other non-American ideogram using programmers write code primarily meant to be used and understood within their own culture. I see nothing wrong with that.
Also what about the fact that while Mandarin is the largest language/dialect it's far from the only one in China? Using English/ASCII means all the programmers in China can understand each others code...
I mean, yeah, this can be made a requirement of programming languages: After all, it was such a requirement for a long, long time. But it doesn't have to be anymore.
BTW, full disclosure, I'm French and I find code written in french completely fucking unreadable. And that's 100% ASCII. As I said I believe code should be written in English, but I also don't think we should have essentially-artificial barriers for people to enter something as important as programming; those barriers only end up eroding the culture in question.
Compare итератор and 迭代器, which are complete mysteries to me produced by Google Translate. If my intention were to reach as many people as possible, I'd use "iterator" (which, coincidentally, works for English and my native German).
If once worked in finance and there is a difference between GAAP accounting and German accounting rules. If my algorithms used English terminology to be consistent with technical terms this would be confusing inneach review. Using German terms (even combined with English "get" or "set", like "getBetriebsertrag") there was beneficial, even though it always confussed new members of the team.
(Indeed, we should arguably move away from the notion of a single character string as the only human-facing semantics that an identifier is associated with-- there should be a higher layer, perhaps with multiple choices of e.g. native language, formatting and the like. Human facing semantics are closer to "literate" documentation than to anything that compilers should have to deal with. Yes, the "native", underlying representation should still be something that we can somehow make sense of - I'm not saying that our identifiers should be GUIDs or anything like that! But it will only be resorted to in a pinch.)
I don't see a dichotomy (much less a false one). What are the two options I separate artificially?
I'm saying just don't impose regional alphabets (other than AZaz that's already par for the course with the syntax of all major programming languages anyway) and regional words into source code.
That's also a problem in most other languages; just a couple of days ago, someone else's C++ code didn't compile on my machine because they had accidentally included <something/whatever.h> when the file was actually named <something/Whatever.h>, because macOS is case insensitive. I had the same experience with JavaScript some months ago, that time because they were running Windows.
On the "filenames must be valid identifiers" thing; I really wish more languages would start allowing kebab-case in identifiers. That's also absolutely not a D thing, more of a common complaint about most languages.
Exactly which operators should require whitespace and which don't is up for debate, but in my personal opinion, requiring space around infix operators and letting prefix/postfix operators not require a space would be appropriate. Nobody wants to have to write `myArray [i]`, but I think most people would be willing to give up `i-1` and instead write `i - 1`.
t[x] = t[x-1] + t[x-2]
looks more readable to me than this : t[x] = t[x - 1] + t[x - 2]
Another example : y = a*x1 + b*x2 + cWould adding whitespace sensitivity really be a problem? You already need whitespace to separate identifiers, so it's not a totally foreign concept in mainstream languages.
It seems like we've been making a weird trdeoff, by disallowing kebab-case just so we can smash our operators together with our operands.
Some people don't want to bother putting a space between operators and operands, and proponents of kebab-case just don't want to push the shift key to get an underscore.
I agree that nobody would want to write `i ++` or `foo [10]` or `myvar . mymember`, but I think a lot of people could get behind `10 - 20` and `foo && bar` instead of `10-20` and `foo&&bar`.
Also, the reason to prefer kabab-case for me has nothing to do with avoiding a keypress. It's that I find kebab-case easier to read.
x = (-b + sqrt(b**2 + 4*a*c) / (2*a)
x = (- b + sqrt(b ** 2 + 4 * a * c) / (2 * a)
... Although, one could argue that allowing tightened multiplication and division are enough. x = (-b + sqrt(b**2 - 4*a*c) / (2*a)Aside from C-style type declarations ("unsigned int x;"), C-style syntaxes seem to always have ways other than whitespace to separate identifiers.
Like (using some JavaScript in a hypothetical example) I can't think of many concrete reasons why this is easier to parse:
let first_number=2, second_number=2, answer=first_number-second_number;
...than this: let first number=2, second number=2, answer=first number-second number;
Although, of course, some languages—most Lisps, Tcl, and Red/REBOL come to mind—actually do rely on whitespace and whitespace alone to separate identifiers in many situations, and something like this would likely be unworkable there. let let x = 5;
let x = 6;
// should this set the variable "let x"?
// or define a variable named "x"?
One could potentially design around situations like this, but allowing whitespace in identifiers likely does require being much more meticulous about treatment of reserved words, identifiers, and whitespace than more traditional syntaxes, and this is likely why not many people attempt this.I think the idea is worth experimenting with, though, and that a good implementation of it could be convenient enough for end-users to outweigh the implementation inconvenience.
CL-USER 115 > (let ((first| |number 10)
(second\ number 20))
(+ first\ number |SECOND NUMBER|))
30I think the best way to get identifiers with whitespace to work in a Lisp would be contrive a syntax for S-expressions that uses something other than whitespace to separate things. Perhaps letting (first rest-1 rest-2 ...) be written as as (first: rest 1, rest 2, ...) or (first, rest 1, rest 2, ...), so that example could be written as:
(let: ((first number: 10),
(second number: 20)),
(+: first number, second number))
I imagine it would be possible to write a macro in Common Lisp to transform this into runnable code, or a language in Racket to do so—although, I'm not sure how many people would actually want to make or use something like this.TXR Lisp:
1> (list 1"a"'(b(c)d(e)))
(1 "a" (b (c) d (e)))
Here we just have one space that prevents list 1 from being list1. foo-bar # kebab-case
foo−bar # subtraction
foo minus bar # subtraction? (infix identifier "minus")
foo - bar # subtraction (infix identifier "-")
foo − bar # subtraction (operator symbol)
using \u2212 as a explicit subtraction operator for people who really can't stand having 'extra' whitespace?Small Intro: https://perl6advent.wordpress.com/2015/12/05/day-5-identifie...
this could be solved by allowing strings in qualified imports
import "illegal identifier"."some more weird unicode" as someLib; const thing = @import("relative/path/to/thing.zig");
const package = @import("packagename");`extern crate foo`
`extern crate "foo-bar" as foo_bar`
This is all legal identifiers but
`use std::path::Path;`
`use std::path::Path as int; // For maximum confusion`
And not that anyone uses this part but:
`mod bazz;`
`#[path = "bazz-bar.rs"] mod bazz_bar;`
import A -- compiler reads A.hs
import A.B -- compiler reads A/B.hs
Thus, two semantically related modules are now in different directories.Python, IMO, handles this correctly by having __init__.py support inside directories. It's theoretically less elegant because of the special name, but in practice leads to better file organization.
Same for Rust, but even better, because one can define nested modules in the same file. So you can either define a new module in the same file, put it in a different file named by the module, or put it in the file `mod.rs` inside the directory named by the module.
> import A -- compiler reads A.hs
> import A.B -- compiler reads A/B.hs
That has to happen at some level, assuming subdirectories; otherwise, what would "import A.B.C" refer to?