https://github.com/julialang/julia/issues/558#issuecomment-4...
In general, much of what's chosen for language design is influenced to some degree by convention. This is a valid design choice, because you're maximizing use of existing domain knowledge of your audience.
In this specific case, I would argue that since in many languages where an array isn't a small shim over memory management (Ruby, Python, Perl, JS, etc) where easing user friendliness is at a premium over performance, 1 based array indexing would make better sense... if most people adopting those languages weren't already fairly well acquainted with 0 based indexing from CS programs or exposure to other languages (and it wasn't a simple concept to understand). Since many new users are familiar with it, and it is easy to understand, it's simple to just go with what some percentage of people already know, so convention wins out.
http://wiki.c2.com/?WhyNumberingShouldStartAtZero
What I find most persuasive is that 0 based indexing brings measurement into accord with enumeration.
Being able to easily port code between languages (and math articles) is important, when there's no clear winner, I like consistency.
In mathematics 1...N is a perfectly good notation. Julia expresses the same idea with 1:N. This is consistent with spoken language (1 denotes the 1st element). Consistency, if anything, would not come down in favor of the software engineering choice here.
Then I realized that 90% of my mental gymnastics with indices got simplified. Explaining indexing to newbies is now immediate where before it took a slide, several examples, and a clever picture. Slicing is also more natural:
"1:3 picks out index 1, 2 and 3, which are the first, second and third element of the array."
Instead of the python version of
"0:4 picks out index 0, 1 and 2, which are the first, second and third element of the array."
I don't think that's right? It should pick out four things, not three.
Of course, you could argue that getting that wrong proves your point! I personally still prefer 0-indexing, but I definitely agree with you that having an exclusive upper bound is confusing.
0:3 picks out index 0, 1, 2.
So 1:3 picks out the second and third element, which are indexed by 1 and 2 respectively.
My guess is that what one prefers probably depends entirely on the data structures one works with most....
Edit: I think I screwed up that analogy.
0-based indexing won in programming because it is simply more to purpose in programming. We rarely operate on ranges, but we operate on offsets all the time.
And 0-based wasn't solely "C won so 0 won". We had 1-based languages for a LONG time, and, if they were sufficiently superior, they should have displaced C. They did not.
In addition, in proper programming languages you don't count--you iterate, fold, accumulate, etc.--and avoid the index altogether because it is error-prone.
1-based indexing causes all kinds of havoc in circular ranges. In particular when you try to access things in a circular manner (very common in programming--uncommon in mathematics), it causes grief.
// 1 based
new_index = index % N + 1
new_index = (index - 1) % N // Careful: the parentheses are REQUIRED
new_index = (new_index == 0 ? N : new_index)
// 0 based
new_index = (index + 1) % N
new_index = (index - 1) % N
Range discussion from Dijkstra in 1982: https://www.cs.utexas.edu/users/EWD/transcriptions/EWD08xx/E...
I still prefer 0 based indexing (for example for calculating the size of an array), but the worst thing is inconsistency between languages. 0 won, that's it.
Python:
>>> -1 % 20
19
C: printf("%d\n", (-1 % 20));
-1
Even worse, it still fails even on unsigned: printf("%d\n", ((unsigned int)0 - (unsigned int)1) % 20);
15
So you still need to add N: printf("%d\n", ((unsigned int)0 - (unsigned int)1 + 20) % 20);
19
Thanks for the reminder of humility.(Side Note: For those reading this, the C operator isn't "broken", per se. There are three properties that modulo can adhere to but two of the three are mutually exclusive.)
The Dijkstra discussion is only aesthetic preference, nothing more.
Edit:
In Julia the examples would also idiomatically be written in terms of the provided mod1 function:
new_index = mod1(old_index + 1, N)
new_index = mod1(old_index - 1, N)Possibly, but then 1-based indexing certainly isn't enough of a positive to overcome the other stuff. And that's evidence, too.
> The Dijkstra discussion is only aesthetic preference, nothing more.
Dijkstra's comment says that people using the other 3 conventions were committing more errors--that's data.
> In Julia the examples would also idiomatically be written in terms of the provided mod1 function:
Agreed. The proper way is to encapsulate that behind a function so you don't have to think about it.
However, if you have to unpack that and repack it all the time (for example, Lua calling C), then you can't just encapsulate and forget about it.
I would disagree with that assertion. Maybe it was true historically, but just look at how often ranges are used today - the fact that many languages have an abstraction for them in the core library is a testament to that.
I would also argue that having ranges (and underlying iterators) as opaque abstractions is preferable to conflating them with indices. Then you can have your cake and eat it too - the elements are counted naturally, but if you have an iterator to the first element, you can deal with 0-based offsets just as naturally.