Why numbering should start at zero (1982)
cs.utexas.edu
cs.utexas.edu
:-)
Dim myArray(1 to 10) As Integer
Don't know if it still works, though.I'm thinking this is like people who drink coffee will cite studies saying it is good for you, and people who hate it will cite opposite studies.
my @arr = (1, 2, 3);
for my $i ($[...$#arr) {
...
}
It is however, advised not to change $[.I think it's only ever been a matter of preference as you say. And Julia was otherwise such a pleasure to write code in that it was easy for me to pay the small price to learn that.
Regarding some of the advantages of zero-based:
- Indexing backwards from the "end". Your language can always add an `end` keyword like Matlab does, and this stops
being an advantage.
- Indexing cyclically using modular arithmetic. Yes, this is an advantage. Albeit a rare one for me.
I commit less off-by-one errors with 1-based, and I don't have to double-check as much -- so on balance, I prefer it.[edit]
The amount of karma this comment is getting is undergoing something like Brownian motion.
for i = 1:numel(s)
foo(s[i])
end
That is fine, but what if you want to process in blocks of N? Now you have to do: for i = 1:(numel(s)/10)
foo(s[(i-1)*N + (1:N)])
end
Ugh no thanks. This is as simple as it gets too. For more complex array manipulation... enjoy adding and subtracting your 1's. 0-based indexing is just way more natural.You could also do
for i = 1:10:length(s)
foo(s(i:i+9))
endforWith numpy the 9 turns into a 10 by the design choice that s[a:b] only goes to s[b-1], this is sometimes convenient, and sometimes it costs a +1.
This is totally subjective. Starting at 0 is more intuitive to me.
My intuition is very simple. `x` is the name of some memory location. `x[k]` is the location `k` spaces away from `x`.
If we have a sequence a_1, a_2, a3, ... we can talk about a_3 by calling it the third term. If we have a sequence a_0, a_1, a_2, ..., the third term is actually a_2.
Whether or not we should index starting at 0 or 1 is probably dependent not only on intuition, but the application at hand. For most analytic purposes it's generally more useful to talk about the nth term, and we don't need to know a specific number to reason about the distance between any two indices. For other purposes, such as programmatic ones, it is useful to know e.g. the traversal distance between two items in a list.
In my opinion it's best to first consider whether you're working in more of a mathematical or programmatic context, and then secondarily who will have to read it later on.
The fact that our ordinal numbers are closely connected with the off-by-one cardinal numbers (e.g., "third", meaning the element of a sequence in position 2, is closely etymologically related to the word "three") is an unfortunate defect of language.
It's still the "third" term. People commonly refer to the "zeroeth" term in a list as the "first" term, the "first" term as the "second", and so on. Admittedly, the usefulness of "zeroeth" depends on how often you think about the mechanics of array traversal, and that's probably not often if you don't program computers.
0 and 1 are both perfectly fine starting points. They both have advantages and disadvantages.
0 is more common since a lot of languages use C as a starting point in some fashion. C chose 0 because then it can create syntactic sugar for sequential memory access via pointers. Or arrays.
Having used Matlab a bunch, in practice this sucks, because “end” is treated as this weird special case in the language grammar, and many reasonable and convenient expectations of syntax that should work turn out not to. Folks writing complicated Matlab projects end up needing to work around it, and the workarounds are brittle and confusing.
Using negative integers is a whole lot easier to reason about and work with.
Everyone learns in grade school how to do arithmetic with negative integers (adopted by European mathematicians in the 17th century). Even Matlab experts don’t always understand Matlab ‘end’ arithmetic.
Only when counting. Why should 'numbering' be the same as 'counting'? I don't see why that would lead to a generally more intuitive when most of what you calculate with 'numbering' is offsets and indexes. Offsets begin at zero. Indexes could be by ordinal and not by offset, I suppose, but I don't see how that could be a benefit in typical index calculations--mostly they're just offsets in my world, distances between indexes.
I rarely ever refer to specific, literal indexes, so I can't say that I have much of an opinion on them, but that does seem to be the area where indexing by ordinal would make the most sense. Perhaps this is the main operation that dominates numerical calculations? Rails even adds english ordinal methods (e.g. first, second, third, fourth, fifth) for this purpose.
Wow, it really IS a treat. I've always heard people say how much a person's handwriting can tell you about that person. Mine is a mangled mess, and I'm pretty disorganized in most areas of my life.
http://lucacardelli.name/indexartifacts.html (go to Artifacts > Fonts tab).
E.g. if you divide a matrix of 100 columns into 20 vertical bands of width 5 each.
Mathematicians use 1-based indexing for both the element index and the band index, so there band n would start at coordinate "(n - 1) * 100 / 20 + 1"
For a programmer, band n would start at "n * 100 / 20"
That's two correction terms that you need to add in math which programmers don't!
I had to use Matlab for microphone arrays once and it was full of + 1's and - 1's everywhere due to that.
Another example of mathematics and off by one errors: a polynomial. They call it "degree n" if the highest power is n, except I see n+1 coefficients in there and need to allocate an n+1 sized array to contain its coefficients, so why not call its degree the amount of terms, including the "x^0" one. The powers themselves in the polynomial are already hinting at 0-based indexing in this case.
Mathematicians, please use coordinate "0,0" for the top left element of a matrix :)
an array with elements at index 0,1,2,3 has size 4
I don't see how you can view calling a linear polynomial degree 2 an example of 0 based index.
> To denote the subsequence of natural numbers 2, 3, ..., 12 without the pernicious three dots
What's so pernicious about them? I don't see it. It seems like a clear and intuitive way to communicate a sequence to me. I even wrote a little range generator in JS to explore parsing declarations like that. https://github.com/chrisbroski/iterize It seems to work fine.
For example, does 3, 5, ... 11 mean the odd numbers between 3 and 11 inclusive, or the primes between 3 and 11 inclusive?
Just decide on some rule for how decide and let people make their own standard if they don't like it.
2,3,4,12 is also a "subsequence of natural numbers", but probably not what was meant by 2,3,...,12. You might counter that ... means the complete subsequence, but what do you mean by "complete"?
If I write 2,4,6,8,...,24, you probably want the ... to mean 10,12,14,...[1], not 9,10,11,12,...
Basically by the time you make ... precise, you are better off just writing the mathematical notation.
[1] See what I did there? ;)
Aside, Antony Jay cowrote of Yes Minister and Yes, Prime Minister.
https://en.wikipedia.org/wiki/Antony_Jay
I'm familiar with the shows but didn't recognize his name.
I wanted the best of both worlds, and so overloaded the () operator so that if arr is an array, then arr(i) = arr[i-1].
This worked reasonably well, especially for arrays where it was always clearer to go 1-based, or arrays where it was always clearer to go 0-based, so that the array was always accessed with the same operator.
The only places it was questionable whether or not it made the code clearer was where I used both in the same section of code. E.g., something like b = a[i] + a(i) is arguably less clear than b = a[i] + a[i-1] or b = a(i+1) + a(i).
[1] I think it was implementing an arbitrary precision integer library using algorithms from TAOCP, but I don't remember for sure.
for(uint8_t i = 0; i < 256; i++) { ... }
What's even worse, the following loop has undefined behavior because signed integer overflow is not defined: for(int8_t i = 0; i < 128; i++) { ... }For example, Python:
for i in range(0, 256):
Rust: for i in 0..256 {
and many other languages have similar things.Proper ranges also mean that the compiler never needs to do complicated reasoning that may depend on signed integer overflow having undefined behavior in order to prove that the loop has a finite number of iterations or that `i` doesn't wrap around.
We commonly have syntax for the first, or sometimes some variant of it designed to avoid the `-1' in the common case, at the cost of complicating some other cases. It's also common to have some rather over-general iterating mechanism called "for" that covers both, or tries to.
But I don't think I've ever seen anything special for the second, for some reason. And the iteration counter and the loop value could have different types - so you could count 128 int8_t values if you wanted. Though in most cases you'd just use it as a replacement for "for(i=m;i<m+n;++i) {do stuff with i}" or "(for i=0;i<n;++i) {do stuff with m+i}".
(One possible exception: the ancient 1980s computer I used as a boy had two syntaxes for saving blocks of memory. You could enter the range as "A B" (save from A to B inclusive), or as "M+N" (save N bytes starting from M).)
for(uint8_t i = 0; i <= 255; i++) { ... }
then it still doesn't halt.You can do it with a do { } while loop, which only needs N evaluations of the termination condition for N iterations.
As far as I can recall, I've seen both 0 and 1-based indexing in statistics.
https://github.com/hypotext/notation#mathematical-notation-p...
There might be some other tricks you can play depending on your instruction set, but nothing comes to mind.
It's clear that the author structured his essay carefully to avoid this exact assumption. He explicitly avoids considering whether the natural numbers start at 0 or at 1 until after he has chosen "a)".