Why do arrays start at 0?
buttondown.email
buttondown.email
Whether or not the language presents the concept of "pointer" to the user is independent of whether or not it uses pointers internally. And if it exposes arrays as a concept, it has to implement them somehow.
The simplest possible implementation of arrays is having a start address and putting all elements next to each other in RAM. To get the address of a particular item, this layout naturally leads to the formula "base address + index * element size", with "index" being 0-based. If you want to expose other indexing schemes in your language, you'll have to add more logic to convert the user-visible index back to 0-based before you can obtain the address.
All of this is completely independent of the fact whether your language exposes pointers to the user or not.
Even the yacht story sort of hints at this:
> To keep people from having their jobs cut short by yacht-racing, Richards designed the language to compile as fast as possible. One optimization was setting arrays to start at 0.
If all indexing schemes were equal, why would this change even be an optimisation in the first place?
If 1-based indexing was used in assembly, then "mov 1(%ebx),%eax" would be the equivalent of "mov (%ebx),%eax".
at compile time, not runtime.
You'd still have the odd subtract here or there, but for many use patterns you could probably ignore the overhead.
You can, but don't need to. Just have compiler store array pointer constant as (arr-1) instead of arr, et voila, zero runtime overhead.
Or, for modern(-ish) ISAs, often you can add/substract a small constant at runtime, with no extra cycles taken. For example, for x86_64:
# rbp contains "true" pointer to arr
# rax contains 1-based array index
mov rax, QWORD PTR [rbp-8+rax*8]Yes it is. It was evaluated, carefully weighted and discarded, as it was not worth it.
A benefit is that you can then use 0 as a sentinel value; for instance if you have a find() routine that surely can fail, it can just return 0 instead of having e.g. -1 (which can introduce minor issues).
In my experience, though, I am so used to 0-based index that switching schemes can cause stupid off-by-one bugs. I guess that's the main reason behind complains about Lua. It's not that "natural" arrays are thought of as bad, but mixing both schemes (often C an Lua) is error-prone.
Whether that's what actually happened in the cases of HLLs that adopted zero-based arrays or not, I don't know. But today, to me, zero-based array indexing feels very natural, and the idea that zero-based arrays being simpler in assembly carrying over to HLLs seems at the very least plausible.
int foo[n]; // what the user wrote
// what happens behind the scenes, after a fashion
int* foo = malloc(n * sizeof(int); // or sp - n * sizeof(int) if stack allocated; subtraction since stacks usually grow "down"
foo = foo - 1;
All references into foo will now work just fine so long as they are within the [1,n] range (same issues as with 0-based there since C doesn't carry size information for checking array bounds access). This adds one extra instruction per allocation (which includes allocation on the stack) for all non-0 offsets, but then all access will have the same cost whether 0-based, 1-based, or arbitrary-based. That's a non-zero cost, but it's not exorbitant since you'll be accessing much more often than allocating (and if it's reversed, something weird is happening).I agree with you the subtraction instruction is not exorbitant, especially considering malloc is much more expensive and always has been.
// 11 arrays are created here
int foo[10][10];
It doesn't make sense to create pointers for the 10 inner arrays, so the subtraction would presumably happen when referring to each inner array: int *z = foo[x]; // (char *)foo_sub_1 + (sizeof *foo)*x - (sizeof **foo)
foo[x][y]; // (char *)foo_sub_1 + (sizeof *foo)*x - (sizeof **foo) + (sizeof **foo)*y
In any case, this is more work, both for the computer and for the human. It's analogous to trying to figure out "which year of which century is X in?" (2022 is the 22nd year of the 21st century): year(x) = (x - 1)%100 + 1
century(x) = floor((x - 1)/100) + 1
The above formulae work for the 1-based convention, because they do the necessary 1 subtraction/addition in the right places. If we instead counted from 0, things would be a lot simpler (2021 would be year 21 of century 20): year(x) = x%100
century(x) = floor(x/100)
Unfortunately, the Romans started counting years before anyone knew about 0. I think this is basically why we have a convention of counting from 1: 0 simply wasn't discovered until relatively recently in human history. Earlier number systems such as Greek/Roman numerals don't have a way of representing it.or how would you count, say, 3 Apples? or three legionaries?
You have to remove all desks from #20 to #30, how many desks are there?
Oh, it's 11 desks, I see. This is because we used INCLUSIVE indexing, instead of semi-open interval.
If we use semi-open intervals, we don't include the last item, so that we can write
match index {
0..10 => println!("first ten"),
10..20 => println!("next ten"),
_ => println!("something else"),
}
but this forces us to start our intervals at 0, otherwise we would have to write 1..11 which would be awkward`a..(a+n)` is an n-length range, not an (n + 1)-length range (oh look, an "add 1" operation!).
And an empty range is denoted by `a..a`, not `a..(a-1)` (oh look, a "subtract 1" operation!).
Looking at the Swift standard library, end-exclusive ranges seem to be far more commonly used:
~/build/swift$ git grep -e '\.\.< *[a-zA-Z0-9_(]' --and --not -e '^ *//' -- stdlib/\*.swift | wc -l
599
~/build/swift$ git grep -e '\.\.\. *[a-zA-Z0-9_(]' --and --not -e '^ *//' -- stdlib/\*.swift | wc -l
112I don't think it's inherently more natural to start from 1, just conventional. Disregarding history/convention, I think it would be more natural to use the lowest available natural number.
Back in Roman times, the lowest natural number that people were aware of was "1", so obviously they started counting from that number.
Our understanding of and use of mathematics has evolved since then, and accordingly there are fields such as computer science and combinatorics where there are clear advantages to starting from the smallest number (zero). In virtually all other cases, the reason "1" seems more natural is because that's the way it's been done historically.
It seems that when labelling those apples using a 1-based count, the logic is basically: each apple is identified by the number of apples that precede it, plus 1. The reason for the "plus 1" is that that was your starting number, but it could have easily been 2 or 3. If you instead start from 0, you can omit the "plus X" logic, just as I omitted the "+ 1" and "- 1" logic in my year/century formulae when moving from 1-based to 0-based counting.
> It seems that when labeling ... number ... that precede it, plus one.
I don't think so. It's the number that you have counted once you've counted that one.
It was introduced into western mathematics through Fibonacci in the 1200s at the same time that Hindu-Arabic numerals were adopted, which use "0" as a placeholder (compare this to earlier Greek numerals which work similarly to the system we use today but without placeholders and using different sets of symbols for the different places—and of course no way of representing zero).
It'd be a pretty huge waste of processor design space to do that, though: if an HLL wants to support n-based arrays for n != 0 it can just store the base pointer with the initial index offset already applied.
Early 8-bit processors had a pretty poor set of addressing modes: you couldn't even expect [R1 + R2] let alone [R1 + R2 * scale + displacement], so random access into an array would be a multi-instruction task. By the time of the 80386 base scaled index addressing with displacement was in the CPU core, but it's not _free_ - there's an additional byte in the instruction encoding for the displacement, so even if the CPU computes the displacement with no additional clock cycles you'll have a latency cost for fetching the byte.
Note, for example that the mighty 6502 succeeded in the market because its designers had heard from customers that the $100 price for the 6800 was too much, so they removed some instructions and addressing modes to get the die small enough that it could be sold for $20.
No, you misunderstood. You subtract one from the pointer at allocation time. You don’t add an instruction every time you access.
Yeah today memory access on a cache miss is really expensive. I think memory was normally running at a different clock rate than the CPU even 50-60 years ago, so has always been something where you can’t touch memory every instruction. The latency of a cache miss wasn’t nearly as bad in the past. When the data cache was invented, a main memory access was like 4 times slower than register access. Today, main memory access can be more than a hundred time slower, sometimes it can approach a thousand times slower, which is why we always have multi-level caches now. And memory latency is still getting worse. If someone figured out how to make memory faster without increasing the cost or the energy consumption, they’d be rich! ;)
Not necessarily so for translation of an index that is evaluated at run time. PL1 (back in the 1960's) compiled arrays into 'dope vectors'. Each array's dope vector included highest and lowest valid subscripts for each dimension of the array and the address where the element with all zero subscripts would be found, if it existed.
Which is sad, because all code is still executed as machine instructions even when the developer does not see it or does not want to care.
The code the programmer wrote is compiled to something such as LLVM IR, LLVM IR is further compiled by LLVM to Assembly, this is further compiled by an assembler into machine code, and then the c.p.u. further compiles this to it's internal code as it executes it. “machine code” really is no more special in this chain of events than, say, LLVM IR.
I think this is the core point. Even though you are mechanically feeding machine code into the CPU, modern CPUs deeply transform your program before (or even while) running it, with potentially deep performance implications. With modern CPUs, the only choice we truly have is which compiler we trust to optimize our program.
x86 has a developers manual for this purpose.
We are way beyond MS-DOS days and Michael Abrash's books.
Why is the c.p.u. special?
But some people, yeah, they could do it.
That has not changed.
Certainly one can bang out code their entire career without ever having a clue how machine code works, but I really wouldn't advise it. At worst it leads to total ignorance, and at best you accumulate a disconnected set of "best practices" as inscrutable lore handed down from on high.
But it probably ran in like 100us anyway
Anyway, complexity stays relevant, no matter what language you use. That's the real deal.
it's not useful to know how the browser allocates memory or garbage collects it, it's more useful to know that it has web sockets. Don't get the idea that things you know are relevant to other people.
Abstractions are good, and while we should be careful not to completely trust a brand new abstraction, high level languages are established enough that we don't need to worry about it.
A pointer is just one kind of array-like indexing scheme. Good pointery languages will distinguish Address from Offset from Integer.
Is that so? I wasn't aware of that.
Now you'll just bring out the people who start the naturals at 1.
Why not? "Every cardinality, except of the empty set, is a natural number" isn't very convincing to me, if we're making cardinal-based arguments.
Dijkstra's argument (https://www.cs.utexas.edu/users/EWD/ewd08xx/EWD831.PDF) is that we count the number of predecessors. This fits well with the mathematical usage, at least among mathematicians who care to dig into the order-theoretic foundations: the von Neumann construction of ordinals (and of cardinals as least ordinals of a fixed, well, cardinality), each finite ordinal is equal to its cardinality: 0 = \emptyset is a set with 0 elements, 1 = {0} is a set with 1 element, 2 = {0, 1} is a set with 2 elements, etc.
If you allow 0 as a valid ordinal, it unifies ordination and measure. That's the best reason to do so.
There. Djikstra as haiku.
Just to be clear, I wasn't making some sort of bizarre appeal to authority; whether something makes sense is a matter of taste, not of mathematics, and can't be proven or dis-. I meant only to respond to:
> In maths they always talk about the first element in a vector, not the zeroth. It is a bad argument to point to mathematics for using zero based indices.
I don't know for sure which convention is more common, though I suspect you're right that it's the 1-based convention, but I do know that it's false that 'they'—meaning, among others, we!—always talk about the first element/component in a vector. I probably will do that in a first Linear Algebra class, because the textbook often does and I don't want to introduce unnecessary confusion; but usually in my own work, and sometimes when I teach upper-level math classes, I use 0-based indexing when it's necessary to choose.
I think an age-based way to phrase it: in your 1st year, your age is 0; in your 2nd year, your age is 1; and so on. We can assign people numbers indicating what year of their lives they're in, or how many years they have lived, and both are fine, but we've settled on the latter.
That's not an argument. It's a coincidence. There are applications where what you care about is the number of predecessors (indeed, that's what the compiler cares about, which is why we have 0-indexing in the first place), but they are a tiny minority of all indexing.
> I think an age-based way to phrase it: in your 1st year, your age is 0; in your 2nd year, your age is 1; and so on.
But that isn't even true. No one ever reports the age of their new child as 0; instead, they will report a positive number of months, or -- if it's an extremely new child -- of weeks or days.
I think there's nothing to say to the first two sentences but that I regard it as an argument that may be more or less convincing. I don't know exactly what it means for something to be an argument vs. a coincidence; it is a coincidence that, say, my name is what it is, but it is nonetheless correct for me to argue that that is my name.
These are all conventions anyway, and there is not much use (or even meaning) in arguing about which one is the right or wrong convention, just which one makes more or less sense; and this is one way to make sense of 0-based indexing, though of course there are also ways to make sense of 1-based indexing.
I'm not sure I buy that these cases are a tiny minority, but I'm certainly in no position to produce any data to the contrary.
> > I think an age-based way to phrase it: in your 1st year, your age is 0; in your 2nd year, your age is 1; and so on.
> But that isn't even true. No one ever reports the age of their new child as 0; instead, they will report a positive number of months, or -- if it's an extremely new child -- of weeks or days.
I can believe that ages aren't reported that way, although I think that a child who will be 1 year in 1 year should logically be said to be 0 years; but we can avoid that debate by considering future years: in, to pick the example that applies to me, my 42nd year of life, I am 41 (I could say 41 and 1 month, but I don't—in fact, there's a Seinfeld joke about that). Similarly, the 42nd entry in a 0-indexed array is indexed 41. It doesn't have to be that way, but I don't think one can argue that there's anything logically amiss about it (nor about 1-based indexing … but we do have to pick one).
Sure. But we're talking about whether to index arrays from 0 or 1. It is true that naming an array element after the quantity of its predecessors will tell you the number of predecessors the element has. But that's not an argument for why you should do it; there would need to be some kind of benefit to having that information. Without a benefit, it's just something that happens to be true.
That's the difference between an argument and a coincidence.
A dichotomy between "ℤ+" on the one hand versus "ℕ" on the other is just plain more convenient than a dichotomy between "ℤ+" and "ℕ" on the one hand versus "ℤ\ℤ-" on the other.
I prefer to start by zero; I think it is more useful in general, and makes more sense mathematically for many (although not all) purposes, and systems that you will find objects and operations that have these properties too.
Well, no, obviously you want to index by ordinals, and they start at "first".
the positive integers (whole numbers) 1, 2, 3, etc., and sometimes zero as well.
It all makes perfect sense in the context of measuring.
Maybe a null terminated string could have length 0 if you don't count the terminator. But that 0 is a property of the "string" abstraction. The actual "array" would still be length 1.
Nope. The moment you "start" somewhere you occupy 1 unit of memory. Thus no longer an "array of length 0".
There is no such thing as an array of length 0. It absolutely does not exist. You cannot write source code to represent it.
ie assuming we are talking about the pure data structure of an array. Some languages may have some abstraction built on of arrays (ie c-strings) that can have length 0 but these are not "arrays" and if they are defined they still occupy 1 unit of memory for the terminator.
https://play.rust-lang.org/?version=stable&mode=debug&editio...
it's an array of length 0, but it can't be "used" because it has a length of 0 so it occupies no RAM
But such an array does not "start" at 0 as it does not exist. You need an array of at least length-1 to "start" at 0.
`for(int i=0; i < 0; ++i)` and it will run 0 times, the exact amount of times I expected
This loop is not processing any physically existing array. In fact the compiler would remove this block of code as it is a no-op.
iopq points out Rust as a language that does allow zero length arrays in a sibling comment.
While ISO C does not allow zero length arrays, GNU C does. There’s also flexible array members, which can be length 0.
Sure a language can define some abstraction that says "hey I'm length-0 array!". But this is more like a statement of "hey I don't exist yet, but if i do exist in the future I'll be of type array and at least length-1".
You can't access index 0 of a length-0 array.
You can only access index 0 of length-1+ array.
To take your ruler example and apply array logic to to. Say you are measuring a 2cm thing. If the ruler follows the index then it would show 1cm (the last index), and so you would have to add + to get the size/length.
Having the index not align with the size is a pain that results in a lot of +/-1 code that a compiler could just as easily have handled.
I'm not saying it has to be this way and it can't start with 1, I'm saying what you are pointing out as a flaw is actually the mixing of systems.
in computer programming teams it's best to adopt a standard unambiguous language to communicate to one another so as not have to constantly say "do you mean...?"
0..4
the second int occupies bytes
4..8
you can see the array up to the second int occupies the contiguous space 0..8 which 0..4*2 and it has TWO elements
The easy way to do that is by shifting the base address. In pseudo-C (I think that computing the b pointer is non-conforming, even if the code never tries to access b[0], but compilers can do this without problems):
int a[100]; // array of 100 integers, zero-based
int * b = a - 1; // array of 100 integers, one-based
Things only get costly when you want to check array bounds or when you have multi-dimensional arrays. There also may be non-standard architectures where this kind of stuff isn’t possible.“In other words, if the expression P points to the i-th element of an array object, the expressions (P)+N (equivalently, N+(P)) and (P)-N (where N has the value n) point to, respectively, the i + n-th and i − n-th elements of the array object, _provided_they_exist”
[…]
If both the pointer operand and the result point to elements of the same array object, or one past the last element of the array object, the evaluation shall not produce an overflow; _otherwise,_the_behavior_is_undefined.
In this case, the minus-one-th element doesn’t exist, so the expression
int * b = a - 1;
triggers undefined behavior.I think some compilers use this in practice to produce faster code (that, often, will not do what the programmer expects it to do). Start reading at https://stackoverflow.com/questions/56360316/c-standard-rega... if you’re sure they don’t. I expect that will change your opinion.
Not that I would want to do it, I think zero-based addressing is not very taxing for the convenience of being closer to how we think of memory addressing.
It feels like you're introducing a million edge cases.
[0] https://en.wikipedia.org/wiki/Performance_Handicap_Racing_Fl....
The element at offset 0 is the first element.
Precision in naming is important.
It's also a definition commonly used in math, through not always. It depends a lot of the area of mathematics and even the cultural context.
The only reason we sometimes feel 0-index is wrong IMHO is because the english language describes entities in a sequence as 1st, 2nd, 3rd etc.
But I wouldn't be surprised if there is some human language somewhere which does it differently. Like something which roughly matches to english like "head, head+1, head+2, ...".
Depends which country your building is in, and more!
My apartment building is 1,2… but my mall is G,M,1,2… (same country different architects). For many years I lived in Europe where it’s usually G,1,2… (where G can also be E or Fsz or whatever) and when I was in the US I had to remember that 1 is G, though L,2,3… is also common.
City people live with index ambiguity all the time and our brains manage just fine.
Confused me when I visited USA.
> Confused me when I visited USA
It's applied inconsistently enough to be confusing for natives too!
Do you count the machine floors in high rise buildings?
In AU, the ground floor is effectively "zero", the first floor is the one above ground. Depending on the design of the building, there might be a "mezzanine" floor between ground and 1st.
Basements usually follow the same pattern, so B1 is the first floor below ground etc.
0.5 ... yep that won't work
In some buildings build around 1900 where I live the "ground" floor is not on the ground but around 1/2 a floor above the ground floor. While the cellar is just half below the ground. Because of this the cellar has above ground level windows allowing legal cellar apartments (oversimplified in Germany apartments require windows which are above ground).
This creates a lot of confusion.
By convention this 0.5 floor still counts as ground floor with the floor number 0. But people confuse it all the time for the first floor. Especially if in some cases you do have a few utility rooms exactly at ground floor.
Now things get worse if you live in a cellar apartment in such house: By convention floor -1.
Because nearly everyone will thing you mistyped as normally you legally can't have apartments below the ground floor.
Another fun index ambiguity is for apartments build on a mountain side/cliff.
In this case depending on which side of the house you look at the ground floor might be at, above or below the ground level. Often the floor level with the main entrance counts as ground floor, but sometimes it is not the case and e.g. the lowest floor not below ground or the highest floor even with a "ground".
So yes, index/floor level ambiguity is everywhere and much worse then just is ground floor 0 or 1 ;)
But it's less precise than offset and obviously introduces confusion. Whoever originated the term "0-indexed" should have just suggested we use a better word rather than keeping the inaccurate word (proven by the fact you're currently looking for a way to remove some confusion around the currently chosen word) and prefixing it with a digit, which itself is a confusing thing to do to an English word, thus adding to the confusing while subtracting none.
Also it's not less precise, as there are many different kinds of offsets too. First you have element wise and byte wise and similar offsets, then offsets are inherently relative to something, and that's not always the first element. But could be an offset from the back and I have even seen a offset from one element before the first element in some very unusual use-cases.
It really isn't. The human race gets by without it the vast majority of the time.
It works better with the modulo operator: `array[i%length]` vs `array[(i+length-1)%length+1]`. Or you would have to define a modulo-like operator that maps ℕ to [1..n].
It works better if you have a multi-dimensional index, for example the pixels in an image. With 0 based indexing, pixel `(x,y)` is at `array[x+widthy]`. With 1 based indexing it is at `array[x+width(y-1)]`. You might argue that programming languages should support multi-dimensional arrays, but you still need operations like resizing, views, etc.
for (size_t i = 0; i < length; i++) ...
The following causes an unsigned integer underflow and is an infinite loop: for (size_t i = length - 1; i >= 0; i--) ... for (size_t i = 0; i < length; i++) {
size_t j = (length - 1) - i;
...
}
EDIT: change i to j const j = (length - 1) - i;
in that case. Much safer.Makes my brain hurt, but I think this will only run through the loop one time looking at the last element of the array.
for (size_t i = length; i-- > 0; )...
It's surprising how often the issue pops, it works well with both signed and unsigned integers.(edit) I've started with one based indexing (basic)... mixed with 0 based (assembly), more 1 based (pascal), then more stuff (all zero based). I am, yet, to see a real advantage of a one based indexing... after the initial process.
If performance matters, you should experiment with __builtin_prefetch, which is available in clang and GCC.
for (size_t i = n; i-- > 0 ;) ...
Or count from `length` to 1, but subtract 1 in the loop body, or count up and subtract the length in the loop body. Any modern compiler should be able to optimise these to be equivalent.In the majority of cases, counting down is not necessarily. Nor is ordered iteration. Most languages have a `for each` style syntax that's preferable anyway.
size_t i = length;
while (i--) ... for (size_t i = length; i-- > 0 ; ) ...
that several other people posted would not execute for index 0. Shouldn't it be this instead? for (size_t i = length; --i > 0 ; ) ... for (size_t i = length - 1; i < length; i--) ...
You are also free to start at any other (not necessarily in-bounds) index, just like with ascending iteration.[1] https://gustedt.wordpress.com/2013/07/15/a-praise-of-size_t-...
Seriously though, while the idiom does work for unsigned integers, it's a bad idiom to learn [makes code reviews harder]. The post-decrement one in the loop body works with everything (signed/unsigned), and it's well known.
why doesn't that loop end instantly?
I mean length - 1 < length should always be true, right?
Or does it only terminate when the number underflows? Terribly confused here
for(initialize; condition; increment) { ... }
to: initialize;
while(condition) {
...
increment
}
(more or less, some scoping things not encompassed by the above; this is also how pretty much every for loop in a C-syntax language works) The condition of a for loop is equivalent to a while loop's condition. So yes, length - 1 < length will be true on the first iteration, which is fine because the loop continues as long as that condition is true.What the above approach takes advantage of is that when underflow eventually happens you'll have this condition:
MAXINT < length
Which will terminate it for all possible values of length.Also, that behavior is not guaranteed. The programmer would need to be aware of how the particular machine in question actually handles that.
Then again, that's C.
for (size_t i = length; i > 0; i--) {
// ...
item = array[i - 1];
(This is how reverse iterators work in C++.)Not hard to figure out.
> Not hard to figure out.
While that's true, "you just shift by 1" is equally good at all arguments for or against 0-based indexing, so deploying it here probably won't convince.
That said, the effort of one versus the other is so trivial that there is no point in ever using effort as an argument either way. Doubly so because what seems like effort to us is simple unfamiliarity.
What is important is which one leads to more careless errors in practice. As a trivial example, consistent indentation takes effort, but failing to do it leads to more careless errors. Therefore everyone indents code.
The only data point I've seen on that is the side remark about Mesa in https://www.cs.utexas.edu/users/EWD/transcriptions/EWD08xx/E.... That remark, therefore, is the only argument that I care about.
That would be because any game worth playing has at least one player... and so it's natural to continue from there. (In terms of language.)
Using the term "zeroth" is basically some form of showing off (even though it's kinda fun), but will be utterly confusing when you get to the fifty-second element which is the last in a group of 53 elements.
That said, not all languages have given up on this. For example Julia allows it. https://docs.julialang.org/en/v1/devdocs/offset-arrays/
Ironically I learned this from a discussion of Julia bugs. Apparently changing offsetting of arrays has proven to be a source of bugs in Julia. So maybe someday they will come to the same conclusion as languages like Perl and stop allowing it.
Therefore changing the base leads to more bugs than either base alone.
That said, the more you can just use a foreach to not worry about the index at all, the better.
Of 0-based and 1-based, the only data point I have is a side comment of Dijkstra's that the language Mesa allowed both, and found that 0-based arrays lead to the fewest bugs in practice. I'd love better data on that, but this is a good reason to prefer 0-based.
That said, I can work with either. But Python uses 0-based and plpgsql uses 1-based. Switching back and forth gets..annoying.
Search for OffsetArrays in https://yuri.is/not-julia/ for practical problems encountered in trying to make this feature work in a language whose compiler does try to be smart.
The existence of a mechanism does not guarantee its correct use. As that link demonstrates.
and the example is in 24hr format - which needs to have the 00 to differentiate it from being 12:30. But if you write in 12hr format, you don't ever use 00 - it's always 12:30am or 12:30pm
But yea, I do agree with you. we don't say 0, we say 12, no matter if it's noon or midnight. That's because we humans avoid saying zero when we mean zero.
It's the same with other comments here, talking about counding seconds, we say "ok go, one, two,.. and not "zero, one, two,..". We say 3 months old baby, and not 0 years old baby.
We use 0-index in so many things, we just avoid saying the word "zero", and we use other names or other units to avoid that word.
Consider:
Characters h e l l o
Cursor index 0 1 2 3 4 5
Char index 0 1 2 3 4
Range [0,3) [0,1,2]
Range [2,5) [2,3,4]
Range [1,1) []
If we used 1-based indexing and exclusive ranges, it leads to ranges where the end index is greater than the string's length... Characters h e l l o
Cursor index 0 1 2 3 4 5
Char index 1 2 3 4 5
Range [1,4) [1,2,3]
Range [3,6) (!) [3,4,5]
Range [2,2) []
but if we use inclusive ranges, it leads to ranges where the end index is less than the start index... Characters h e l l o
Cursor index 0 1 2 3 4 5
Char index 1 2 3 4 5
Range [1,3] [1,2,3]
Range [3,5] [3,4,5]
Range [2,1) (!) []
Also: Characters h e l l o
Cursor index 0 1 2 3 4 5
0-based range [0,3) [0,1,2]
1-based range [1,4) [1,2,3]
for the 0-based range [0, 3), the left array bracket is at cursor index 0, and the right bracket is at index 3. With 1-based indexing it doesn't work like that because the range is [1, 4)Iterating an array or adding to the end are fine, we have ipairs and insert for that, but ranges on strings I'm constantly having to think harder and write more code than necessary.
I love the language, wouldn't trade it for another, but the 1-based indexing on strings, which represents an empty string at position 3 as (3,2), it's egregious.
Not as egregious as a dynamic language where 0 is false though.
well, C also considers 0 being false (and you can argue that C is "dynamic"!).
To write 4 numbers (including zero) you only need two bits
0: 00
1: 01
2: 10
3: 11
To write 4 numbers if you avoid using the number zero, you need three bits 1: 001
2: 010
3: 011
4: 100
If you extrapolate that a little bit, you'll realize that you'll need two bytes (1 Byte + 1 bit from another byte) to store the indices of an array with 2^8 elements, which is just dumb.Some of you might be thinking: you don't need to store it the same way it's written, you can just substract 1 from whatever the user typed and convert it behind the scenes.
Yes, you could subtract 1, but then you would be making the whole system more complex, opaque, and unelegant, while hiding information from the programmer for no good reason.
0-indexed decimal grid:
0 1 2 3 4 5 6 7 8 9
10 11 12 13 14 15 16 17 18 19
20 21 22 23 24 25 26 27 28 29
30 31 32 33 34 35 36 37 38 39
40 41 42 43 44 45 46 47 48 49
50 51 52 53 54 55 56 57 58 59
60 61 62 63 64 65 66 67 68 69
70 71 72 73 74 75 76 77 78 79
80 81 82 83 84 85 86 87 88 89
90 91 92 93 94 95 96 97 98 99
1-indexed decimal grid: 1 2 3 4 5 6 7 8 9 10
11 12 13 14 15 16 17 18 19 20
21 22 23 24 25 26 27 28 29 30
31 32 33 34 35 36 37 38 39 40
41 42 43 44 45 46 47 48 49 50
51 52 53 54 55 56 57 58 59 60
61 62 63 64 65 66 67 68 69 70
71 72 73 74 75 76 77 78 79 80
81 82 83 84 85 86 87 88 89 90
91 92 93 94 95 96 97 98 99 100The first year CE being "1" resulting in the new millennium starting at 2001 instead of 2000 also seemed idiotic, and the 1900s being referred to as "the 20th century" was something I always had to consciously compensate for. Numbering items starting at zero would have given us more elegant/less confusing ways to communicate those things.
Zero-based indexing makes things inherently simpler because it matches the way we write numbers (and becomes much more noticeable once you have more than one digit). It's not just an optimization for computers.
By AD 150, Ptolemy, influenced by Hipparchus and the Babylonians, was using a symbol for zero.
https://en.wikipedia.org/wiki/0#History
To keep it confusing: the traditional proleptic Gregorian calendar (like the Julian calendar) does not have a year 0 and instead uses the ordinal numbers 1, 2, ... both for years AD and BC. Thus the traditional time line is 2 BC, 1 BC, AD 1, and AD 2. ISO 8601 uses astronomical year numbering which includes a year 0 and negative numbers before it. Thus the ISO 8601 time line is −0001, 0000, 0001, and 0002.
https://en.wikipedia.org/wiki/Gregorian_calendar#Dual_dating
They will get confused because every day counting got it wrong. In a perfect world we would count fingers from 0.
> These all point to one reason why 0-indexing might be preferred: it matches machine semantics more
ldr x, [addr, idx]
it just does sub t1, idx, #1
ldr x, [addr, t1]
If we're talking about C, accessing invalid/OOB indices is undefined and so if idx happens to be zero and unsigned, we'll overflow and hit something unexpected, which is fine.> It's not the reason you think. No, it's not that reason either.
Indexing is a language level concern, it has nothing to do with the generated code.
Unless you write machine code I guess.
IIRC, Visual Basic 6 took a why not both? approach to array indexing.
https://www.cs.utexas.edu/users/EWD/transcriptions/EWD08xx/E...
“… you probably know that arrogance in computer science is measured in nano-Dijkstras.” — Alan Kay
I suppose the 0,1,infinity principle argues that including the empty set and smallest number is more important that including the largest number, so that notation should be preferred by default. And when writing using A<=i<B it is clear enough, but once you remove those symbols and start using them in indexing or function call syntax, like [A:B] or range(A,B), the inconsistency of bounds really breaks my brain, and having to pass a number larger than the largest array index as an upper index range bound gives me nervous ticks. I'd rather use inclusive bounds everywhere and have special syntax for empty set.
https://en.wikipedia.org/wiki/Edsger_W._Dijkstra#EWD_manuscr...
https://www.cs.utexas.edu/~EWD/ - Pick random ones from different decades and you'll see very similar handwriting.
(Didn't read these, picked two at random):
"Natural" zero-based counting at its best.
All room numbers were 5 digits, and skipped numbers if there were multiple windows in one room, based on the theory that such a room might later be sub-divided, and wanting to avoid renumbering of rooms down the hallway.
(I encountered this around 2003, so might have well changed since then; didn't find any references only with a quick search).
But then again people measure distance in feet and write dates as month day year :)
So there was a floor with rooms numbered 1xx and 0xx. There was also a floor below the 0xx floor. They just numbered it 00xx. And you could enter the building on any of those floors depending on what door you picked.
The class I taught was in 0005. Room 005 was someone's office. So on the first day of every semester we would have students waiting outside the door of 005 instead of 0005...
Floors 1 through 10 are "regular" floors. L is the lobby. P1 through P3 are the parking (basement) levels.
So you have three types: regular, lobby, and parking. And the elevator labels are of type regular | lobby | parking.
(Now you can't disagree with me because sum types are popular on HN, and I have framed this in terms of sum types.)
up: (regular | lobby | parking) -> (regular | lobby | parking)
down: (regular | lobby | parking) -> (regular | lobby | parking)
and which_way: (regular | lobby | parking) x (regular | lobby | parking) -> direction
If you use subset of integers you can just use operators inc, dec and <.Yeah I know it was a joke.
Hmm, are there actually an infinite number of floors? Do we need to switch to floating point?
Why separate the floors that are on the same level?
> Hmm, are there actually an infinite number of floors?
I assumed there isn't and that we use a subset of integers. If you have infinite number of floors - feel free to use full set of integers (if you have enough memory that is).
> Do we need to switch to floating point?
No point, if you have more than aleph0 floors - floats won't help you.
0 is ground floor.
This is a great example.
I wonder how many multiples of people who like arguing about starting from 1 vs 0 and how many people were confused for more than a day
So your initial assumption is incorrect from my experience.
I've seen setups like you mention, but they definitely aren't the majority in places I've lived.
To add to the fun the only continuous corridor through the entire T-shaped building was on Level 10, one end of which (the base of the T) was at ground level on its part of the hill. There was a famous 'kink' in the middle of the top / cross-bar of the T (apparently the longest corridor in Europe) because they started building from each end and were slightly off.
The elevator, however, goes from floor '1' (the basement), to floor '5' (the top or fourth floor).
This really confuses visitors.
The Elevator Mafia want $10,000 to fix it.
The 0 makes a lot of sense in a C pointer world where memcpy and other alike functions can be written very thight.
The 1 makes a lot of sense in a human world, when we count, we start at 1, we talk about the "1st", counting on finger starts with 1, etc.
I once were at a Lua (1 indexed language) conference where this was discussed, and Luis started explaining why Lua was 1-index with this sentence: "The 1st argument ...." :)
99% of the times, the correct approach is to use iterators. When you really need indices (and you almost never do), 0 is more practical, because it matches the "including start, not including end" convention.
Although often with an implicit zero. Under typical North American culture, your 1st birthday, for example, is more accurately the first anniversary of your birthday. Your birth is zero indexed.
In some cultures you are considered 1 the moment you are born, so the zero indexing isn't universal here, but typical in North America as noted earlier.
You can't possibly be serious.
Is the empty set countable? (Yes.)
Dictionary:
nat·u·ral num·bers
the positive integers (whole numbers) 1, 2, 3, etc., and sometimes zero as well
Countable:
https://en.wikipedia.org/wiki/Countable_setSet theory:
> Equivalently, a set S is countable if there exists an injective function f : S → N from S to N; it simply means that every element in S corresponds to a different element in N.
Defining N is usually done via a successor set, on which case 0 makes no sense to include.
Standard construction of ordinals is that each ordinal is the set of all its predecessors. (0 has no predecessors , hence 0 is the empty set.) (And so finite ordinals have the same ordinaliity as cardinality).
Birthdays are clearly 1 indexed, the first birthday is indexed with 1, and not with 0.
Birthday[0] gives you an out of range exception, since there is no birthday called 0th. Birthdays are [first,second,..] indexed from 1: brithday[1] = first, birthday[2] = second, and so on. That's what indexing a series from 1 means.
Suppose your job is to count some events. You check in for work at 8:00 a.m., but the first event has not registered until noon. By your logic you should not be paid for four hours, because you're paid to count, and counting started at 1.
Your pedometer?
Your traffic clicker?
We are talking about the he numbering, not the work-doing.
You start waiting at 8, and you wait between events. But you only count (increment) when an event happens.
The 0 comes for free when you are ready to count (hello golang and C++, as well as human intuition).
When you count stars in the sky you don't say "0, 1, 2".
You say "..., 1, 2", or if no stars show up, you say "there are 0" after timer expires.
Increment what?
> When you count stars in the sky you don't say "0, 1, 2".
No, you say, "I'm going to start counting stars now. Okay, 1, 2, ...".
The preparation part is the zero. You counted stars before and reached some number; you're not starting at that number.
If you just look at an empty space, the # of apples is equivalent is equivalent to the # of dinosaurs, but they're only equivalent by their absence.
I then put it in an initially empty bowl
The milliliters in a baker's cup are a scale based on 0!
nitpick: unless you were born on Feb 29.
The nice part is that now February 29 "just works" the downside is the impedance between how months and days are numbered and a how an offset from the epoch(beginning of the year) is defined. January 3rd(1-3) is stored as "0 months 2 days" as when it hits, 0 months have passed and 2 days have passed.
So the specific case of February 29 hits when 1 month has passed and 28 days have passed. 3 years out of 4 this will be the same as "2 months 0 days"(3-1) but every forth year this will be(2-28). As an aside, and the specific reason I went with interval types, every event past feburary 29 works just fine with or without a leapyear, that is, the extra day in the middle does not mess up the offset to days after it.
Honestly I curse a little as I wish months and days were 0-based. At least clocks get this right, almost, 12 hour clocks are a special breed of stupid. start at 12, then go to 1 and proceed up to 11. 24 hour clocks properly start at 0. The worst part about 12 hour clocks is that it is almost correct, replace the 12 with a 0 and it every thing be the same but now it makes sense from a moduler math point of view.
A special curse is reserved when I think about how there is no year zero. https://www.postgresql.org/docs/13/functions-datetime.html#F...
If that's the case, then how did the ancient Greeks or Romans count before zero was an acceptable concept in their counting system?
There is a narrowing from "nothing at all" to "zero apples" which doesn't happen 'in the world' but is a necessary precondition to counting apples. The existence of any apple is a requirement for there to be zero of them before you put anything in the basket.
Another definition might be from conception. But birth being “one year old” is illogical. The sperm didn’t even exist yet, one year before birth.
Your first birthday is the first anniversary of your birth day.
We celebrate the anniversaries of our birth day.
That is not the case with age or birthdays which are, thankfully, zero indexed. The first year of human life is age=0, birthdays=0.
0. https://en.wikipedia.org/wiki/Astronomical_year_numbering
1. https://docs.oracle.com/javase/8/docs/api/java/time/temporal...
So as far as your job is concerned, that's when they start. Hope that helps.
If you have 4 classes in school today, you would never talk about the 1st class as number 0, the last one is the 4th, not the 3rd.
I am assuming that the reader is biased towards the average HN user. That won't always work for everyone who will come across my comment, but close enough for an unpaid contribution. I'm not about to write a novel to make sure I catch every edge case.
Most human counting uses end-index. I.e. "1" is after 1-thing (has passed, is physically obtained, etc.).
Most computer counting uses start-index + length, for efficiency and to better generalize. I.e. "1" is at the memory address immediately before the "1"st item.
Which ultimately creates the "0 index is 1st thing" linguistic confusion.
PS: Also, language predates computers by a few years, and the concept of zero is hard.
i think a more accurate and complete idea is that humans refer to a thing in its entirety, with things being lined up and scanned in order as only a potential convenience
if you ask someone to identify an object, they'll point to the middle (or center of the most important component) of the object, not to the 'start' or 'end' of it in their field of vision..
that said, the human perspective is subtle and convenient in completely different ways to how a computer manages memory, and this 'end-index' idea seems like a useful way to map the human whole-object perspective to a linear memory-index perspective
The assumption is that the thing is its entirety, as you said.
Hypothetically, I imagine an array index reference, in human terms, would be communicated as "this thing starts here" or "this is the beginning of this thing."
Which isn't a concept or phrasing we have much occasion to use, other than for routes or long length-measured objects?
Indeed they are, which is why I literally said so in the previous comment. Did you forget to finish reading it?
> They do not include the event itself.
That is true because the event itself is implied information. There is no value in speaking of it. If you stand before us, we can be certain that you had a day of birth (your 0th anniversary). We don't necessarily know how many times you've gone around the sun following that, however, so that is where we find value in communicating additional information.
If you were counting apples, there is the state where you have no apples (index 0), the state where you have one apple (index 1), the state where you have two apples (index 2), etc. When counting you don't need to worry about index 0 because the no apple state is naturally implied. It only becomes interesting and worthy of communication when you have at least one apple to speak of, thus you start at 1. The state found at index 0 is still implicitly there, though.
0!=null
I say it has less weight as birthday is a compound word, the root words of which suggest that your first birthday could logically be the day of your birth rather than a year after it.
My first weddingday was not a year after I got married. My first graduationday was not a year after I graduated.
There's no culture where you're born at 1 year old and turn 2 on New Years when everyone gets older?
You sure?
So, if you're born on December 31st, you're two years old the next day. (I see no problem, but apparently some people are hung up on such minor details. I can't fathom why.)
That makes no sense to me. It's also your first century. Does that make you one century old? Of course not! The moment you're born, you're not even one hour old, let alone one day.
It could be though that Korea simply never encountered the Mayan or Arabic civilizations as most did encounter one of those two in history, who are famously known to have independently discovered the concept of zero.
Birth of an array can only logically be defined as index zero, the same as a child. The concept of zero was not universal globally and Korea is pretty isolated.
Americans and their alcohol laws…
In fact, this whole discussion is more about a choice of phrasing rather than the numbers. When indexing arrays, sometimes we're talking about an offset from the first element (starting with "0"), and sometimes we're talking about ordinal element numbers (starting with "first"). Some programming language designers found that offsets are more useful (because that choice tends to simplify the underlying arithmetic), while others found that ordinal numbers are more useful (because the word "first" should mean "1", to simplify communication between people).
On the other hand a computer/car/house can also be 10 years old but not 0.
Really? When I buy a new PC, to me it will get 1 year old only after a year. Before that it is just NEW. Is that what you meant?
That's why zero-based indexing is good!
But gregorian epoch itself is 1-based. 1AD (0001-01-01) goes right after 1BC (-0001-12-31). There was no 0000-mm-dd. That's why 3rd "millenium" and 21st century started at 2001-01-01 and not at 2000-01-01. YYYY means not how many whole years already passed, but which incomplete year goes right now. On the other hand, your age means "whole years passed since birth [plus maybe a few months]".
And the birthdays are definitely 1 indexed: you denote your first birthday with 1, and not with 0, and so on. You don't denote any birthdays of yours with 0. You may denote something else with it, but birthdays[0] gives an out of range exception. (Especially true in French, where birthdays are called anniversaire, but in English too.)
BTW on which level is the 1st floor?
in Europe or in the US?
In English you count "floors" and floor is a usable, hard surface on which you can put something, like a chair. That's why a floor on the ground level is treated the same as the floor above it - it is equally good on accommodating chairs, beds, and other stuff.
Etymologically, étage comes from the Greek στέγω (and gave the English word “stage”); it is a typically wooden cover. Since the first floor was often instead a continuation of the outside road (way back!), it was not considered a “stage”.
Again, this is more a spoken language/culture thing, and this goes back to what premises we use to communicate with the machines, ours or the machines...
No elevator I've seen has yet taken a cue from UI design: simply put the buttons within an outline of the building, along with the local numbering scheme.
No more visits to the serial killer lurking in the basement.
The players in this market evidently have been operating at T'ump levels of intelligence. /s
Maybe distances should be human-focused, so that no displacement is equivalently expressed as 1 cm, 1 m, 1 km, ...
Then converting a distance d from m to cm is (d - 1) x 100 + 1.
Just think of 0-based as offset-based and 1-based as index-based. Both are intuitive just like that. I never get why people arguing over this bring pointers and memory (or anything computer related) to the table. No normal person is going to understand that, but everyone understands that if you don't move at all (0 offset) you stay at the first (index 1) item. Add one and ... you get the point.
[ ] [ ] [ ] ... [ ]
0 1 2 3 n-1 n
We can regard this n as a "virtual zero", and then make it possible to index the n - 1 element also indexable as just -1. The index -n then aliases to 0.English is spoken as a native language in many countries outside of America.
99% of the instances I've seen "ordinal" outside of this thread has been in code/documentation. It is not a common word in everyday language.
If you move zero from the first element, you're still at the first element... but if you move zero from the fourth element you're still at the fourth element. If you move one from before the first element, you're at the first element.
I think you're more or less right, but I think we still need something to motivate the first element as the point of reference, and the machine focus is one way to do that.
Other languages are more sophisticated (like humans are!) and can say that position 0 does not exist. An array of size 0 has no elements. An array of size 5 has elements 1at through 5th. An array of size N has 1st through Nth, inclusive.
I suppose saying that does have an advantage over explicitly talking about pointers, in that the word "pointer" is a piece of jargon that has a lot of baggage. That's just avoiding jargon, though, not really using a different model.
The advantages in measuring from origin can accrue to the person choosing to do it, because there are other reasons to do so which aren't satisfying the CPU.
Calling that defined location "origin" doesn't suddenly make it not a pointer.
[0]- NUMA and cache effects aside, as they don't matter for this purpose
In Google standard SQL, to access an array it is simply not allowed to put a number inside square brackets. You must specify which way you mean. So `SELECT some_numbers[OFFSET(1)], some_numbers[ORDINAL(1)]` is allowed but not `SELECT some_numbers[1]`.
This type of verbosity makes sense in a big organization where the left hand doesn't talk much to the right hand much, so communication naturally evolves to happen at the code level.
I like for example using enum types to ensure that what’s actually passed is the expected value even though fundamentally the semantics do not need to be much more complicated than integer values. There could be the same thing with a distinction between offset and indices as two different integer numerical types to avoid any ambiguity.
That is false; counting begins by initializing an accumulator to zero. When you register the first item, the count jumps from 0 to 1.
Or someone who simply becomes motivated to count something, without making any utterances or gestures to that effect. The motivation is followed by the persistent awareness that nothing has been counted yet, which then changes to 1.
Or do you believe that the answer "none" or "zero" is then given without the activity of counting having taken place?
What do we call the meta-activity then: the procedure that results either in the "empty" answer or "one", "two"? Whatever that activity is called, it starts with a concept of zero. Let's call that activity "quanting". Quanting starts with a motivation to enumerate items, and an initially empty result. When no items are present, quanting terminates, reporting that zero/nothing/none result. Otherwise quanting branches into a subprocedure called counting, and that begins at 1.
A smart ass kid might reply there are zero oranges in the basket, or zero miniature unicorns. Since the basket is empty, it could have potentially had anything if we're just going to imagine things in baskets. But we don't enumerate over all possible zero items in the basket. And anyway, the basket isn't really empty since it has N air molecules, N fibers or whatever.
The pedantic point I'm making is that counting at zero is a convention we developed for mathematical reasoning when appropriate, but not a starting place for counting things in everyday language.
The floors of a building might start at 1, but you go up 1 flight of stairs from the 5th to the 6th floor, not two.
It is more like counting from 1 is just a leftover from times where zero was not commonly considered as number. Once one have zero, it makes sense to use it as an initial ordinal (see e.g. set theory, where zero is both initial ordinal and initial cardinal number, way before computers).
Another example is time and date, we start counting of days from 1, but counting of hours (at least in 24-hour notation) and minutes from 0.
> Luis started explaining why Lua was 1-index with this sentence: "The 1st argument ...."
Note that for spoken language, it is "The first argument ..." and 'first' is etymologically unrelated to 'one', but related to 'foremost', 'front', so it make sense to use 'first' for the initial item in the sequence even when using counting from 0.
This is of course different for continuous quantities. When counting seconds, for example, we should absolutely start from 0.
As a kid I learned "one one-thousand, two one-thousand, three one-thousand" when counting time out loud, but at some point I realized this was incorrect. The prefix is the start of the nth second but it isn't complete yet, so for example stopping in the middle of saying "two one-thousand" you actually haven't reached two seconds yet.
My fix was to move the prefix to the end, so I say "one-thousand one, one-thousand two, one-thousand three".
In retrospect maybe I should have used "zero one-thousand, one one-thousand, two one-thousand".
Doesnt sound half bad
The only way to avoid the speed penalty would be either to have a wasted element at offset 0, or to maintain the array base address as (address - (1 * element-size)) to avoid having to subtract 1 from the index when accessing. In the latter case for dynamically allocated arrays the code would still have to do a subtraction to adjust the pointer returned by the memory allocator, but at least that would be a 1-time penalty rather than per-element-access.
Of course this is supposing a high level language where an array is abstraction, not one explicity aliased to a chunk of memory such as C where an array and a pointer to it's first element are interchangeable.
Do you know the answer to that question? (I don't, but if someone does, it will settle this issue).
OTOH, for arrays who's contents are size 2^n (char, short, int, long) I'm sure the generated code isn't using multiply in the first place.
Anyways, an optimizing compiler could certainly remove much of any overhead added by 1-based indexing .. for an array access in a for loop it could, if necessary, calculate the "base-1" address once at start of loop.
Personally, having grown up with assembler and C, and still using C++ today, I'm quite happy with 0-based.
Also, the time jump from 11AM to 12PM to 1PM! So I think more human friendly sometimes means more confusing.
Other than that - binary AND and power of 2 sized arrays are the backbone of any hashtable. Overall modulus (binary AND) is actually useful.
Doing the reversal test, if programming languages had all been 1-indexed, then I doubt we'd hear much from people, in 2022, saying "I think the first element should be 0, and the second 1".
TBH humans would have been better of if we were 0-based, it's just a convention. And we have the confusing language where "20th century" means 1900's. If we wanted to bring the 1-indexing we use in language to the fullest extent here to fix that particular issue, time counting would have to start at 1111. Except that won't work once reaching 5-digit years.
If we would start with "zeroeth" instead of "1st", then this would have solved itself and 20th century would mean 20xx's.
to this day I can't seem to be able to explain to people that Jan 1 2000 was NOT the start of the new millennium, but rather Jan 1 2001
https://www.latimes.com/archives/la-xpm-2000-dec-26-mn-4810-...
TBH humans would have been better of if we were 0-based, it's just a convention. And we have the confusing language where "20th century" means 1900's. If we wanted to bring the 1-indexing we use in language to the fullest extent here to fix that particular issue, time counting would have to start at 1111. Except that won't work once reaching 5-digit years.
If we would start with "zeroeth" instead of "1st", then this would have solved itself and 20th century would mean 20xx's.
I especially don't understand why mathematicians use 1-based indexing (for matrix rows/columns etc...). Like programmers, they should see the advantages of starting at 0 (e.g. the coordinates of subdividing into block matrices are simpler if starting at 0). Mathematicians do start at 0 for the origin of plots, after all.
That's funny, but there's actually an elegant way to do it called Bijective numeration (https://en.wikipedia.org/wiki/Bijective_numeration). We're currently living in the 1A22th year.
If I ask you to count the number of red balls in a bag with only 3 yellow balls, then the initial count in your head is 0, you inspect the balls one by one, never encountering a red ball, and thus never incrementing the count. And then you pronounce your final count of 0. So that's counting starting from 0.
What you call "starting at 1" is not so much the start as it is the first increment, which need not arise.
The problem it that most people, even programmers, don't understand how computer arithmetic works and have fantasies of mathematical number lines that the hardware only partially simulates. You see this consistently in the post-Java crowd who think that unsigned integers are some sort of unholy aberration because the languages they've grown up went further to maintain the fictional number line semantics.
Personally, I've always found zero-based more intuitive and satisfying, absent arbitrary indexing (in which case I'd chose zero-based for most purposes where there wasn't a compelling solution based reason to use a different index base). But then spent years writing SPL ... I
But I don't think the fact that many of us developed that intuition, makes it inherently better. It just makes it more intuitive for those of us who developed that inuition. I've written oodles of codes with both paradigms. Either works. Zero is just more natural, to me.
I have Martin Richards book on my shelf, and for a while shared an office with his coworker. I'm also one of the few folks here who has coded in BCPL.
Ok, so that said, this article is obvious nonsense. Zero indexing comes from the CPU itself. It's how you generate an indirect address by adding the offset. Many CPUs have this in hardware. They don't support non-zero offsets. The reason some languages have zero-based indexes is that their developers had the mind set of doing the same thing you do when writing machine code. They probably also wanted to directly interface between their language and libraries written in assembler (and the kernel was typically written in assembler too). Also they wanted to be efficient. Any offset other than zero is not efficient because you can't use the hardware addressing modes.
So then why do some languages _not_ have zero based indices? The obvious answer it that Mathematics by convention always used 1 as the base (or some arbitrary value). Many languages were conceived by mathematicians or people with a strong mathematical background.
Both have inconsistencies, like in Spain you might have a "middle ground" (entresuelo) which is neither 0 nor 1, but sits between, and is normally commercial or non-livable, reserving the 1 to the first floor where people live. I like that system better though because it usually goes 1 => 0 => -1 (underground), while in Japan you go 2 => 1 => -1, so I feel like it's missing a floor. You could justify it though as 0 being the ground line, so +1 is "the first above the floor and -1 is "the first below the floor", but I still prefer having each floor to be a natural consecutive number.
In older buildings, the ground floor is usually "1" - which means the floor above has to be "2".
type Day is (Monday, Tuesday, Wednesday, Thursday, Friday, Saturday, Sunday);
type Hours is array (Day range <>) of Natural;
V : Hours (Monday .. Friday);
Which index type you should use depends on the problem that you're trying to model. You can find out the lower/higher end of a type/variable with the 'First and 'Last attributes.IIRC there's an RFC though to force the lower end of the index to a certain value like 1 or any other number/enum.
Turns out both were invented by the DoD.
>Due to the Department of Defense requiring as much of the syntax as possible to be based on Ada, in order to avoid re-inventing concepts that had already been thoroughly tested in the development of Ada,[citation needed] VHDL borrows heavily from the Ada programming language in both concept and syntax. - https://en.wikipedia.org/wiki/VHDL
Maybe I should pick up Ada soon. That could be a fun journey! (I really love writing VHDL)
The old way of doing it was setting the $[ variable to 1. You can actually set it to any value, in case you prefer to start your arrays at 42. It is now deprecated but you now have Array::Base that offers similar functionality.
Use it if you absolutely hate the person who will read your code, as if having someone read Perl code wasn't hateful enough.
> new Date()
Wed Aug 24 2022 ...
> new Date().getFullYear()
2022
> new Date().getMonth()
7
> new Date().getDate()
24
So 2022-08-24 comes out as (2022, 7, 24). One-based indexing for the year and day, zero-based indexing for the month. tm_mday: The day of the month, in the range 1 to 31.
tm_mon: The number of months since January, in the range 0 to 11.
I wonder where C got it? It goes back to at least 1973's V4: https://github.com/dspinellis/unix-history-repo/commit/92779...Convenience. Look at how the days in the months are stored and accessed. Using 1-based months would introduce an extra calculation (-1) on all searches or an unused value in the 0-index. Also look at how printing is handled for weekday and month names. They, again, take advantage of 0-based indexing.
Day and year are already represented as numbers, so it's natural to keep them as the "correct" (conventional) number as used by most people. Since the months aren't being stored as strings but as an index, this saves them from having useless data (entries in 0) or doing an extra calculation.
Or, if you're willing to give up three bytes, index into "ErrSunMonTueWedThuFriSat".
Or, and this is mildly insane but perhaps in keeping with early C, have your pointer be three bytes before the beginning of "SunMonTueWedThuFriSat" and use one-based indexing.
"When dealing with a sequence of length N, the elements of which we wish to distinguish by subscript, the next vexing question is what subscript value to assign to its starting element. Adhering to convention a) yields, when starting with subscript 1, the subscript range 1 ≤ i < N+1; starting with 0, however, gives the nicer range 0 ≤ i < N. So let us let our ordinals start at zero: an element's ordinal (subscript) equals the number of elements preceding it in the sequence. And the moral of the story is that we had better regard —after all those centuries!— zero as a most natural number."
https://www.cs.utexas.edu/users/EWD/transcriptions/EWD08xx/E...
e.g. len(range(3)) == len([0, 1, 2, 3]) == 4
1 <= a <= N -- N = length, no delta
0 <= a < N -- N = length, no delta
You only calculate the length when dealing with other than 0- or 1-based ranges. There, the inclusive-exclusive convention is very handy as you point out. So if we're fixing the initial offset at either 0 or 1, then use the appropriate convention for that offset. If we let the initial offset float, then the inclusive-exclusive makes sense.E.g. in Java, a typical example is that OutputStream has the following two methods, where the first can delegate to the second, and the second (which write the specified subrange of the array) can easily calculate the number of bytes to write:
int write(byte[] bytes)
{
return write(bytes, 0, bytes.length);
}
int write(byte[] bytes, int start, int end)
{
int length = end - start;
...
}> If we let the initial offset float, then the inclusive-exclusive makes sense.
But you replied to someone using inclusive-inclusive for 1-based arrays and complained about the delta not matching the length. Which is a nonsensical complaint, you have the length already why would you need to calculate anything?
I guess the majority of the people are not like this, but arguments like "just look at it how strange it looks" don't do it for me, because it looks natural, and the other way looks complicated.
If I am working with closed ranges, [i:i-1] looks like the list [i, i-1]. Like [5:2] would be [5, 4, 3, 2].
With [b+1:c] I would feel like I needed to insert a check to ensure b+1 <= c. With the closed ranges the invariant "left is <= right" is maintained automatically. Though I guess it doesn't matter so much if you accept any list with left > right as the empty list.
The issues with compositionality become even more noticeable with floats. Then you would need [a:b] + [b + minimum_float : c], or something like that.
1. Optimizing compilers exist.
2. Addressing at fixed offsets is cheap (a single instruction): https://en.wikipedia.org/wiki/Addressing_mode
0-based indexing is superior (as a default) because it simplifies a lot of common math, as discussed in TFA.
4. Macro expansion is often used to store indexes and other 'magic numbers' in Assembler, and this pattern originated either when Assembly was the primary programming language, or even earlier when humans directly punched out cards with machine instructions. Compilers in any remote sense of the luxury we have today did not exist or were not common.
char at(char matrix[10][10], char i, char j) {
return matrix[i][j];
}
still has more computation on 1-indexed than on 0-indexed, I believe.memory_address + array_width*i + j
or
memory_address + array_width*(i-1) + (j-1) = (memory_address-array_width-1)+ array_width*i + j
so you can just absorb the extra computation into the pointer.
If that was true, Fortran, Julia, R and Matlab would use 0 instead of 1-based.
I'd be surprised if the use of zero vs one actually made any difference, from a number of operations point of view -- from the compiler's point of view, the first element in the array is the first element in the array, no matter what we call it.
I mean in an extreme edge case maybe if you are computing an index, and it happens to be zero, then in your math maybe some identity related to zero could be exploited (go to element simple_integer*complicated_function(), where simple_number might be zero) but that seems a bit silly.
In x86 you could probably hide it in the addressing but that does not mean it does not to be computed
Same pitched notes are called "interval 1" and there goes thirds, fifths, sevenths... All off-by-one in my base-offset-addressing mind... And then major vs. minor which creates all kinds of "aliased addresses"...
I'd really love a BASE-12 floating number representation. Like 4.00 for the middle C; Chords can be then represented by a tuple of such numbers -- major = [+0.04, +0.07] (some sequencers already do something like that and I'm far better at reading that kind of sequencer data than a sheet)
I don't see how 0x0-0xf could be called base 17. 0xf is 15. Did you mean base 15? I think if mathematics terminology developed differently it could be called base 15. The same way binary is 0 and 1 but we call it base 2 because there are 2 digits, but we could totally call it base 1 too, who cares, it's all convention.
It's possible to have a speculative argument and also refute other arguments that claim certainty. Pointing out another answer as being wrong does not mean one needs to know the correct answer or claim to have a precise answer.
⸻
1. As did all the other contemporaneous BASICs that I encountered. For added fun, when creating an array with DIM, you gave the highest index and not the number of elements so, e.g., DIM A(20) created a 21-element array. The 1980 Apple ][ manual I found which discusses both Integer BASIC and AppleSoft doesn’t admit that 0 is a valid index for an array, but elsewhere I saw it indicated as such which leads me to suspect that Integer BASIC has 1-based arrays.
This was true of even TI-99/4A BASIC, which was closer to Dartmouth BASIC and not written by Microsoft.
fruits = ['apple'] \\ Hello I have fruits
x = len(fruits) \\ Only one kind tho lol
fruits[x] \\ OK here is what I have
SUBSCRIPT OUT OF RANGE
I have always hated zero indexing for this reason (other than thinking in assembler where it is beautiful). It is of course useful in many contexts, albeit with tradeoffs; more importantly everyone is used to it and it's predictable. But I was absolutely thrilled to discover Julia defaults to indexing from 1 like a normal person.Essentially what we have here is a mismatch between the tool (computer that does everything in binary) and the task (mathematical abstraction of quantities). Now it's completely understandable that people work within the limitations of the tool when there is no other choice, but that's the same sort of path dependency that creates technical debt.
If zero indexing were so great, mathematicians would have made it the default centuries if not millennia ago; but mathematicians don't want to do off-by-1 adjustments on all sorts of common operations because it makes things unnecessarily complicated. To be honest, I think this has become a moat to keep people out of programming even though it's a shallow one.
It could be interesting to test this, say by taking two classes of schoolchildren and teaching one Julia and the other Python (or...). By not having to take on the idea of zero-indexing at the same time as the concept of an array, the Julia group can get into collections using their intuitive understanding of the natural numbers. I expect that picking up language elements quickly will have a compounding effect and that at the end of the evaluation period the Julia group will be able to make significantly more complex programs than a control group.
Edit: in your heart you know I'm right
They did. For any context where the index contributes to the math[1] (i.e. is not simply a label convention like "the x (first) component") the index always start at zero. Polynomials, Fourier series, geometric series, Bessel functions (both their order and their series expansions), etc.
[1]: Except, of course, when a divergent 1/0 term shows up.
Here is a nice discussion on stackoverflow https://cseducators.stackexchange.com/questions/5023/why-do-...
The first answer nicely retells the dijkstra argument: integer ranges should be described using half open intervals, and [n, m) is nicer than (n, m]. Furthermore, [0, n) is nicer than [1, n+1), and that's that.
The second answer makes the observation that there is a difference between indexing and counting, and that even in daily life often the first element is indexed by 0.
Still, I rather like indexing the first element of a sequence with 1.
Yeah, I think this exactly why it is zero based. The index became a simple multiplier for sizeof(int).
Don't get me wrong, it's why I like C.
Maybe in the very early days, but that necessarily had to end as soon as compilers started doing non-trivial optimizations. Today, C is very far from being "portable assembly".
If your language doesn't support zero-indexing, then you are stuck with 1-indexing, and that's that.
So, zero-indexing is the natural thing to put in a language, as it accommodates everybody.
In the past, I've written roughly half a million lines of Lua. Arrays don't always start at 0. I've regularly shaken my fist at Lua and cursed it's wicked ways, but it's really damn useful in a lot of contexts.
Arrays in programming use offsets.
for(int i=0;i<10;i++)
{
array[i] = x;
}
The error here is calling it 'i' for index. It should be 'o' for offset.Not all do [1].
1. https://therenegadecoder.com/code/which-programming-language...
The point is that offsets and indexes have different meaning.
The symbol changed over time as positional notation (for which zero was crucial), made its way to the Babylonian empire and from there to India, via the Greeks (in whose own culture zero made a late and only occasional appearance; the Romans had no trace of it at all)
Another meaning is "nothing" but we're talking about position, not nothingness
The mathematical zero and the philosophical notion of nothingness are related but are not the same. Nothingness plays a central role very early on in Indian thought (there called sunya), and we find speculation in virtually all cosmogonical myths about what must have preceded the world's creation. So in the Bible's book of Genesis (1:2): "And the earth was without form, and void."
https://www.scientificamerican.com/article/what-is-the-origi...
a[z * h * w + y * w + x]
But with 1-based indexing you need to substract one first from each component except the last: a[(z - 1) * h * w + (y - 1) * w + x]You know what would make all this confusion go away? Zero-based indexing.
Another example is if I'm getting directions from someone, and they say, "Walk 3 blocks that way", in my mind I conceptualize that I'm presently on the 0th block and I don't yet count it.
I don't think every day people disagree with these examples so much (well, maybe the birthday one, but that's because we use the same word for the literal day as well as anniversary), and they're all fairly intuitive. I never realized that 0-based indexing required such a deep background to be justified.
That depends on whether you think of "birthday" as synonymous with "the actual day of my birth" or "a celebration of the day of my birth". I think most of us actual consider it the latter, which is why the first one is after you've been born for a year. If you look up the definition of "birthday" (one word, no spaces) you'll see that there is often some acknowledgement of this in the differing definitions offered.
> "Walk 3 blocks that way", in my mind I conceptualize that I'm presently on the 0th block and I don't yet count it.
You have to decide whether the person is saying something is on the third block from your position, or you need to walk about three blocks of distance. You probably do this automatically based on how close you are to the edge of the current block. If I'm 30 feet from the edge of a black, I'm probably not going to include the current block in that distance. I imagine you probably won't as well.
> they're all fairly intuitive. I never realized that 0-based indexing required such a deep background to be justified.
I think they're only intuitive because we're all context sensitive. The issue is when people don't have the same context, or the context no longer strictly makes sense when the terminology is used in contexts that make less sense.
As an offset, and in programming languages where it's easy to see it's an offset (C), it makes perfect sense. In languages where that's all hidden from you, and you're really just referencing the nth item in a list, the 0th item doesn't make a lot of sense. People in these different contexts will likely have different ideas of what "intuitive" means with regards to this.
Something like the following IIRC: Type IntArray = Array [-10...10] of Integer;
0 based index makes sense to me. The memory location is offset 0 bytes from where the first data byte is stored (without getting into endianism).
I started programming in FORTRAN, then COBOL, then (Lord help me) RPG II. All had 1 based arrays. I got paid to program in RPG II eight hours a day, five days a week, for two and a half years.
A common bit of code was to create an array of month names. Then take an input date, say 20220825, and split that into $Year (first four characters) and $MonthNumber (next two) and $Day (next two).
$MonthIndex = strip the leading zero (if present) from $MonthNumber
Then, this code worked: $MonthName = @Month[$MonthIndex]
When I ran into my first language with 0 based arrays, and @Month[1] returned February, I was pretty much convinced you all had gone insane. December is an out of bounds condition now? Really?
Much later, a friend explained the stuff about multiplying the index by the array element length, and easily getting the memory location of the item in that array. It does make sense; but, I've always hated the code I've had to write that says $MonthName = @Month[$MonthNumber - 1]
You can now ask, "why does everything start with 0"? I guess because it gives you nice round (base-2) numbers. With 4 bits you can have 16 different "numbers", either 0 to 15 or 1 to 16. However, representing the number "16" in binary requires 5 bits, so you need an additional bit for the same number of elements. So representing 4 bits as 0 to 15 makes more sense
Constructing the naturals without including 0 and starting with it is incredibly awkward. Addition doesn't even have an identity.
Look at a clock, is 1 at the top, or is it the next one after the top.
Look at a ruler. Does it start at 1?
Do you want to have to shuffle everything around when you go from whole numbers to halves? Do you start at 0.5 or 0? Depends if it is "halves" or the real number 0.5
This is a recipe for pain and off by ones any time you're dealing with time steps (which is why it's so stupid that matlab and fortran are 1 indexed by default, if anything there's a better argument for C to be 1 indexed than scientific languages).
Now index and compose a bunch of ranges. A half open range of two elements is incredibly stupid 1 <= i < 3.
So closed ranges go with 1 indexing. Composing and splitting them is incredibly awkward [a,c] is split into [a,b], [b+1, c] ... that's kind of okay, but not unique as you could use [a,b-1], [b,c]. Now do [d,e], [e,f] compose at a glance?
Half open ranges go with zero indexing and they just work.
Either one is something you get used to with practice. I do prefer 0-based, but 1-based isn't as big of a deal as many people treat it in the vast majority of situations, and almost every situation where you'd be using something like Lua in the first place.
0 is a number, and just because the west ignored it for quite some time (read: "Zero: The Biography of a Dangerous Idea") we have these odd built-in aversions to it but in the end it makes plenty of sense.
0th element is 0, so why don't we just do a better job of educating kids to start counting with 0 and thinking about 0.
in the end - as a CS person, i think it's good - but i also don't mind if you disagree with me. :)
That would fix so many of these discussions.
...Of course, the answer is offset, because offsets are beautiful, composable, elegant things that harmonize with bitmasks, modulo, array slices, and everything else, while (one-based) indexes are nothing but "offset plus one", no more useful than "offset plus two".
In fact, it's so fundamentally more elegant that I've come to the conclusion that the real issue is that it's actually the mathematicians and language itself that got indexing wrong. Instead of "first" being associated with 1, it should be associated with 0. We should give athletes 0th place, and talk about the 0th man on the moon etc.
That's where the cognitive dissonance for the 1-based-indexing people is coming from. They can't deal with the fact that their "normal" way of indexing is wrong.
It's a bit like how pi should really be tau (2*pi), or the electron should really be positive. We got it wrong, but we're stuck with it because it's too much of a hassle to change it. Fortunately computing got it right! But now there's a mismatch with everyday life where people are used to the wrong thing.
- closer to hardware
- makes index == element wise offset from the first element
- there are both algorithm which are (very slightly) easier to implement with 0 and 1 indexing, AFIK more of the "bread and butter" algorithm are easier with 0 then 1 (e.g. indexing of C-style arrays used all the time, heaps etc.)
- 0 is not special => less bounds checks in typed languages (if you have an unsigned int index you only have to check the upper bounds, with 1-index you also have to check for 0)
- 0 is not special => with signed integer you can have negative indices (from the end) in which case all signed integers are valid (but potential out of bounds), with 1 indexing you have a "gap" in the middle which for some (rare) use cases can be very painful to handle
So independent of what the history was. It's just much more pragmatic to use 0-indexing in many common use-cases. Especially given that 1-indexing produces more special cases for very common use-cases makes it sub-optimal for both usability and performance.
>// if one prisoner fails all are dead (1st try -> tries = 0)
https://gist.github.com/amalic/c107b5289d946b758418778a95fdc...
That being said, it seems to be a junior software engineer issue where the brain isn't wired yet for "0-based" indexing. I don't hear such complaints from more experienced developers who have simply learned to appreciate it because like in my example above it is more intuitive to me.
I have added "and eventually future older me" above because I don't do much coding these days, and my career trajectory looks like it's going to be even less in the future.
However I did not understand why they would have a range in the early days of programming. On one hand it would have been hard to make a language and its compilers and then you add arbitrary ranges?
When you index starting at 1, you either have to add/subtract 1 internally, which sometimes but not always can be optimized away by a smart JIT/compiler, or you have to "waste" the first element in memory, trading memory for performance.
Lua always subtracts 1, while LuaJIT instead went with the latter approach.
print("Finger {}'s motor has overheated".format(i+1))However, it can also be useful for many purposes to allow arbitrary ranges, so a programming language probably should allow that; if the index of the first element is not zero, then the base address will not be the address of the first element (if the first element index is positive, then the base address may actually point into a different array, or a different variable). Then, if you have the base address and add the index, you will have the proper address of that element, whether or not the first index number is zero.
Zero is not always the most useful starting index; sometimes other numbers (which may be positive or negative) are useful. But, I think in general, zero is better.
Excerpt:
By a general rule, in B the expression
*(V+i)
adds V and i, and refers to the i-th location after V.
Both BCPL and B each add special notation to sweeten such
array accesses; in B an equivalent expression is V[i]
C was designed to be as close to the machine as possible. Prior to C the only programming language for systems programming was assembly. If the compiler-generated code subtracted 1 from the index every time you indexed an array that would have been seen as not as efficient as assembly.It would be easy for languages to offer two array access syntaxes, one for 0 and one for 1 like array[i] for 0 index and array{i} for 1 index access, then the accessing code could use the right tool for the job. Unlike defining the indexing type as a property of an array you shouldn't have problems with accidentally using the wrong one at time of use.
... but to really know if it's a good idea you'd have to test it, and sadly very few concepts in programming are subjected to any kind of rigorous scientific study.
If you think of arrays just as a kind of abstract list, then beginning at 1 makes much more sense. No one who looks at a to-do list at home talks about the 0th thing on their list to do. But the spacial nature of arrays makes starting at 1 confusing. After I take 1 step, I am no longer at my starting point. And cycling through arrays in code for me has a very similar feeling to taking actual physical steps through a data structure.
Like other languages, it is a collection of contiguous elements that can be selected by indexing.
Unlike other languages:
- any discrete type can be used for indexing, not just integers
- bounds can be any value (doesn’t have to start at 0 or 1. First index could be 13)
- one consequence of the aforementioned is that arrays are types that map to the problem domain, rather than being tied to the computing domain
- and more…
I'm happy to hear any dissenting opinions if this is inaccurate.
What? Is this some kind of trolling challenge? If you don't care about math, just say so. Most people don't. But making up a nonexistent convention is just silly.
If your arrays start at 1, you have indexes.
"0-indexed" shouldn't be a thing. I like the screenshot quote that refers to this as either "1-origin" or "0-origin".
Here's a discussion from some years ago that starts with some of Guido's thoughts on array indexing and moves on to the QuickDraw coordinate plane and points and rectangles, including how grid lines and points are infinitely thin/small, but pixels occupy the space between the grid lines.
In a minecraft clone I made, each chunk consists of 16 subchunks so that I only need 12 bits to represent the block position in the shader. Technically, the block position ranges from INT32_MIN -> INT32_MAX, but since each subchunk only consists of 16x16x16 blocks I can get away with storing the position in 12 bits and then passing in the subchunk's world position once as a uniform.
Lo and behold, it's written by the same author.
"technical benefits to zero-based arrays" -> "why zero-based arrays are a strong convention"
Article mentions that Hoye says as much but is dismissive of it. I am not sure why he would be so quick to dismiss.
The origin of zero-based array indexing - https://news.ycombinator.com/item?id=6879478 - Dec 2013 (107 comments)
Zero-based arrays and the mythology of programming - https://news.ycombinator.com/item?id=6708409 - Nov 2013 (1 comment)
Citation Needed - https://news.ycombinator.com/item?id=6595521 - Oct 2013 (2 comments)
Threads about the EWD mentioned by jaapsen01:
Why numbering should start at zero (1982) - https://news.ycombinator.com/item?id=22162705 - Jan 2020 (220 comments)
Dijkstra's argument on why numbering should start at zero [pdf] - https://news.ycombinator.com/item?id=17850441 - Aug 2018 (1 comment)
Why numbering should start at zero (1982) - https://news.ycombinator.com/item?id=17765034 - Aug 2018 (63 comments)
Why numbering should start at zero (1982) - https://news.ycombinator.com/item?id=13186225 - Dec 2016 (216 comments)
Why numbering should start at zero (1982) - https://news.ycombinator.com/item?id=9761355 - June 2015 (47 comments)
Dijkstra: Why numbering should start at zero - https://news.ycombinator.com/item?id=777580 - Aug 2009 (71 comments)
Also these. Others? I'm a little surprised there aren't more.
Why do we count starting from zero? - https://news.ycombinator.com/item?id=17923391 - Sept 2018 (1 comment)
Why C Arrays Start at Zero: I Don't Know - https://news.ycombinator.com/item?id=11228267 - March 2016 (3 comments)
Why C Arrays Start at Zero: I Don't Know - https://news.ycombinator.com/item?id=11114704 - Feb 2016 (2 comments)
I'm sympathetic to its author's approach - we talk about indexing and it's a strangely fascinating topic but the history of it hasn't been all that well dug out. We do indeed tend to respond to articles like this by going "well obviously [x]-based is good because" - you can see that at work in this discussion here - but those are not necessarily the true historical reasons. But I think the author made a leap too far, the article was a bit too brittle, and it landed in a slightly too argumentative spot.
0-based or any arbitrary offset of your choice so that your data model can more closely map to what you need it to. And use eachindex instead of making assumptions and you can work with arrays using 1-based, 0-based, or arbitrary-based indexing.
VB is the only 1-indexed language I have used. VB is bad. Therefore 1-indexed languages must be bad.
I feel like I've also seen algorithms that work out better with 1 based indices, but not as many.
My manager thinks it was done to make the language easier to understand for people just getting into programming.
Lua's arrays are 1-indexed, and it is a terrific interpreted language! Array[0] is nil, makes sense. v5.1 has a solid JIT, and it's got a great 2D game engine called Löve2D.
Toddlers love this invariant.
If you don't understand this, you should not be writing computer software.
I've always considered the index and the power to which a base has been raised to be equivalent in many circumstances.
Thus starting at 0 makes sense.
The array indexer was syntactic sugar for pointer arithmetic.
1 is just as arbitrary as 2 or 79823 etc. Math wouldn't work without 0.
Thats good enough for me.
I've been annoyed recently by this obnoxious neckbeard behavior when they make things zero based for no reason (other than to conform to their supposedly existent philosophy), like Ratpoison and Screen selecting windows by a zero based index, which means you have to move your hand to the far right of the keyboard to select the first window with the 0 key, the far left to select the second window, with the 1 key, and one to the right to select the third window, with the 2 key.
I so happen to have been trying to work out a tangible explanation on paper of which system is better on trips a few weeks ago, and could not come up with one (and none are given in the article nor the top comments in this thread).
On a side note, the moment you say, "zeroeth", you can no longer be sure anyone knows what you're talking about because at any moment you may count in English or this pretentious UN*Xtard dialect of English. Seems like another pedagogical stumbling block held on to boomers and boomer-wannabes who just want to have their little counter counter counter culture or whatever.
BASIC and Pascal I am sure would have been considered a high-level language in their day and naturally would have arrays be 1-based.
C on the other hand....
You are confusing arrays with strings. Those used to start at 1, but since Unicode took over (~2007) strings are implemented as zero-based too (even though you can still think of them as 1-based when writing code).
Also in Delphi, when creating a cross-platform application, strings are treated as 0-based, with even multiple locations in documentation stressing this importance due to having enough differences between classic VCL and the new kid on the block (aka FireMonkey).
Especially the DBAD comment.
https://www.google.com/search?q=arrays+start+at+zero+meme&tb...
Sometimes you have a data structure called an "array", which means "a bunch of things that are the same size next to each other in memory". You can store the address of the whole array as the address of its first element, and offset into it by an offset; clearly, the offset of the first element is zero, but it's not the "zeroth" element.
And finally, sometimes you have cargo cult behavior by people who think that the way offsets into arrays work is because "numbering starts at zero in computers!" or some similarly wrong rationalization. This is when you get stupid things like (nth 0 list) in a lisp.
Because you predominantly think of indexing as the first definition as shown here[1], does not mean everyone else does. Be careful when using absolutes, especially when talking about languages.
A journey starts 0 metres from the origin.
If you count the number of _ you are along your journey and start with 1, then at the second metre you are at the 101st centimetre.
When someone says how many dragons have you slain? Just as you leave and haven't heard of any dragons yet, you don't say "I'm slaying my first dragon now"
Every construction of the natural numbers I have seen starts with 0.
The first hour of the day has passed when the clock strikes one. The hour preceding that is the same day. English calls it 12 because it was invented before we had a firm grasp of zero and english is inconsistent. 24 hour clocks call it 00.
1 based indexing can be correct, but 1 based ordered discrete sets map incredibly poorly to any ordered set with a different number of elements (such as the reals). Additionally dealing with ranges is less easy to make consistent.
0 based indexing is always correct.
However, when you examine the origin of the words, that is not the derivation. The derivation and path to English is quite complicated. You may feel a sense of righteous indignation or righteous repudiation on this topic. I would gently suggest deferring that feeling because I think the story is very interesting to follow.
That is not the end of the story though. The reason those pairs of words stuck around is almost certainly because it makes a consistent mental picture.
Bringing it back to the topic at hand - "Why we use 0" - because it works. There are many times when it has been evaluated, and at least to the people who design languages, it is more mentally appealing.
https://www.etymologynerd.com/blog/man-vs-woman
https://medium.com/interesting-histories/interesting-histori...
The linked post is not just about a technical decision, it's about a social history.
The conclusion of the linked post:
> "Lessons
Things can have more than one cause. Don’t trust easy explanations.
Always look for information that refutes your theory, not just information that supports it. Hoye stopped as soon as he had a satisfying answer and didn’t keep researching.
Don’t be a dick.
It is tragically easy to trick me into doing free research."