More succinctly, in pseudocode:
while(++digits[i] > '9')
digits[i--] = '0';
By starting the increment at the non-rightmost digit, it also allows for easy addition of powers of 10: 10, 100, 1000, etc.I found this example very illuminating
sub AH,AH ; clear AH
mov AL,'6' ; AL := 36H
add AL,'7' ; AL := 36H+37H = 6DH
aaa ; AX := 0103H
or AL,30H ; AL := 33H
because aaa instruction clears the high nibble of AL, it works with either unpacked BCD or with ASCII (but always outputs BCD).http://service.scs.carleton.ca/sivarama/asm_book_web/Instruc... (page 8)
But iiuc aaa didn't make it in the transation to 64bit processors.
Indeed, the Z80 itself, which people like us know because it was very common in 8bit home computers, is an enhancement of the Intel 8080, which itself is an ancestor of the current Intel CPUs.
Incidentally, Z80 CPUs are still manufactured 40+ years later ( https://en.wikipedia.org/wiki/Zilog_Z80 ) and benefit from modern C compilers like SDCC and software development environments like z88dk or the one I wrote for the Amstrad CPC: https://github.com/cpcitor/cpc-dev-tool-chain (written for Linux and tools like GNU make, wget, etc).
That said, writing C code that compiles to packed BCD and the use of DAA at assembly level does not seem very practical. If few numbers are manipulated, unpacked BCD is simpler (closer to string, less code to handle) both in assembly and in C.
slow path:
seq -f '%.1f' inf | pv > /dev/null
...[ 12MiB/s]
fast path: seq inf | pv > /dev/null
...[ 491MiB/s]The most common way to label rows is with sequential numbers.
2.508u 0.152s 0:02.66 99.6% 0+0k 0+0io 0pf+0w
and the best-of-3 time for "less -N filename > /dev/null" is: 2.568u 0.159s 0:02.73 99.2% 0+0k 0+0io 0pf+0w
That is, it doesn't seem like printing sequential is the limiting factor in performance.This is with "less 458", "Copyright (C) 1984-2012". I downloaded and compiled stock 487 and the best-of-3 times went up to 0:02.94 for both cases.
Checking the source code, it does not appear to use knowledge of the previous output index in order to save time. The relevant code is:
static int
iprint_linenum(num)
LINENUM num;
{
char buf[INT_STRLEN_BOUND(num)];
linenumtoa(num, buf);
putstr(buf);
return ((int) strlen(buf));
}
where #define TYPE_TO_A_FUNC(funcname, type) \
void funcname(num, buf) \
type num; \
char *buf; \
{ \
int neg = (num < 0); \
char tbuf[INT_STRLEN_BOUND(num)+2]; \
register char *s = tbuf + sizeof(tbuf); \
if (neg) num = -num; \
*--s = '\0'; \
do { \
*--s = (num % 10) + '0'; \
} while ((num /= 10) != 0); \
if (neg) *--s = '-'; \
strcpy(buf, s); \
}
TYPE_TO_A_FUNC(linenumtoa, LINENUM)My implicit argument is that there's there's no reason to believe that line counting adds anything more than trivial overhead to the less output, so there's no reason to even consider this optimization, much less getting to the point to make a pull request.
Sequential IDs permeate everything. Array indicies. Memory offsets. Line numbers, track numbers, page numbers, check numbers... the dates of your daily logs, your calendars... heck, enumeration attacks can be as simple as "what happens when I increment this number by exactly 1 when making this GET request?"