Stop using 'short' for line and allocation sizes (2013)
git.kernel.org
git.kernel.org
I guess there's one constant with humans and infrastructure - people will always find ways to hit its limits.
Typing in machine code programmes for the ZX81, and ZX Spectrum consisted of REM statements with 1000s of zeros, which then you had to poke the machine code from data statements. Converting magic numbers, to magic.
The original code:
struct line *b_dotp; /* Link to "." struct line structure */
short b_doto; /* Offset of "." in above struct line */
struct line *b_markp; /* The same as the above two, */
short b_marko; /* but for the "mark" */
struct line *b_linep; /* Link to the header struct line */
The new code: struct line *b_dotp; /* Link to "." struct line structure */
struct line *b_markp; /* The same as the above two, */
struct line *b_linep; /* Link to the header struct line */
int b_doto; /* Offset of "." in above struct line */
int b_marko; /* but for the "mark" */
int b_mode; /* editor mode of this buffer */If a pointer is 8 bytes and an int 4, the old structure typically (this is compiler dependent) would be 40 bytes, the new one 36.
But yes, that should have been two commits, and he should have updated the comments.
This also is somewhat of an argument for giving the compiler the freedom to choose filed order, as the order that is optimal for memory usage isn't the best for human understanding.
Ah: I actually did work out the alignment math, as I considered that as a reason to reorder the fields, but I was hoping that the goal would have been only to do that if the change was making the struct even larger (which it doesn't) and failed to think about a 64-bit computer :(. However, I personally don't consider "save four bytes per window" to be a particularly compelling reason to rearrange fields (which seem to have been put in that order more for semantic purposes than compression), and yeah: the real issue is that this is all in one patch and scrambled the comments :(. It is essentially at best a an unrelated optimization that seemingly didn't even notice how it broke the comments. (BTW: an issue with having a compiler decide the field order is that you then make separate compilation and sharing structures between libraries really really hard.)
Also, since Linus worked on kernel code a lot, that reordering probably was like a knee reflex; it didn't involve his brain. His brain wrote that we shouldn't be bothered by this in an editor, but his spine made the edit, anyways, and, apparently, his spine doesn't read comments.
off_t for file offsets
ptrdiff_t for memory offsets (vs. size_t for unsigned)
int for return flags
A problem with this is that printf does not have sizes to match these.
Silently increasing the size of an int is more likely to break things than it'll be helpful. If you really want an int to be 32 bit on a 32 bit machine and 64 bit on a 64 bit machine, make it explicit (e.g. intptr_t for pointers)
That said, the code bases that target both 64- and 8-bit architectures probably aren't that many and both our points are insignificant in those other 99.9% of cases.
For size_t: "%zu"
There's no format for off_t, though. The best you can do is to typecast it to intmax_t, and then use "%jd".
I doubt either of these situations come up when dealing with line sizes in a text editor.
Neither of these situations is happening in a text editor application though.
A uint that actually threw an exception or something if you tried to underflow it would be useful, but most unsigned ints nowadays aren't all that useful.
And it will be a bit faster for certain weird architectures since the overflow behavior of an unsigned is prescribed but signed overflow is undefined. But I really doubt anybody cares about that here.
This is why I like languages like Lisp: by default, integers are only limited by the amount of memory available (as an example: (- (expt 2 1024) (expt 3 27)) → 179769313486231590772930519078902473361797697894230657273430081157732675805500963132708477322407536021120113879871393357658789768814416622492847430639474124377767893424865485276302219601246094119453082952085005768838150682342462881473913110540827237163350510684586298239947245938479716304835356321998626652229).
For efficiency one can of course limit things to fixnums, and in a kernel one would want to be careful not to let reading a file eat all memory — but that's all doable in a Lisp.
For structs where there are thousands, or millions, in memory using the smallest type required to get the job done will improve performance because you will fit more data into the cache.
In any case, these particular structs don't have any special packing applied.
Edit: apparently it has no undo. The rules:
1) Linus does not make mistakes
2) In the event of Linus making a mistake, see rule 1
3) ...
4) git reset --hard HEAD
"Linus has been polishing a turd for two years"