The secret life of NaN
anniecherkaev.com
anniecherkaev.com
> the 9007199254740990 (that is, 253-2) distinct “Not-a-Number” values of the IEEE Standard are represented in ECMAScript as a single special NaN value. (Note that the NaN value is produced by the program expression NaN.) In some implementations, external code might be able to detect a difference between various Not-a-Number values, but such behaviour is implementation-dependent; to ECMAScript code, all NaN values are indistinguishable from each other.
> The bit pattern that might be observed in an ArrayBuffer (see 24.1) or a SharedArrayBuffer (see 24.2) after a Number value has been stored into it is not necessarily the same as the internal representation of that Number value used by the ECMAScript implementation.
[0] https://www.ecma-international.org/ecma-262/8.0/index.html#s...
Playing around with it in Chrome 65.0.*...
const buffer = new ArrayBuffer(12);
const uint8s = new Uint8Array(buffer);
const floats = new Float32Array(buffer);
floats[0] = NaN, floats[1] = NaN; //127.192.0.0, 127.192.0.0, 0.0.0.0
uint8s[0] = 104, uint8s[1] = 105; //127.192.104.105, 127.192.0.0, 0.0.0.0
floats[2] = floats[0] + floats[1]; //127.192.104.105, 127.192.0.0, 127.192.104.105
console.log(String.fromCharCode(uint8s[8]) + String.fromCharCode(uint8s[9]));Of course NaN-boxing is the reason that the ECMA spec specifies that all NaNs are treated equally. If you are encoding all of your values in NaN, you still have to have one NaN value that is actually NaN. The spec doesn't specify which one that must be.
BTW, I highly recommend reading TFA and the links within, because it's really fascinating (the "tagged-pointers" used in V8 and other languages like Guile are also really interesting).
0 - https://stackoverflow.com/questions/43129365/javas-math-rint... 1 - http://webassembly.org/docs/rationale/#nan-bit-pattern-nonde...
I have seen data passed through the NaN payload in C in a signal processing application. It was both vile and genius at the same time. Vile because of how hackish it was, genius because it avoided a larger redesign of the application.
Just do:
void some_function(float number) { if(number != number) number was passed in as <NaN> }
Took me a little bit to get my head around it but if the number is not equal to the number, then a problem occurred. I used this to stop camera code (driven by floats) from crashing when receiving non-sense input coordinates. It still doesn't work properly, but it doesn't crash now either! (at least not for that :)
What about a comment that notes what part of a spec some code implements (i.e. something outside the actual behavior of the code)? A comment that answers "why" can be helpful, and can sometimes be worth the high cost that an unchecked, unexecuted part of the program inherently carries.
If Fowler's claiming that all comments are bugs, I'd call that damaging.
But when you can make the code clear enough to need no commenting, you should do that instead.
// ISO/IEC TR 18037 S5.3 (amending C99 6.7.3): "A function type shall not be
// qualified by an address-space qualifier."
if (Type->isFunctionType()) {
S.Diag(Attr.getLoc(), diag::err_attribute_address_function_type);
Attr.setInvalid();
return;
}
The comment justifies the code in a way that the code itself never could. if (isTypeFunction) {
showErrAttrDialog(Attr);
setInvalid(Attr);
}Later, if a bug is filed saying that the compiler isn't compliant with the requirements of "ISO/IEC TR 18037 S5.3", how can you be sure to find the code where the behavior is implemented?
And of course the function in question only deals with "ISO/IEC TR 18037 S5.3" so it's easily tested.
If someone files a bug saying the compiler isn't compliant with the requirements of "ISO/IEC TR 18037 S5.3", then they would provide a test case showing not compatible. Add that case to the existing unit tests and you will see which function fails. No need for searching the code to see where the behaviour is implemented. With clean code it's obvious, the test will show this method to be at fault. Even without any comments.
Code in functions should also try and stay at the same level of abstraction, moving low level stuff to abstracted methods that describe the intention (And then the low level method does the how). That way your code reads like a story.
So...it sounds like you'd have one broken implementation of the code, and one working implementation, that might just cancel out the effects of the broken one. You're assuming code that has been cleanly written for its whole history, or a lot of love poured into it to develop quality tests for each requirement. How commonly does that actually happen?
Clear code provides a clear "how". Good test coverage can act as documentation of the proper behavior and help prevent regressions. But I don't see how it follows that clean code makes the location of each implemented feature obvious. It doesn't seem like it would be inherently true.
[0]: https://blog.codinghorror.com/code-tells-you-how-comments-te...
/begin rant I'm working on a C++ code base developed by contractors that were lazy and doing things the expedient way instead of the correct way (thousands of circular dependencies between libs - and even apps - apps depending on source from other apps), multiple copies basic functionality with slight changes (largely bug fixed in one place, but they forgot the other X places it was copied).
I've gone so far as to write a "cleanup" script in python that runs a number of transformations on the source. Fixing things like inconsistent line-endings (via dos2unix), inconsistent formatting (via clang-formatting), limited conversions of certain Boost uses to std C++, a few pervasive spelling errors (contracts were eastern European, non-native English speakers). Another thing I'm considering is removing all comments. Lots of dead, commented out, code. Also, invalid UTF-8 characters in comments, from an ANSI code page I've not been able to identify (which of course creates problems with the python script treating the source as UTF-8).
The comments that are currently present that I've seen are: 1. stupid/moronic/redundant (like merely marking constructors as ctor and destructors as dctor with no more info - like I can't just figure that out by reading the source to begin with) 2. Just plain wrong. As-in, comment no longer matches the code. 3. In a foreign language (Ukrainian/Russian), in an unknown code-page, so also worthless at face value (yeah, I could lookup the code pages for Ukraine and Russia have give it a shot and run the result through a translator), but it's likely the result will also fall under #2. 4. Unnecessary demarcation between functions. e.g. "// -------" out to ~80 characters wide between functions with no other white-space. Want things to be clearer? Don't use K&R notation and but open & close braces on the same indentation level (e.g. don't put the opening brace at the end of your loop/function/if/else/other statement, but put them alone on the next line). 5. Dead code. Don't commit commented-out code. Delete it. That's what version control is for. Want to know what it used to do? Look at the history, not the chain of commented-out code. Commenting out code is fine for quick testing. But don't commit that. People that follow you will see that and wonder "why is this here? is this significant?" 6. Random BOM (byte-order-marks) for UTF-8 source that is actually entirely ASCII. A lot of linux tools don't expect or handle BOM in UTF-8 sources (I'm looking at you, psql).
In this particular project I'm working on at my employer, I don't think I've seen a single useful comment in the source. Sadly, the Jira tasks filed by the now-fired contractors are also nearly universally useless. All headline and mostly zero description on what the problem is, usually no steps to reproduce.
/rant
Sorry for the rant, when I started writing this, I did not intend it to be so. But, as I wrote, I started remembering more and more things that have been driving me nuts.
We had an iron clad contract with that contracting firm, but after I started reverting every single subsequent change they made, they stopped sending changes ;-) Luckily my boss was a director and could pave the political fallout from my rather brash (but justified) actions.
I remember Solaris had a problem because it would put anonymous mmap()s in one 48-bit range and the heap and stacks in the other range, and this broke some ECMAScript implementation that used NaN-coding with the assumption that all data would be in the top 48-bit address space.
Which is precisely why the implemented in that way, to prevent "clever" uses like this from eventually breaking on chipsets that define more address bits than the early default.
https://groups.google.com/a/groups.riscv.org/forum/#!topic/i...
The C standard does not define and/or require this. What you link and refer to is the POSIX standard.
That is genius! It'd be a little weird using a system where the most positive fixnum is 51 bits, but it wouldn't be terrible — and of course that's why bignums exist. And 50 or 49 bits are certainly more than large enough for realistic RAM sizes for quite awhile — the latter is roughly half a petabyte.
It seems from the article that it's a pretty common technique; I'm surprised that I've not heard of it before.
How can you check if a variable is NaN then? Well, if it's not equal to itself then it must be NaN!
For completeness: Linux supports 57-bit address space on x86-64: https://lwn.net/Articles/117749/
I don't know if there's any real hardware that supports this, though.
In my youth, I spent days wondering how in the world numbers in animation scripts are not "NaN" strings the moment something causes zero division, but are still numbers called NaN.
Still, `5 + 2 * NaN == 5 + 2 * NaN` would return false, because it becomes `NaN == NaN`, and NaN is not equal to NaN.
NaN will come up when doing
Inf - Inf
But clearly: 0 * (Inf - Inf)
cannot be said to be zero.NaN really is not a number.
func sum(a, b float64) interface{} { return a+b }
It will allocate memory for the float, because an interface type always contains two pointers. [1]
It's pretty crazy that the second word in a Go interface has to be a pointer. But I suppose if dynamic types were used more, they would be optimized more.
[1] https://www.darkcoding.net/software/go-the-price-of-interfac...
If it may not be a pointer you have to branch, which can be add some overhead to the lucky path (every pointer deref will have to wait for type comparison in machine code)
If the float has just been written, the data should be in cache so the performance penalty is probably not too bad in most cases.
In principle, using NaN boxing or a similar technique to store non-floats, you should be able to store floats in an []interface{} as efficiently as in a []float64. Scripting languages do this, but not Go.
https://github.com/mruby/mruby/blob/d6cb4f9cf2027eb20f67238a...