Why ['1', '7', '11'].map(parseInt) returns [1, NaN, 3] in JavaScript
medium.com
medium.com
- I know all that, have been programming for 15 years, yet I still would have done the mistake.
- Simple unit tests may very well not catch this bug the first time.
- The language design allows your brain to ignore the index parameter because JS accepts superfluous parameters, which is a terrible decision.
- map() is a mapping primitive. It's supposed to adapt a type to another type so that you can pass it to a monoid. JS breaks this convention.
- Even traditionally not functional languages know better than that and separate iteration from indexing. E.G: python map() just maps, and if you want numbering, you use explicitly enumerate(). Ruby has each() and each_with_index(), etc.
Just another of the numerous sucky things in JS. They are not big, but they accumulate very quickly and make it one of the worse language existing. Certainly the worst of all modern stack languages. It's a shame it has a monopoly on the most awesome platform in the world: the web.
In fact, it's such a big problem the most popular JS projects are all things to avoid coding in JS or workaround JS deficiencies: typescript, coffeescript, jsx, babel, webpack, lowdash ...
I have not done it with parseInt per se, but I have made this precise mistake at least twice with passing a function `f` with an optional second argument in `some_array.map(f)`, and it took a while to figure out what was happening each time.
Now that Javascript has proper iterables and iterators and generators, I have been enjoying using versions of `map` and `spreadmap` which take in a callback function and iterables and produce an iterator. Then I can explicitly use `enumerate` if I want it. https://observablehq.com/@jrus/itertools#map
Even with all of JavaScript's deficiencies, it's still far easier to just deal with them than to switch languages entirely.
This obviously doesn't explain every Javascript fan, but the tendency is strong enough to shift the dynamics of the community.
I can't say I relate to the religious fervor with which people defend programming languages. C++ is probably my favorite language to work in, but I absolutely understand where the hatred for it comes from. I readily admit that it's grotesque in many ways and my preference is likely for lack of sustained exposure to certain other modern languages.
So have you not used PHP, or are you counting it as not-modern? :P
Without that we would not be able to have variable length argument lists in the past.
Spread has existed in Python forever under the name of "splat operator", default values as well. Same with ruby.
That what I meant when I said "they accumulate". A bad design decision not only affect the user cognitive load and productivity, but it also cascades to the rest of the language and shapes it.
People came at Guido relentlessly for adding new features, debating his decisions, trying to change the philosophy and aesthetic of the language.
The guy stood up to them for 20 years, staying polite but resolute at doing things his, at the time controversial, way.
To me, that's even more impressive than being a good language designer.
The name “splat” is not commonly used in Python. That name comes from Ruby (or Perl?). In Python it is usually called “star” or similar.
Here you can see Steven D'Aprano, one of the dev of the Python stdlib, naming it "splat": https://mail.python.org/archives/list/python-ideas@python.or...
I use "splat" all the time myself.
On the other hand, itertools.starmap() is named this way because it does:
def starmap(func, iterable):
for e in iterable:
yield func(*iterable)
Things rarely have one name in computing. Hell I call curly brackets "mustaches" all the time. It's as hard as cache invalidation apparently.Obviously occasional people are going to pick up terminology from other communities, and the name “splat” has been gaining popularity recently (I had literally never heard that term before a few years ago, and have been writing Python code since 2002). I occasionally hear British expats in the USA calling elevators “lifts” or baby carriages “prams” or lines of people “queues”. Doesn’t mean those have been common American terms.
This was certainly not called the “splat operator” “forever” as claimed in the previous comment.
> I call curly brackets "mustaches" all the time
I have never heard anyone call these “mustaches”. The common terms are “curly brackets” and “braces”. Nobody is going to have any idea what you are talking about.
Mustache is a Ruby library from 2009. Here’s the first commit https://github.com/mustache/mustache/commit/6ee6bcf21d381554...
But more to the point, someone calling their template library “mustache” because a curly brace has a vaguely mustache-like shape doesn’t remotely imply that people regularly call curly braces “mustaches” or would have any idea what “mustache” meant in the context of someone pronouncing their computer code.
Using multiple parameters for those are so uncommon that it's easy not to realise that it's supported.
_If you know all of that_ it may not be surprising to you, but in most cases there is no need for the average developer to know it, making it very surprising when it crops up like this.
This may come off as arrogant but.. I'll proceed nonetheless. There is a need for the average developer to know the interfaces of the functions they're choosing to use. They're well documented, and the documentation is neither hard to find, nor difficult to read.
There may be an argument here for strongly typed languages -vs- weakly typed ones, where often some of the ambiguity around optional function parameters is eased, but most languages have optional function parameters, this is not in any way unique to JS. And parseInt and [].map are very common JS methods: this is not some obscure API not all devs would be unfamiliar with, these are built-ins.
But one of the safeguards here is that clean, intuitive code makes it possible for any engineer who's a little more thoughtful to catch little bugs like this. God knows I've done it dozens of times over the course of my career: while reading code for some other purpose, a block catches my eye as having something off about it, and I dig in and find a bug.
The problem with unintuitiveness,especially when it's baked into the language, is that you don't scrutinize every line of code you encounter with your full brain and attention. Without going on the hunt for "map and Javascript in general are horribly designed, parseInt has a rarely-used second param, root out likely bad uses", my brain on another task would skim right over a map-parseInt call like the above. It reads as if it's correct, and in any sane language it would be.
You might be a dev who isn't very familiar with Javascript, but you still need to fix it. Or you might be having an off day. Or you're writing a critical fix under a lot of pressure and simply forget about the issue. And I'll bet if this code was being reviewed, most devs would still overlook the issue.
There are many real world scenarios where it isn't so easy or clear cut. We're all human, and we make mistakes. In other fields, we try and reduce how easy it is to make these mistakes. Some things will always be dangerous like table saws. But a programming language?
Anyway, the "deal with it"/"get good" mentality just seems like a lack of empathy. Not to mention the arrogant in-group thing of "oh, you don't know <wart x> of my favourite but flawed language? you must not be a good dev". And the fact that fixing something like this could cumulatively save a huge amount of brainpower for current and future devs.
You can say the same thing to Pythonistas who are fussed by memory management, but the tradeoff of that intuitiveness is performance (which is being somewhat obviated by newer languages like Rust, which have their own challenges to learn). The actual task you're trying to do is more complex, so the knowledge required is too.
In Javascript's case, the thing you're becoming an expert in is the path-dependent history of dumb language decisions by people who never should have been let anywhere near a language. It's beyond me that people don't understand why this would bug some, at the very least because it was so easily avoidable.
ES5 deprecated that behaviour and it's been widespread enough that code which started in the IE9+ era now looks like the very old buggy code omitting the radix, so anyone getting started relatively recently quite reasonably may never have needed to learn about it:
https://developer.mozilla.org/en-US/docs/Web/JavaScript/Refe... https://eslint.org/docs/rules/radix
But TypeScript doesn't catch this specific bug because it's technically correct[1] from an API stand-point, as the second parameter to parseInt is a number and the second argument to map's callback is a number.
For as many situations as TypeScript genuinely helps you, there are just as many situations where it gives you false security, like in this case, and I'm starting to consider not using it anymore. [2]
[1] not actually "the best kind of correct" despite popular television quotes
[2] https://sdegutis.com/2019-06-20-considering-removing-typescr...
This is, of course, why famously machine shops simply put up signs saying “don't touch the whirring metal” rather than using safety guards and similar measures.
I don't disagree that it's useful to have the extra arguments for map but the part which is really deadly is ignoring extra arguments, even if it's obvious why backwards-compatibility made changing it unlikely. I really wish that JavaScript had implemented keyword-arguments when the benefits had become obvious because when features like map were added a decade or two later they could have been defined as passing arguments by keyword and causing an error when given a function which didn't accept those names.
Seems like you'd only want to use parseInt if you expect to need radix changes at some point, e.g. converting between hex strings, decimal values, and binary strings
['1', '7', '11'].map(Math.round)
// => [1, 7, 11]
[["00000001", 2], ["00000111", 2], ["0x0B", 16]].map(x => parseInt(...x))
// => [1, 7, 11] ['1', '7', '11'].map(x => +x) ['1', '7', '11'].map(Number)['1', '7', '11'].map(x => ~~x)
A bunch of confusion points in Javascript (including this one) have been well understood for a long time, best practices have been developed to deal with them, and mature tooling exists to easily enforce them.
Put another way: If you are unsure, go download eslint along with the plugins for your environment, and find/create a ruleset that turns on almost everything.
> You just have to know this
Hidden, silent, unintuitive behavior that you "just have to know" (and remember each time you might read or write the code) is absolutely strange and surprising in an engineering context. Javascript's deadly combination of unnecessary unintuitiveness (map's ludicrous API) and permissibility (ignoring extra arguments) strikes again.
The permissibility of guessing what malformed code means is actually somewhat defensible in Javascript's case, since the incentive ecosystem of the Web is complex and frontend parsing has long had a tradition of best-effort parsing instead of throwing an exception. But the absurd signature of map here is half the problem, and there's really no good explanation for that than yet another instance of "Javascript is a dumpster fire that should only be used when forced to". Similarly "easy-to-use" languages like Python can afford strong typing and rejecting extra args when not specified because they're not bound by the Web's permissibility. But Python also, for all its flaws, has adults making language decisions and does the sane thing with APIs like map without loss of expressiveness. Adding loop metadata like indices requires an explicit call to enumerate(), trading off an iota of verbosity for intuitiveness, which is one of the most important things in writing code that can remain productive and bug-free.
Surprise! You need to know the basics of how your language works.
The correct error is something along the lines of:
let numbers = input.iter().map(parseInt).collect::<Vec<_>>();
--> src/main.rs:10:32
|
4 | fn parseInt<T>(input: T, radix: u32) -> Option<i32> where T: AsRef<str> {
| ----------------------------------------------------------------------- takes 2 arguments
...
10 | let numbers = input.iter().map(parseInt).collect::<Vec<_>>();
| ^^^ expected function that takes 1 argument
This isn't rocket science, it's basic language design. From time immemorial engineers have been making mistakes as they write software, so compilers evolved not to pretend otherwise and hope for the best, but to help engineers catch them. Except for one. ['1', '7', '11'].enumerate().map(([val, idx]) => parseInt(val, idx))
I understand that this pattern of having extra "helper" arguments is pervasive in JS and might seem less exotic to the community, but I feel it is a problematic approach, as it is foot-gun prone.No.
> expected function that takes 1 argument
`map` expects a function that takes one, two, or three arguments.
0: Going by JS naming conventions anyway; I'd use somthing like map, mapi, no-that's-useless personally.
array.stride(2).map(|(x,y) ...|)
Both error checking and ergonomics are preserved.If you think it should be ok for all devs to believe that parseInt === parseDecimalInt then maybe all languages should be decimal-only. That isn't the case though.
Seems intuitive that if parseInt allows omitting the radix parameter, that it should default to whatever radix a standard integer primitive would default to.
But in this case, the radix is _not_ omitted. The map function passes the index of the current iteration as the second parameter to the function it is passed. It is kinda like this: ``` ['1','7','11'].map((item, index) => parseInt(item, index))
[9, 10, 11].sort()
[ 10, 11, 9 ][...].map(Number)
?
I would never have guessed that it was passing the index as a second parameter, I would have expected a compiler error.
> ['1','7','11'].map(console.log)
1 0 [ '1', '7', '11' ]
7 1 [ '1', '7', '11' ]
11 2 [ '1', '7', '11' ]
[ undefined, undefined, undefined ]
> parseInt(1,0)
1
> parseInt(7,1)
NaN
> parseInt(11,2)
3
The correct way is: > ['1','7','11'].map(x => parseInt(x))
[ 1, 7, 11 ]
same as: > ['1','7','11'].map(x => parseInt(x, 10))
[ 1, 7, 11 ]Using 0 is the source of many interesting bugs, like when someone puts in an IP address like so:
192.168.010.001
and the client tells you it can't connect (to 192.168.8.1).
(That behavior contrasts interestingly with IPv6 where the numbers [segments] are only supposed to ever be hexadecimal (Base-16) and leading zeroes are to be ignored.)
JS map doesn't behave like anybody else's map.
JS argument passing prefers doing the wrong thing to interrupting the programmer.
These two things conspire together and poor parseInt does its best with the resulting garble.
JS map behaves sanely - It adds extra params, but I've had them be useful every now and then, and I've never had them cause a problem.
Arguments behave... Like arguments behave in JS. Arguments have ALWAYS been variadic and accessible through the "arguments" variable within a function. That's not the "wrong" thing, it's just a thing. If you want named arguments - put them up top. Otherwise you'll get an array-like with everything passed. This is entirely consistent with the language.
Finally - These two things conspire together to do what exactly? This isn't a subtle bug where it works correctly 99% of the time and blows in prod late on a friday. This is an obvious error with even dead simple test cases that clearly show the dev has messed up.
Unit test your shit, or hell, just run it once or twice before using it and you're fine.
---
Basically - you should know how your std library works. That includes JS. I think this is pretty trivial. Worse, I actually find this behavior far more intuitive than something like ConfigureAwait(false) in C#, for example.
I much prefer languages like C#, C++ or TypeScript where the compiler warns me of such problems.
Though as I said in another comment, best-effort parsing was a decision that arose out of the Web ecosystem and its one of the few things about Javascript's language design that can't br blamed on incompetence.
"Variadic arguments" (really optional arguments with defaults/signature overloading, which is only arguably the same thing) existed long before JavaScript, and, indeed, the web. Optional arguments are also widely considered to be good/useful.
Honestly, most languages STILL won't catch this, because the typical pattern is to use method overloading to provide multiple signatures if default arguments aren't supported.
parseInt(in) parseInt(in, radix)
Any sane language having optional arguments is using keys for opt args. That's what common lisp and OCaml do.
In OCaml you would write:
let parse_int ?(radix = `Dec) string =...
val parse_int : ?radix:[`Bin | `Oct | `Dec | `Hex] -> string -> int
and call it like parse_int "42"
- 42
parse_int ~radix:`Bin "111"
- 7
List.mapi parse_int ["1"]
- Static error: this expression has type `string -> int`
but expected int -> 'a -> 'bNot saying one is better than the other, optional parameters can make for much less verbose code, I especially like parameter defaults introduced in ES6.
[1]: https://www.typescriptlang.org/play/#src=console.log(%5B'1'%...
parseInt is defined as accepting a string and an optional radix value, which is numeric. map is defined as providing the value and its index, which is also numeric. Would any of C#, C++, or TypeScript catch that without redefining either parseInt or map to require a more specific type, breaking compatibility with many millions of lines of code around the web?
C# would have a compile error with that map and parseInt definition because it can't coerce the types.
https://developer.mozilla.org/en-US/docs/Web/JavaScript/Refe...
https://developer.mozilla.org/en-US/docs/Web/JavaScript/Refe...
The thing which would actually catch this would be the mismatch in the number of arguments (modulo someone declaring that third argument as optional) or breaking compatibility to change one of them not to be a basic integer.
You don't have to know if you don't use implicit type-casting. Make it explicit and you won't really have to worry about it.
"Explicit is better than implicit" is part of the Zen of Python for a very good reason.
Correct. Unfortunately a lot of people use implicit casting a lot and have no idea that that's what they are doing.
FTFY.
I know exactly how map works: https://en.m.wikipedia.org/wiki/Map_(higher-order_function)
This is map:
> a higher-order function that applies a given function to each element of a functor, e.g. a list, returning a list of results in the same order.
Since that is not what Array.prototype.map does, it is not map, it is some similar thing that is misnamed as map. If I wanted the index and the array as arguments, I would ask for them.
This reminds me of "array set" in Tcl (a language many probably haven't used in a long time -- or ever). Unlike the normal "set" command in Tcl, which overwrites the full value of the variable, "array set" is actually a merge operation -- it merges in new keys and doesn't get rid of anything. I've seen experienced programmers use "array set" in a loop, leading to awful bugs. This too could be dismissed by saying that they "don't even know how array set works," but actually I place the blame on the language for making it so easy to misuse.
I will not be convinced that an implementation of map that gives different results when there are identity functions in the middle is doing the right thing.
The official docs for parseInt says this:
>> An integer between 2 and 36 that represents the radix (the base in mathematical numeral systems) of the string. Be careful — this does not default to 10. [1]
I just found it confusing whether the author meant the default value is 10, or if a falsy parameter (not undefined) turns out to be 10.
[1] https://developer.mozilla.org/en-US/docs/Web/JavaScript/Refe...
parseInt("0x10") == 16 parseInt("0x10", 10) == 0
To clarify for others (because this probably sounds crazy): The default is implicit like the rest of JavaScript, it will change base depending on presence of the prefixes 0x (16) and 00 (8), it doesn't seem to have included 0b yet. The confusing bit was octal because as you can imagine some sources might have base 10 padded with zeros, ES5 basically removes implicit octals in parseInt to avoid this issue.
Arguably this is more of a problem with the ambiguous octal prefix than the concept of using prefixes to determine base.
parseInt('1', 1) // same error for '0' or '|'
> NaN
The result is intriguing still as "NaN" seems to indicate a invalid input in the first parameter instead of a invalid second parameter.I've found that apparently there is no consensus [1] on Base-1 notation or parsing, although my primary intuition is correct [2] in that a parser could be written that would parse a "1" as a 1 base10, "11" as a 2 base10, "111" and so on. The parser would probably look a lot like a simple length() function, but that could vary with certain base-1 encodings like the ones used by the Golomb Rice compression algorithms, which have each string end in "0" (unary coding).
[1] https://math.stackexchange.com/questions/371972/what-would-b...
['1', '7', '11'].map(numStr => parseInt(numStr));
I think you'd learn something much more useful with function radixParser(radix) { return numStr => parseInt(numStr, radix); }
['1', '7', '11'].map(radixParser());
> [ 1, 7, 11 ]
['1', '7', '11'].map(radixParser(8));
> [ 1, 7, 9 ]
['1', '7', '11'].map(radixParser(2));
> [ 1, NaN, 3 ] const radixParser = (radix) => (numStr) => parseInt(numStr, radis); function foo() { ... }
is a clearer way of indicating "I'm defining a function named foo". I tend to use arrows only for anonymous methods.Thanks Medium! Great readability.
Not too weird IMO.
That's hilarious. Everybody loves the syntactic sugar that makes things easy, until the unexpected point where it makes things very hard.
Using directly like this a function with map() is just incorrect, the correct way to do it is:
['1', '7', '11'].map(x => parseInt(x))
edit I'm getting downvoted, it doesn't matter much but I don't understand it when the most upvoted comment seems to say more or less the same?More generally, general purpose languages don't provide safety guarantees in the problem domain. They can provide certain guarantees in the solution domain, but it's left to the programmer to compose the elements of the language into a correct solution.
(In my experience, by far the biggest barrier to a correcty solution in the problem domain is that no one actually knows what that is, much less has expressed it. Instead, people express certain specific behaviors they think the system should have, from which the programmer needs to extrapolate the actual requirements -- not straight-forward since to a greater or lessor degree the expressions will be vague, self-contradictory, self-defeating, and/or incoherent. Then they need to compose those requirements in the solution space. BTW, the language is just a part of the solution space and "safe" languages are usually only referring to static checks, which the solution space often consists of distributed components which aren't strictly controlled in lock-step by the same static sources, meaning the static guarantees are useful, but in a limited way.)
[5 * 5] * 2 // 50
[5 + 5] * 2 // NaN
1 + [5 * 5] * 3 // 76
1 + [5 * 5] - 1 // 124
Interestingly, even Typescript will not ( by default ) catch this class of bugs.Even languages that do have different versions of it on different objects (e.g. Scala's zipWithIndex)
You can crank up the settings in TsLint and never worry about things like this again
https://www.typescriptlang.org/play/index.html#src=alert(%5B...
TSLint similarly reports nothing even with the tslint:all ruleset:
https://palantir.github.io/tslint-playground/?saved=N4Igxg9g... https://palantir.github.io/tslint-playground/