So you think you know what a number is
chris.improbable.org
chris.improbable.org
Decimal characters include digit characters, and all characters that can be used to form decimal-radix numbers, e.g. U+0660, ARABIC-INDIC DIGIT ZERO.
Rigid designators vs definite descriptions for numerals, anyone?
>>> int("۲۶۷۹")
2679
Thats pretty cool. $ ruby -v
ruby 2.2.6p396 (2016-11-15 revision 56800) [x86_64-linux-gnu]
$ ruby -e "puts \"۲۶۷۹\".to_i"
0
$ python2.7 -c "print(int(\"۲۶۷۹\"))"
Traceback (most recent call last):
File "<string>", line 1, in <module>
ValueError: invalid literal for int() with base 10: '\xdb\xb2\xdb\xb6\xdb\xb7\xdb\xb9'
$ python3.5 -c "print(int(\"۲۶۷۹\"))"
2679
$ node -v
v6.9.1
$ node -p "parseInt(\"۲۶۷۹\")"
NaN Python 2.7.12 (default, Dec 14 2016, 13:32:53)
[GCC 4.9.3] on linux2
Type "help", "copyright", "credits" or "license" for more information.
>>> print(int(u'۲۶۷۹'))
2679 $ python2.7 -c "print(int(u'۲۶۷۹'))"
Traceback (most recent call last):
File "<string>", line 1, in <module>
ValueError: invalid literal for int() with base 10: '\xdb\xb2\xdb\xb6\xdb\xb7\xdb\xb9'
$ python2.7
Python 2.7.13 (default, Jan 03 2017, 17:41:54) [GCC] on linux2
Type "help", "copyright", "credits" or "license" for more information.
>>> print(int(u'۲۶۷۹'))
2679
Looking at what the string is: $ python2.7 -c "print(list(u'۲۶۷۹'))"
[u'\xdb', u'\xb2', u'\xdb', u'\xb6', u'\xdb', u'\xb7', u'\xdb', u'\xb9']
$ python2.7
Python 2.7.13 (default, Jan 03 2017, 17:41:54) [GCC] on linux2
Type "help", "copyright", "credits" or "license" for more information.
>>> print(list(u'۲۶۷۹'))
[u'\u06f2', u'\u06f6', u'\u06f7', u'\u06f9'] main()
…
setlocale(LC_ALL, "")
…
argv_copy[i] = Py_DecodeLocale(argv[i], NULL)
…
mbstowcs() or mbrtowc()
…
setlocale(LC_ALL, oldloc)
…
Py_Main(argc, argv_copy)
While the REPL’s encoding (sys.stdin.encoding) is set to UTF-8 due to LANG/LC_CTYPE settings. You can get the same error when invoking the REPL as: LANG="en_US.iso8859-1" python2.7
So the shell isn’t doing anything to the text—it’s providing UTF-8 bytes in both cases, it’s just that Python is interpreting them differently. isdecimal isdigit isnumeric
12345 True True True
១2߃໔5 True True True
①²³🄅₅ False True True
⑩⒓ False False True
Five False False False
Use isdecimal if you want to call `int` (though it's EAFTP). int("一万三千二百六十九")
Traceback (most recent call last):
File "python", line 1, in <module>
ValueError: invalid literal for int() with base 10: '一万三千二百六十九'
int("一三二六九")
Traceback (most recent call last):
File "python", line 1, in <module>
ValueError: invalid literal for int() with base 10: '一三二六九'
Well, that was disappointing. For the record, my site http://ichi.moe/ can handle both (and Arabic numerals too).Apparently the Unicode consortium chose to exclude those characters because they weren't encoded in a contiguous sequence so they're defined as Numeric_Type=Digit rather than Numeric_Type=Decimal. They've apparently realized that this is not helpful but chose to apply the updated policy only to new ranges.
https://github.com/navarr/JapaneseNumerals/blob/master/Japan...
I haven't yet found a security vulnerability based in this, but I keep checking. :-)
… and, yes, that was definitely something I tried to find when I first noticed it.