In Python, `[0xfor x in (1, 2, 3)]` returns `[15]`
twitter.com
twitter.com
If you neglect the outer list that's only there to fool you into thinking this has something to do with list comprehensions, same as the x, and if you don't use hex to make it seem like there is a "for" in the middle, it boils down to something like: "15or whatever()" which doesn't seem all that confusing, even if "whatever" uses an uninitialized variable like x, or is undefined, because we're in Python and it's only evaluated when it runs. Then we are left with the 'confusion' boiling down to a Wat about why it's legal to do "15or 3" without space before or, and especially why "x0for 3" works. This is documented as other comments mention and is due to the way the parsing works.
I don't know if the parser could be changed to require a leading space before the or operator, but it's pretty clear to me that this is only confusing if you intentionally do you very best to try to add confusing structure around it in an attempt to fool the reader into thinking something very different and weird is going on.
This has come up from a Core Developer before [1] so it's not just code golfers having a laugh.
Yes, the docs note the tokenization behavior [2], but Guido's response today in the above mail thread is also pretty unambiguous:
> I would totally make that a SyntaxError, and backwards compatibility be damned.
1: https://mail.python.org/archives/list/python-dev@python.org/...
2: https://docs.python.org/3/reference/lexical_analysis.html#wh...
Sure, but that's half the fun :-) I frequently see Ned's name pop up with interesting things like this.
[15or x in (1, 2, 3)]
As another person pointed out, Python's lexer sees 0xfor (or 15or) and splits it into "0xf", "or" ("15", "or" for the other case) then parser processes it as usual. int x = 10;
while ( x --> 0)
[0] https://stackoverflow.com/questions/1642028/what-is-the-oper...https://docs.python.org/3/reference/lexical_analysis.html#wh...
Also Python: Meh, you can leave out the whitespace, I'll figure it out
I've never encountered this issue before.
The reason I've never encountered this issue, or even needed to know it was an issue is because I use good development tools. 1) Pycharm highlighted the or in orange, making it clear that it was being interpreted as a keyword, 2) I use a linter (as we all should) which explicitly highlights the lack of whitespace around the token as an issue (PEP 8: E225), and 3) I use a code formatter (which we all should) which, again, highlights this statement as requests that I fix it.
Python has a lot of flaws. This one is down near the bottom of ones that are even really worth talking about.
You're taking my statement way too personally and out of context. You and everyone else can express your opinions all you want.
As a topic in the list of "flaws with Python," I don't think it's a very important or insightful topic because it's not something people really run into. This just doesn't fit there.
This is a novelty, pure and simple. Filed under "weird programming hacks you won't expect" This would be fine. It's like, `([]+![])[+!+[]+!+[]+!+[]][([]+{})]` in javascript. Weird, kind of interesting, but it's not really a "flaw," at least not the sense of something that will burn you unexpectedly.
Reminds me of Fortran.
When we were in college learning FORTRAN, students would ask for help from the teaching assistants at the computer center. One big problem was FORTRAN allowed horrible spaghetti code because GOTO statements could be used anywhere. It was easy to jump into and out from loops. The TAs had a tough job.
It didn't help when people would deliberately mess with the TAs by asking about code with language features similar to those in the article you linked. Something like IIRC after 47 years:
DO 15 I = (1, 100)
purported loop stuff goes here
15 final statement of loop
That is also not a loop control statement. Instead it is an assignment to a complex number. % cat 0xfor.py
0xfor 1
% python -m tokenize -e '0xfor.py'
0,0-0,0: ENCODING 'utf-8'
1,0-1,3: NUMBER '0xf'
1,3-1,5: NAME 'or'
1,6-1,7: NUMBER '1'
1,7-1,8: NEWLINE '\n'
2,0-2,0: ENDMARKER ''
Also, literals have to be evaluated before names, otherwise you could overwrite them: >>> 0xzzz = 1
File "<stdin>", line 1
0xzzz = 1
^
SyntaxError: invalid hexadecimal literal
>>> 0xf = 1
File "<stdin>", line 1
0xf = 1
^
SyntaxError: cannot assign to literal
[0]: https://docs.python.org/3/library/stdtypes.html#float.fromhe...[1]: https://github.com/python/cpython/blob/5ce227f3a767e6e44e7c4...
[2]: https://github.com/python/cpython/blob/5ce227f3a767e6e44e7c4...
> the first thing it sees looks like the start of a number
Yes, because it checks if the token is a number before checking if it is a name.
But to know if a token is a name, it has to check, and that check never happens because the tokenizer yields when the number case hits. You can attach a debugger and see for yourself.
> It doesn't matter in which order the rules are checked.
Nowhere did I claim otherwise. The rules _do_ happen in a deterministic order, which I posted the source code for.
> Nowhere did I claim otherwise.
You claimed otherwise when you wrote "Because the hexadecimal parsing rule [0] happens before [1] the rule that parses names [2]", and when you wrote "literals have to be evaluated before names", and when you wrote "the rules can’t be evaluated concurrently". That is, you claimed otherwise in every single one of your posts in this subthread.
You are conflating the fact that there is an order to the rules with a strawman that the rules _must_ be evaluated in a specific order. All of the quotes you pulled are facts (Python checks if a token is a number before it checks if it is a name, Python checks if a token is a literal before it checks if it is a name, and the rules are not evaluated concurrently), but your strawman is making a different claim.
I'm not the one who wrote "literals have to be evaluated before names".
> rules are not evaluated concurrently
Strawman yourself. I didn't say they were evaluated concurrently, I said that your claim that they can't be evaluated concurrently was wrong. They could be, because they are disjoint, and regardless of order only one of them can match at input starting with a '0' character.
Anyway, I'm done here.
This isn't a bug. It's a slightly loose grammar spec that can surprise python users who don't play golf.
0xf = Hex for 15
Expression is evaluated as a boolean
>>> ast.dump(ast.parse('[0xfor x in (1,2,3)]'))
"Module(body=[Expr(value=List(elts=[BoolOp(op=Or(), values=[Num(n=15), Compare(left=Name(id='x', ctx=Load()), ops=[In()], comparators=[Tuple(elts=[Num(n=1), Num(n=2), Num(n=3)], ctx=Load())])])], ctx=Load()))])"* 0xfor1 evaluates.
* 1or 2 evaluates.
* 1or2 doesn't.
* ''or'foo' evaluates.
This is gross.
“1or2” is lexed into “1” (integer) followed by “or2” (identifier), which is valid on the lexer level but then fails on the grammar level.
I guess it’s an operator token after all.
https://docs.python.org/3/reference/lexical_analysis.html#wh...
This is not a bug.
> Whitespace is needed between two tokens only if their concatenation could otherwise be interpreted as a different token (e.g., ab is one token, but a b is two tokens).
The two tokens in this case are "0xf" and "or". Their concatenation cannot be interpreted as a different token, because "0xfor" is not a valid token. Therefore, if I'm reading the rule correctly, whitespace is needed in this case.
"0xffor" is another interesting case. It's also not a valid token, but it could be interpreted as two tokens in two different ways: "0xf" "for" or "0xff" "or". (Python does the latter. I presume it uses something like C's "maximal munch" rule.)
> Whitespace is needed between two tokens only if their concatenation could otherwise be interpreted as a different token (e.g., ab is one token, but a b is two tokens).
Because the concatenation of "0xf" and "or" can't be interpreted as a different token, the whitespace is not needed.
I dislike the rule, and I strongly think that "0xfor" should require whitespace between "0xf" and "or" (I'm sure that influenced my reading), but you're right about what the rule says.
(Apparently I can't edit my previous comment.)
But requiring whitespace between all tokens is not an acceptable solution, since "2+2" should work. Always equiring whitespace between alphanumerical characters in different tokens would make sense.
edit: 'cause it's a bug! https://bugs.python.org/issue43833
The question is whether "0xfor" is a valid way to represent the two tokens "0xf", "or". My reading of the rules is that it isn't, and that whitespace is required.
The rule, which I initially misinterpreted, is:
> Whitespace is needed between two tokens only if their concatenation could otherwise be interpreted as a different token (e.g., ab is one token, but a b is two tokens).
Since "0xfor" cannot be interpreted as a different (single) token, whitespace is not needed.
(I'm not a fan of this particular consequence of that rule, and I suspect it was not intended, but it says what it says.)
[sielicki@dogfruit ~]$ python3.10 -c 'print(0xf or some_undefined_variable in some_undefined_set)'
15This code "should" result in a NameError or something similar. But resolving that identifier to a name won't happen until we decide to evaluate the right-hand side of the `or`.
Thankfully, most python static checkers make some mostly-sane assumptions and will flag this code as an error, despite the subtle possibility of legally creating this name using some other Python/CPython features.
Type annotations are a great idea but in practice I don't see them used much. Should be great for dev teams working on bigger Python projects.
Whitespace in Python is significant to indicate statements and block.
I still can't figure out what the author was trying to do when he stumbled onto that, though...
That is, 0or, 0xor, 0and,etc. in my view, this should clearly be a syntax error.
edit: as noted by the parent comment, 0xor != 0 xor, but 0x or. But this seems to hold true for all data types.
I can see how this could slip by, it's possible that original python used a non-alphabetical token as the logical operator, then it was swapped out at some later time for or and and.
[0x1for x in (1,2,3)]
looks like it should parse as [0x1 for x in (1,2,3)]
but instead parses as [0x1f or x in (1,2,3)] >>> 1or 2
1One of those things that looks helpful at first glance, but is problematic in the long run. Throwing a syntax error immediately would lead to more robust/maintainable code.
I guess people don't remember PL/I very much: PL/I keywords are not reserved words, so it is possible to use them in a program in other than their keyword context. DO DO=1 TO TO BY BY;END=END+1;END; is a valid program.
and even if you do, future versions of the language may include more keywords.
There's an argument to be made that disallowing it could help in the long run as well, but Python has never been so strict. So I'd stop short of recommending it here. Perhaps in a bondage and discipline language. :-D
In most languages, numeric literals, identifiers, and keywords cannot be adjacent, but any of them can be adjacent to an operator. The odd thing here is "0xfor" being tokenized as "0xf" and "or". In C, for example, "0xfor" is an invalid token.
0xf is a numeric literal, and or is an operator, so as you said, they can be adjacent. or isn't an operator in C.
of course, if it were 0xfand ... you have a fun problem, because it could be 0xfa nd, but nd isn't an operator (that I know of), or 0xf and. Ideally, syntax shouldn't be ambiguous.
In most languages, numeric literals, identifiers, and keywords cannot be adjacent, but any of then can be adjacent to a token that consists of punctuation characters. (In some languages, most or all operators spelled using punctuation characters, which is what I had in mind when I wrote my previous comment.)
Yes.
> and if you #include <iso646.h> in C
Sort of. As a preprocessor token, it's an identifier that happens to be the name of a macro. As a token (after preprocessing), it becomes the "||" operator.
>>> 1or x # without x existing
1That said, I have a really hard time imagining this issue coming up in regular usage of the language so I don't really have much of an opinion should they choose to keep or drop it.
https://docs.python.org/3/reference/lexical_analysis.html#wh...
# launch debugger with the program provided inline, consisting of the statement `1`.
# this program is ignored and we just mess with some variables.
% perl -de1
Loading DB routines from perl5db.pl version 1.57
Editor support available.
Enter h or 'h h' for help, or 'man perldebug' for more help.
main::(-e:1): 1
DB<1> @a = (1..4)
DB<2> x \@a
0 ARRAY(0x7faf890418f8)
0 1
1 2
2 3
3 4
DB<3> x scalar @a
0 4
DB<4> sub foo { return (1..4) }
DB<5> @b = foo()
DB<6> x scalar @b
0 4
DB<7> x scalar foo()
0 ''
yes, that's the empty string, aka false, because the .. operator returns entirely different things in scalar context so you can write code like while (<STDIN>) {
next if /^BEGIN$/ .. /^END$/;
# ...
}`11-0xbor-11` is -11
and `11-0xbor-0xbor-0xbor-11` is -11, too
>>> [0xfor d or cambridge]
[15]15