5% of 666 Python repos had comma typo bugs (inc V8, TensorFlow and PyTorch)
codereviewdoctor.medium.com
codereviewdoctor.medium.com
I've moved away from working in Python in general, but I think the #1 feature I want in the core of the language is the ability to make violating type hints an exception[1]. The core team has been slowly integrating type information, but it feels like they have really struggled to articulate a vision about what type information is "for" in the core ecosystem. I think a little more opinion from them would go a long way to ecosystem health.
[1] I know there are libraries that do this, I am not seeking recommendations.
I just wish that the core team would take that same zeal for a "pythonic" experience with small code and use it to develop more scaled-up systems for dealing with larger code bases. My idea is to enforce strong pre-conditions on function calls using type hints, but I am sure there are other ways to do it.
Again - the dynamism of Python means teams can write amazing extensions to Python (like mypy), but that isn't a replacement for the core team having a plan for how they think typing information should be used at runtime. Their current answer seems to be "nothing," which disappoints me.
The problem in the article is more related to syntax, not types, with the problem that both forms are valid syntax with different but still very similar outcome.
Pylint on the other hand can find it with implicit-str-concat check enabled.
This is simply not true - Python with mypy isn’t even as strong as Typescript, let alone Rust, F#, Haskell and so forth.
foo,
to be a tuple? why not make this a syntax error? is there a use for single value tuples?
The thing is, your example is the way tuple should be defined. The parenthesis are merely allowed (and ignored). Why? I see this as a mistake of language creators. But to be fair, it is difficult to make a perfect language (or anything really), and Python is pretty close imho.
char ch_arr[3][10] = {
"uno",
"dos"
"tres"
}; # let ch_arr = [
"uno";
"dos"
"tres";
];;
Error: This expression has type string
This is not a function; it cannot be applied.the parentheses are not for function arguments, are for "invocation".
If function (or method) arguments don't require parenthesis, referring to (rather than calling) a function/method usually requires quite distinct syntax, so it's quite easy to it apart from a call.
It may not be familiar to people coming from languages where no-parens refers to the function and parens call it, but being clear and distinct and being intuitive to people indoctrinated in contrary syntax are not the same thing.
E.g., in ruby (which has methods but not functions in the strict sense) I can call a method with:
thing.square # or thing.square()
Or access the corresponding method object with: thing.method :square # or, thing.method(:square)
Either of the former options are distinct from both of the latter.For other examples of languages where invocations don't use parentheses for arguments, OCaml and Haskell. In fact, I'd argue that if they tried to add that feature to those languages (parens around arguments to a function), it'd make things very confusing given the way functions and tuples work.
I say "had a proper type system", but actually it turns out that it does have something like that: When I use python for anything else than a most tiny script now, I use "mypy"[1] which implements static typing according to some existing Python standard (whether that came about because of mypy or the other way around, I don't know).
It is so, so good to have mypy telling me where I messed up my code instead of receiving a cryptic, weird runtime error, or worse, no error and erratic runtime behavior. Because not knowing that a particular type is unexpected and wrong, values often get passed along and even manipulated until the resulting failure is not very indicative of the actual problem anymore.
https://github.com/UWQuickstep/quickstep/pull/9
https://github.com/tensorflow/tensorflow/pull/51578
Also, I personally don't mind this approach to string concatenation. I think it's a fine compromise between easy formatting and clarity. I was whining about a corner case of tuple construction - which as far as I know is not a feature of any other language.
(
"one",
(
"a very very"
"long long two"
),
"three"
)
And of course a,b should be syntactically invalid. It must be (a,b)In hindsight, singleton tuples are not common or useful enough to deserve their own syntax. If the way to create them was something like this:
t = tuple.single("hello")
we'd thing it's ugly or inconsistent, but definitely not confusing or bug-prone. x = (1,2,3)
#print("the value of x is %s" % x) # breaks if x is a tuple
print("the value of x is %s" % (x,)) # works even if x is a tuple
There is a readable way to create singleton tuples, without the sneaky trailing comma or a new function like tuple.single: tuple(["hello"])
The square brackets can be slightly annoying. I recall writing the following function to omit them: def tup(*args):
return tuple(args)
This basically lets you use the usual tuple syntax, just prefixed with the word "tup". The advantages are that you don't need a trailing comma for singleton tuples, and it's more obvious that a tuple is being created (it can be difficult to distinguish between tuple literals and parentheses used for grouping in a complex expression).I am reminded of a somewhat similar issue with empty set literals: {1,2} is a set, {1} is a set, but {} is a dict. The way to create empty sets is using set().
More broadly, the https://codereview.doctors makers are making the point that their tool caught an easy-to-miss issue that most wouldn't think to add a rule for. A bit of an open question to me how many of those there really are at the language level, but still seems like a neat project.
Also in terms of mistakes codereviewdoctor twice linked to the same issue in their blog https://github.com/tensorflow/tensorflow/issues/53636 and raised the PR to the wrong project https://github.com/tensorflow/tensorflow/pull/53637 (I guess Tensorflow vendors Keras, easy mistake)
Also a factor that bugs in functional code are more visible, both during development and to users once shipped. So there may have been an equal number or more such bugs in the non-test code, that just didn't remain in the code base for this long.
STOP!
This folder contains the legacy Keras code which is stale and about to be deleted. The current Keras code lives in github/keras-team/keras.
Please do not use the code from this folder.
Yeah, not the most obvious notice.The fact they didn't find the same mistake(s) in keras-team/keras (I assume they scanned, it's one of the most popular Python repo) makes me believe these issues have been fixed/removed in up-to-date karas repo.
https://github.com/keras-team/keras/issues/15854
resulting in
"Absent comma results in unwatned string concatenation on line 330"
Bug-ception!
typo in the url (or in HN's markup) btw: it's https://codereview.doctor
> This PEP is rejected. There wasn't enough support in favor, the feature to be removed isn't all that harmful, and there are some use cases that would become harder.
The most common way to split a string in lines is using this concatenation formula.
Preventing a bug that occurs in 5% of observed codebases (and anecdotally, happens to me during development all the time) seems like about as good as reasons get.
Swapping a perfectly fine print statement for a function, on the other hand… that’s the breaking change in Py3k that’s never seemed worth it to me.
2to3 could also trivially add +, and if anything, that would actually help surface these kind of bugs, because if you randomly see a + in the middle of your list of strings, it's much easier to spot the bug than if there was a missing comma.
Is it really? I tend to avoid it in favour of ””” or ‘\n’.join(<list of lines>), because it looks like a mistake.
Triple quotes are kind of annoying if the string is indented, but you can just not indent the string to avoid the whitespace.
https://docs.python.org/3/library/textwrap.html#textwrap.ded...
or inspect.cleandoc.
https://docs.python.org/3.8/library/inspect.html#inspect.cle...
No footgun potential, and as others have mentioned the “good usage” would often be bad simply because it ends up looking like a mistake even if it’s intentional.
I suppose this is also something you could catch with a linter?
"from __future__" is meant to only ever be used temporarily with a specific Python version slated for it becoming the default behavior.
This discussion about flags has come up recently as part of the debate of accepting PEP 649 or PEP 563 or something else continues. If the Steering Council does not accept PEP 563 it will need to be figured out how to deprecate "from __future__ import annotations" without making it the default and how to implement it's replacement.
Explicit is better than implicit.
And yet, s = ["one", "two" "three"] will implicitly and silently do something, that is probably wrong most of the time.
There should be one-- and preferably only one --obvious way to do it.
the author used two different ways of hyphenating (three, if you count the whole PEP 20). PEP 20 is clearly not meant to be taken as law. Nor PEP 8. Nor PEP 257.People frequently mistake "one obvious way" with "one way". There are lots of ways to iterate through something, for example, but there is really one obvious way. And the philosophy here still applies: when you read anyone else's python code, the obvious way is probably doing the obvious thing. I think that is the more appropriate takeaway from PEP 20.
For example, you'll sometimes see people do bad stuff like this:
>>> lst = []
>>>
>>> [lst.append(i + i) for i in range(10)]
[None, None, None, None, None, None, None, None, None, None]
>>>
>>> lst
[0, 2, 4, 6, 8, 10, 12, 14, 16, 18]
>>>
When they should be doing this: >>> lst = []
>>>
>>> for i in range(10):
... lst.append(i + i)
...
>>> lst
[0, 2, 4, 6, 8, 10, 12, 14, 16, 18]
>>>
Or just this: >>> lst = [i + i for i in range(10)]
>>>
>>> lst
[0, 2, 4, 6, 8, 10, 12, 14, 16, 18]
>>> lst = [range(0, 10, 2)] lst = list(range(0, 20, 2))To iterate through it without doing one of those things, you use a for loop.
In “one obvious way to do it”, “it” refers to a concrete task; the same is not necessarily intended to be true of arbitrarily broad generalizations of classes of tasks.
I don't get what you mean by this.
When I read someone else's code, what is obvious to me isn't necessarily what was obvious to the author. For an illustration of this, have a look at the day 1 solution thread from this year's Advent of Code - https://www.reddit.com/r/adventofcode/comments/r66vow/2021_d... (you can search for Python solutions) - and see how many different ways there are to solve a fairly straightforward problem.
No, first, it doesn't use hyphenating at all, it uses hyphens as an ASCII approximation for typographical dashes used to set off a phrase (a distinct function from hyphenation), and, second, in that quote they used one way of doing it: “two dashes set closed on the side of the main sentence and set open on the side of set-off phrase”.
It is an unusual way of doing it—just as with actual typographical dashes, setting open or closed symmetrically would be more common—but it's not two ways.
EDIT: And the third use (in the heading and later in the body) is seperating parts where neither is a mid-sentence appositive phrase, and uses open-on-both sides. So that's not a different way of doing the same thing, it's a different way of doing a semantically different thing.
Actually, I think the dash use makes a good illustration of how the “it” in “one way to do it” is intended.
Eh, I don't think that's the interpretation the author was going for. The author wanted to show two different ways of approximating a dash, and he had limited options.
If he'd done this-- for example-- he would have been showing one way, not two.
If he'd done this --for example-- you would have called it "two dashes set open on the side of the main sentence and set closed on the side of set-off phrase".
If he'd done this-- for example -- it would have been too obvious (on the same line).
I suppose he could have done this-- for example--but I still think that would have been too obvious. You're not supposed to see it on a first read.
> And the third use (in the heading and later in the body) is seperating parts where neither is a mid-sentence appositive phrase, and uses open-on-both sides. So that's not a different way of doing the same thing, it's a different way of doing a semantically different thing.
It's a different use of a dash, but it's still a place where you'd typically use a dash.
-----
Edit: You know what, thinking about it again—perhaps both interpretations are valid. That almost adds to the effectiveness of the whole thing.
There should be at least one-- preferably only one --obvious way to do it.
Kinda funny meta joke considering everybody conflates "one" and "only one" to mean the same thing. Preferably there would only be one obvious way to describe "one". :pIs it really that crazy do set up a figure, axes on that figure, and plot on the axes, returning an artist object for each plotting command?
The Figure is the final image that may contain 1 or more Axes.
The Axes represent an individual plot (don't confuse this with the word "axis", which refers to the x/y axis of a plot).
This is infuriatingly bad and I firmly believe that it makes sense only to people who already know how it works. There's an image, axes (this word alone is a crime), plot, figure... it's like they took a bunch of synonyms and arranged them randomly to put together an API.why so? you prefer something like axiis?
> Axes object is the region of the image with the data space.
In matplotlib axes is not the plural of axis. It has its own meaning specific to the API. And at the same time it's the plural form of another word (axis) which is also relevant in this context and it sounds almost identical when pronounced.
https://www.mathworks.com/help/matlab/ref/axes.html
https://www.mathworks.com/help/matlab/ref/axis.html
https://www.mathworks.com/help/matlab/ref/figure.html
So they emphasize the cartesianess of the axes.
If I was designing something like it, I wouldn't recommend either. The global one has many fewer WTFs per character, but the objects one looks like it works in a multithreaded program or that you can create more than one plot without displaying them (but I've never tested this).
Personally, I like plotting in R way better than in python. It has a lot better developer UX.
Not in comparison to Perl, which usually has multiple ways to do anything, each 'obvious' to different sets of people (each Perl codebase therefore seems to have a distinct dialect based on which 'obvious' alternatives are chosen).
The other direction languages can take that is being contrasted, is there being one non-obvious way to do something.
Python's 'most obvious way' isn't necessarily the fastest/most concise/most efficient/scalable/etc. way to do something in Python, but it will usually be obvious to most Python developers. And although broad styles have certainly developed over time (imperative, functional, OO) as Python has gained power and flexibility, the dictum still largely holds true.
ie, instead of
for line in lines: print(line)
we are supposed to be using
while line := f.readline(): print(line)
I've not been super impressed with this type of thing.
That said, string formatting is better with f strings.
They also rolled back some the forced breakage from trying to force unicode with 3 which made a big difference. 3.3 added back u''
Lots of good cleanups lstrip vs removeprefix etc.
Underscores in numeric literals (10000000 vs 10_000_000)
So lots of good stuff still landing.
> for line in lines: print(line)
> we are supposed to be using
> while line := f.readline(): print(line)
No, we’re not. Walrus, in loops, IME, is more for replacing this pattern:
while True:
myvar = get_it()
if not ok(myvar):
break
# code that uses myvar
with this pattern: while ok(myvar := get-it()):
# code that uses myvarI'm not the only one who looked at the recommended examples of the use case here and went, huh?
https://news.ycombinator.com/item?id=17450890
Recommended new way:
if any(len(longline := line) >= 100 for line in lines):
print("Extremely long line:", longline)
Old way: for line in lines:
if len(line) >= 100:
print("Extremely long line:", line)
break
I prefer the old way. These were examples in the PEP!In your example get_it() might be better as a generator or iterable. A lot of code looks great if you push that type of thing down a bit, and sometimes memory is helped as well. Then you iterate over it, for values in get_it. This keeps python very natural. You start to get a lot of weird line noise type code with := vs the old python style which while a bit longer was basically psudo-code.
All it does is increase line-noise, and for what? So we don't have to write 2 short lines, or save an indentation level somewhere?
There is a good reason why assignments in Golang are not expressions, even though they are in C, and the language is otherwise deliberately close to the mindset of C; The added convenience makes the code much harder to read.
Sure;
char c;
while((c = getch()) != EOF) {
// do something
}
requires less lines than; char c;
while(1) {
c = getch();
if (c == EOF)
break;
// do something
}
but it's also easier to read, because each line carries less information. That's what people call "line noise".IMO, := is a step in the wrong direction, and sadly I see python take more and more of these, going from the deliberately simple and clear language to something that's becoming needlessly hard to read by piling on things it doesn't even need.
What? Something being complex is artificial, we try to avoid it. Problems can be complicated, we try to simplify them, and more complicated the problem is, we tend to develop more complex solutions. So comparing them does not make sense?
Or did I always know them wrong?
Complicated: consisting of many interconnecting parts or elements; intricate.
Nothing specifically artificial about either one. Software that is well decomposed is Complex (made of many smaller connected parts). Software that is is poorly decomposed is Complicated (made of many smaller interconnected parts).
Connected vs interconnected?
Interconnected: connected at multiple points or levels (aka spaghetti code)
Complex: we took all of these simple steps, lumped them together, now we have this
I always took it to mean 'complex' as in having many connected parts, and 'complicated' more as in over-complicated or convoluted - the opposite of 'simple'. In other words, breaking something complicated into a system of intentionally-designed pieces is probably better than a chunk of opaque code to brute-force the current case. A good system is probably also 'simpler', despite having more pieces and interconnects.
My understanding is that: * Complex domains lend themselves to experimentation and emergent behavior. * Complicated domains lend themselves to analysis, expertise, and rule following.
The Wikipedia article offers the domains as containing "unknown unknowns" and "known unknowns" respectively.
I'm trying to think how this maps to Python -- the language is complicated, while the problems we're solving are expected to be complex? Or, maybe, the language lives at the boundary between complicated and complex. We push complicated procedures into the language, and let the programmers deal with complex issues?
from that context it makes sense, because the only goal of python in the 1990s was to be more popular than perl, which was notorious in having many ways of doing the same thing.
but yeah, python had had significant feature creep over the years, it's nowhere near the small clear lang it used to be.
There's match/case in 3.10 - https://www.python.org/dev/peps/pep-0636/
match/case (not a drop in switch statement)
>breaking out of loops
break
>ending scripts early (for explorative programming) exit() or sys.exit()I knew Python wasn't for me in my first foray into it when I fired its REPL and then went to exit it with control-C or whatever and it literally printed out the right way to do it but then didn't do it. Python was more interested in having me do things a certain way even when it knew what I intended to do, just to be a twit.
>>> exit
Use exit() or Ctrl-D (i.e. EOF) to exit
You will get that error response. The goal of this is to have the REPL language the exact same as the scripting language. exit() is supposed to be called as a function to make the language more consistent, so just typing `exit` will do nothingUseful would be, if the default handler for SIGINT would not raise an exception, but have a useful default like eg. terminating the program. Go handles SIGINT this way by default.
If I want an exception, I can just tell the program:
import signal
signal.signal(signal.SIGINT, throwException())
The way it is now, the exception bubbles up to runtime, and if it isn't handled (eg. in the REPL) the program crashes, or worse, hangs if there are other threads of execution running: import threading
import time
def sleepN():
for i in range(20):
time.sleep(1)
threading.Thread(target=sleepN).start()
time.sleep(20)
Press c-C here, and the thread will still run, because the bubbled up Excp only kills the main thread. This is a real footgun in applications which rely on SIGINT being a termination signal, and have long running threads. $ python3
Python 3.9.2 (default, Feb 28 2021, 17:03:44)
[GCC 10.2.1 20210110] on linux
Type "help", "copyright", "credits" or "license" for more information.
>>> exit
Use exit() or Ctrl-D (i.e. EOF) to exit
>>> exit.eof
'Ctrl-D (i.e. EOF)'
>>> exit.name
'exit'
>>> exit = 42
>>> exit
42
>>> exit()
Traceback (most recent call last):
File "<stdin>", line 1, in <module>
TypeError: 'int' object is not callable
>>>
I would have special-cased exit, though.This is a list and you must explicitly place a comma when you want to start a new element in the list. Is there ever a time a new element follows a previous one and is NOT separated by a comma? No, this is explicit.
Whereas, strings also always concatenate in this manner be it in a list context or not. It seems like you're assuming behaviors from other languages would be the same in another.
I'd expect that to be an error.
This is why i like Go/Rust. I detest the implicit warts of these languages.
Sarcasm aside, I'd assume people primarily list things in between [ and ], and sometimes concatenate things in there too. The language should err on the side of doing what people expect, unless explicitly told not to.
> It seems like you're assuming behaviors from other languages would be the same in another.
Rather, I think people expect a language, especially one this big and important, to work for them, and not to be designed with unergonomic features instead.
Yes:
[ "one, two", "three" ]
The comma is not an absolute context-free indicator of element separation.It doesn't seem to have anything to do with typing discipline.
words = (
'yes',
'correct',
'affirmative'
'agreed',
)
Would be a tuple (immutable list) of strings, while words = (
'yes',
'correct',
'affirmative',
'agreed',
)
would also be a tuple of strings.If haskell had for some reason decided to have the same syntax sugar, it also would have caused an issue.
Yes, you'll certainly find somebody who doesn't know what 'not statically typed' means, but ... And yes, there are also C(++) users, that expect strings to be concatenated like that.
foo = 5
fop = 6
Keywords like `let` solve this problem: let foo = 5
fop = 6 # error let foo = a();
let foo = b(foo);
let fop = c(foo);
let foo = d(foo);
(Which is valid, e.g., in Rust.)Auto completion, highlight matching variable, gray out unread variables and warning of unused assignment.
I’ve written lots of python and can’t recall ever having this issue. More likely is a logic typo of two similar variables like length_x vs length_y, where a “let” wouldn’t have saved you anyway if both are already defined.
JavaScript, pre strict TS, on the other hand, where missing var implied global was a real motherload of bugs. Or kotlins “val” vs “var” changing semantics completely…wow. But those are different concepts from basic definition I know.
usage = (
'usage: foo [options] filenames...\n'
' -f force concatenation\n'
' -c for convenience\n'
)
print usage
Edit: forgot to add parenthesesMissing an operator resulting in explicit behavior is much more subtle and not even obvious behavior. For those who use python, it is worse.
it's pretty tough to argue such languages really shouldn't exist
Well, I agree with OP so that is at least two people. I really don't see it as a good trade."Shouldn't exist" is too strong.
Dynamic languages that let you create a new variable via assignment shouldn't be used to create non-trivial software. How about that?
Scripting languages have a place. That place is 100% in creating quick-and-dirty scripts and tools. Or in doing some kind of one-off data transform (as is common in machine learning scenarios). Anything that has a life span of two weeks or less, or a code length of fewer than a hundred lines? Yeah, script languages rock for that.
Explicit/static typing adds vastly more value to large projects than the cost of the overhead. The fact that you can't really gain that value in Python means that Python should be relegated to quick and dirty scripts.
Same for JavaScript, Ruby, and other completely dynamic languages.
You'll note that all of these languages are getting types one way or another, meaning that there are a lot of people who do recognize their value. Though TypeScript is years ahead of the rest in the completeness and sophisticated of its type system; bugs like the comma bug detailed by OP, along with simply every JavaScript "wat" bug, simply can't happen in TypeScript in strict mode. And static types enables entire other categories of bugs to be detectable via a linter as well.
I'd take a project in a dynamic language with a decent test suite over a project without tests in a statically typed language any day of the week.
I'd take the opposite. I've read too many useless tests in python codebases that can be accomplished by a static type checker. "Decent" does a lot of heavy lifting in your comment. And what about a dynamically typed codebase without any tests? I'm sure they exist.
I'd rather dive into a big ball of mud with a compiler that will help point me to my mistakes before I release them, than having to sift through a ball of mud trying to find that mistake with production services flailing.
That all being said I've worked with both types of languages in successful projects. But I prefer the development experience of the statically typed variety.
There are reasons dynamic language (or specifically Python), but I haven't heard one explanation how it helps writing fewer tests.
The cognitive burden of having to memorize and look for which variables are new vs which are being modified is simply not worth it in my opinion, even for a scripting language. Maybe for esolangs, simple math or first time learning programming.
In any case, it's a short coming of the language (IMO) but not a deal breaker. We learn to live with it.
global x = 1
def setx():
if True:
x = 2 # completely different x
print x # prints 2, visible outside of `if`
setx()
print x # prints 1
This led to a funny keyword 'nonlocal', because you can't simply ignore scoping and pretend that you're BASIC in any serious program.(To my opinion, python had a good start, but lost in the woods for no clear reason. It's a movie mutant of a language, which tried to appeal to non-programmers and somehow succeed, and then realized that non-programmers eventually become ones, and it's not hard. Now it's too late to fix this mess. End of opinion.)
In Scheme, for example, this is not an issue.
But that's just what comes with a hyper flexible language like python. You can do lots of things in lots of different ways, but you can also screw things up just as easily, and your IDE won't tell you because technically it's valid code.
Is this "operator" overloadable on each type in Python?
And that scares me a lot. I think I have to reevaluate my position towards Python.
Thanks.
I get the use case as you described it, but it just seems like minimal effort to accomplish and have some semblance of explicit/safety.
So strange that Python has completely different syntax from C, but they chose to copy this obscure syntactic feature _even though they have the plus operator on strings_.
mylongstring = "hello" +
"world"
No idea if python's way of indentations allows this but sounds like it should mylongstring = ("hello" +
"world")
or, without `+` mylongstring = ("hello"
"world") mylongstring = "hello " \
"world " \
"my " \
"name " \
"is"*> The preferred way of wrapping long lines is by using Python's implied line continuation inside parentheses, brackets and braces. Long lines can be broken over multiple lines by wrapping expressions in parentheses. These should be used in preference to using a backslash for line continuation.
It's common in some languages and used the way you use it. I looked in PEP8 and it seems they don't discuss this.
I think it's a perfectly valid use case, but clearly there are two camps to this. If this is so contentious, I would recommend PEP8 be revised to either explicitly endorse it as a way to split long lines or to explicitly discourage it and recommend the + operator instead.
The string concatenation in itself should not be a problem as it's really just string constants. (But again, it might be irony exactly because of this :) )
("foo" "bar", "baz")
and
("foo", "bar", "baz")
I come from a programming platform (C#) where productivity is a key element of language design. I highly doubt that Anders Heijlsberg would have accepted such a error prone concept like a literal free implicit operator on a key type like strings.
I kind of remembered that some languages do support it for braking strings into multiple lines conveniently. I'm a bit surprised that it works even on line (I've never used it, because why would have I), but you'll likely to make the mistake on multiline statements anyway. I've also checked and it doesn't work in java (which I kind of remembered, though I mostly do python these days).
What is inconvenient about just adding a + at the end or beginning of the line?
Even without static typing, argument length verification etc. can be done with a suitable compiler. In python we are left chasing 100% code coverage in unit tests as it's the only way to be certain that the code doesn't include a silly mistake.
One of our products is a universal linter, which wraps the standard open-source tools available for different ecosystems, simplifies the setup/installation process for all of them, and a bunch of other usability things (suppressing existing issues so that you can introduce new linters with minimal pain, CI integration, and more): you can read more about it at http://trunk.io/products/check or try out the VSCode extension[0] :)
[0] https://marketplace.visualstudio.com/items?itemName=Trunk.io
puts "a" "b" == "ab" # true
and puts "a"
"b" == "ab"
prints "a" with "b" == "ab" evaluated to false and discarded. This could create bugs as with Python. However ["a"
"b"] == ["ab"]
is syntax error at the beginning of the second line. The parser expects a ]
It would evaluate to true if it were on one line.# list
list = "a","b",
# function
def foobar
end
=> ["a", "b", :foobar]
The implicit concat of string literals is the culprit here. It really should require "+".
I was just doing some simple refactoring, changing a hard coded sting into a parameterized list of f-strings that’s filtered and joined back into a string.
I’m glad that I had unit tests that caught the problem! I couldn’t figure out why it was breaking, that comma is very devilish to spot with the naked eye. I’m surprised my linters didn’t catch it either. Maybe time to revisit them.
It's both good for those projects and for the company that does the marketing since they reach there exact target group. Plus it gets them on the front page of HN.
https://github.com/YosysHQ/prjtrellis/pull/176
https://github.com/UWQuickstep/quickstep/pull/9
https://github.com/tensorflow/tensorflow/pull/51578
https://github.com/mono/mono/pull/21197
https://github.com/llvm/llvm-project/pull/335
https://github.com/PyCQA/baron/pull/156
https://github.com/dagwieers/pygments/pull/1
https://github.com/zhuyifei1999/guppy3/pull/12
https://github.com/pyusb/pyusb/pull/277
https://github.com/KhronosGroup/Vulkan-ValidationLayers/pull...
It is indeed a very common mistake in Python, and can be very hard to debug. It bit me once and wasted a whole day for me, so I've been finding/fixing them ever since trying to save others the same pain I went through.
EDIT: I will point out that I've found this error in other non-Python code too, such as c++ (see the 2nd PR for example).
Here's the regex for anyone curious:
[([{]\s*\n?(\s*['"](\w)+['"],\n)+(\s*['"]\w+['"]\n)(\s*['"]\w+['"],\n)*
https://chromium-review.googlesource.com/c/v8/v8/+/2629465/3...
Personally, I prefer uniform lists with leading commas, because it's easier to add and remove lines for later, inevitable refactoring. For example, I prefer:
things = [
'foo'
, 'bar'
, 'baz'
]
This drives some people crazy, but I think it's the One True Way. things = [
'foo',
'bar',
'baz',
]
even better? In your case, if you want to add something to the beginning of the list you'll have to modify two lines.ETA: Let me expand on why it's important to put the comma first. Which list is more clear to you:
a
, dog
, weather
, banana
, b
, car
or a,
dog,
weather,
banana,
b,
car
With the leading commas, they all line up, and you can see them in a neat little row. I really prefer it especially in contexts where the trailing comma is not permitted, such as a SQL query: SELECT
name
, date
, operation
FROM
stuff> Apple I was the first product ever announced by the company in 1976. The computer was put on sale for $666.66 at the time.
https://9to5mac.com/2021/11/25/steve-woz-signs-rare-1976-app...
I've been bitten by this one at work, and can't help but think it is an insane behaviour, given that ['foo' + 'bar'] explicitly concatenates the strings, and ['foo', 'bar'] is the much more common desired result.
edit: This also applies to un-separated strings, so ['foo''bar'] also becomes ['foobar']
I don't think it fits well in python
(4 + 5) * (8 + 2)
(this and that) or (theother)
These elements should not become 1-tuples after the interior contents are evaluated. I sometimes add parentheses even around single variables just for visual clarity.Also, this allows you to do dot-access on int / float literals, if you want to
# doesn't work
4.to_bytes(8, 'little')
# works
(4).to_bytes(8, 'little')And there’s no evaluation of importance as to whether these instances are in test files or non-critical code. Packages are big and can have hundreds or thousands of files.
It could be that if these mattered, they would have been detected and fixed.
A good example for unit tests and perhaps checking to see if these bugs are covered or not covered.
I like these kinds of analyses but don’t like the presented like it’s some significant failure.
Python has a few of these things, which is really sad.
the tests are only as good as the code they're written with, and as good as the code review process they were merged under.
It avoids replicating the same category of errors in both the test and the code under test, especially when some calculation or some sub-tests generation is made in the test.
Also, the obvious solution is self-testing code. (Jokes aside, structures like code contracts attempt something like this).
You could say that about anything and everything in software.
It's not acceptable that testing needs to be run for something the language should 100% accommodate.
The whole point of the language is to provide algorithmic clarity and avoid these things.
This isn't really an issue of 'trade offs' is just a bad feature of the language that should have been remedied more than a decade ago.
The lack of proper declaration of variables is even more absurd, there's only downside to that.
Along those lines. I wonder how many of these come from ad-hoc file path handling instead of using pathlib.
test did not work but did not fail either, imagine being that dev maintaining the code that the test professes to cover. Imagine being the user relying on the feature that test was meant to check (if the feature under test actually broke).
{
'key': (
'long string long string long string'
)
}
Using parentheses like this to put long strings on their own line is standard practice. title = 'Hello world',
I, for one, have often used this deliberately.Instead of:
s = ['a', 'b', 'c']
I'll type: s = 'a b c'.split()
For multiline lists where I want to get rid of leading whitespace I'll add lstrip(): lines = """line 1
line 2
line 3
""".split('\n')
lines = [line.lstrip() for line in lines]..."there are perfectly cromulent reasons a developer would do implicit string concatenation spanning multiple lines"...
https://www.merriam-webster.com/words-at-play/what-does-crom...
https://github.com/PyCQA/pylint/issues/1589
Is there usually enough context for a linter to make an educated guess?
Maybe as a matter of linting. As a matter of language design, I think + for string concatenation is a big mistake; using different symbols for numeric addition and string concatenation is something Perl got right.
But my impression using pylint is that its default settings are wildly opinionated, hence the surprise that this wouldn't have fallen under that umbrella.
"put"
, "Commas"
, "first"
, "to"
avoid these kinds of things.A paragraph is repeated and the markdown links at the end are broken because there is a space between ] and (.
Using the plus operator to concatenate strings is just weird.
Think of the usual algebraic properties these operators are supposed to have.
"+" always is supposed to be commutative--so "a"+"b" = "b"+"a", if those mean alternatives (they usually do mean that in mathematics), is just fine.
On the other hand, multiplication is often not commutative--also not here. "a" "b" != "b" "a".
So string concatenation should be the latter. And indeed that's how it's in regular expression mathematics for example.
This might be nice from a math point of view, but I think users are going to be confused using "string"^3 for repetitions (instead of "string"*3). + and * make too much sense to the unwashed masses.
At any rate, explicit is better than implicit.
Well, except if you wanted to support user classes that could duck type as both strings and numbers, which it would make awkward.
Furthermore, Python already uses * for strings to indicate repetition: ("foo" * 2 == "foofoo").
String concatenation really just needs its own separate operator. & is an obvious candidate, if only it wasn't so commonly appropriated for bitwise AND - which is a very poor use of a single-char operator as it's not something that you need often, especially in a language like Python.
On the other hand, D uses binary ~ for concatenation. That has a neat mnemonic: it's a "rope" that "ties strings together".
(Python copies some bad ideas from C. Another one is having to import everything you use. It seems that since Python is written in C, its designer took it for granted that there will be something analogous to #include for using libraries, even standard ones that come with the language.)
Implicit string literal catenation is tempting to implement because it solves problems like:
printf("long %s string"
"nicely breaks up"
"with indentation and all",
arg, arg, ...)
and if you're working in a language which has comma separation everywhere, you can get away with it easily.There are other ways to solve it. In TXR Lisp, I allow string literals to go across multiple lines with a backslash newline sequence. All contiguous unescaped whitespace adjacent to the backslash is eaten:
This is the TXR Lisp interactive listener of TXR 273.
Quit with :quit or Ctrl-D on an empty line. Ctrl-X ? for cheatsheet.
TXR needs money, so even abnormal exits now go through the gift shop.
1> "abcd \
efg"
"abcdefg"
If you want a significant space, you can backslash escape it; the exact placement is up to you: 2> "abcd\ \
efg"
"abcd efg"
3> "abcd \
\ efg"
"abcd efg"
4> "abcd \ \
efg"
"abcd efg"
5> "abcd \ \
\ efg"
"abcd efg"Maybe it is that through my work I use a half dozen languages, where it is hard to remember each in detail.
I have also worked on a javascript project where there were no imports/requires and the build process created one file. So you had to inspect the confusing build script to even know what was what.
Build processes creating one file is the seven decade norm in computing.
Even if you literally don't catenate the .js files into one, they get loaded into one running image one way or another.
And especially how I can choose the best way to indicate the sources of names in my code:
import time
t = time.perf_counter()
import time, my_module
t1 = time.perf_counter()
t2 = my_module.perf_counter()
from time import perf_counter as std_counter
from my_module import perf_counter as my_counter
t1 = std_counter()
t2 = my_counter()
try:
from my_module import perf_counter
except ImportError:
# Fall back to standard implementation
from time import perf_counter
t = perf_counter()
# import time as m
import my_module as m
t = m.perf_counter()Yikes; you're renaming/aliasing global identifiers! Just no.
It's an essential feature used in all sorts of everyday code.
C99 added printf conversion specifiers that are hidden behind macros, and idomatic usage of them relies on string catenation.
uint32_t x = 0;
printf("x = " PRIx32 "\n", x);
where PRIx32 might expand to "%lx" (if uint32_t is the same as unsigned long in that compiler).All sorts of C macrology relies on string catenation. Kernel print messages:
printk(KERN_EMERG "%s: temperature sensor indicates fire!", dev->name);
^ must not have comma hereAlthough I just had a (logging) use case in go where I missed cpp macros - wanted the log statement to get something from the file and just had to pass it in as another parameter.
If I have a uint32_t which needs printing I cast it to (unsigned long) and use %lu or %lx. This requires more typing in the argument list, but keeps the format string tidy. It's important for the format string to be tidy, because that's the reason of its existence: to clearly and concisely convey the shape of what is being printed.
That would be adding 2 pointers, and that's indeed illegal.
However, you can subtract them: “abc” - “def” . Now, the result is not a pointer any more, it's a ptrdiff_t (an integer type), so most compilers will warn if you try to assign that to a char *.
String catenation ("adding") by adjacency (no visible operator) is a thing; "add" doesn't imply that we are talking about a + operator:
$ awk 'BEGIN { x = "abc-" 2 + 2 "-def"; print x}'
abc-4-def execl("/bin/sh", "/bin/sh", "-c"
"echo foo", (char *) NULL);
Here we get one "-cecho foo" argument passed to the shell instead of two, so it can't work.Initializers are another example:
char *strArray[] = {
"how", "now"
"brown", "cow"
};
In non-variadic function and macro calls, you will most likely get an insufficient arguments error, unless another mistake compensates for that.If the language supports json, it should just do that.
1> #J[1,2,3]
#(1.0 2.0 3.0)
2> (get-json "[1,2,3,{\"foo\":true}]")
#(1.0 2.0 3.0 #H(() ("foo" t)))
3> (put-json #(1.0 2.0 t))
[1,2,true]t $ irb
irb(main):001:0> { hello: "world" }.to_json
NoMethodError (undefined method `to_json' for {:hello=>"world"}:Hash)
irb(main):002:0> require 'json'
irb(main):003:0> { hello: "world" }.to_json
=> "{\"hello\":\"world\"}"If it was CLOS with multiple dispatch, it would be easier to swallow. Because it would look like:
(to-json { hello: "world" })
;; error: no such function!
Then load the module, and you have a generic to-json function now, with a method specialized to handle the dictionary object and all. (I still wouldn't want to be doing this if it's supposed to be a language built-in).I regard the ability to add new methods to a class as good, but with a valid use case, like extending some third party piece with new methods in your own application. And the fact of not having to declare methods in a class definition, which is cumbersome. Just write a new method in that class's file, at the bottom, and there it is.
I ideally don't want that third-party piece itself to be divided into three pieces that I have to separately load to get all of the methods. Or worse, pieces from separate third parties that add methods to each other.
I copied a thing or two from Ruby in TXR Lisp. The object system as a derived hook, and that was inspired by something in Ruby:
1> (defstruct foo ()
(:function derived (super sub) (prinl `derived @super @sub`)))
#<struct-type foo>
2> (defstruct bar foo)
"derived #<struct-type foo> #<struct-type bar>"
#<struct-type bar>
3> (defstruct xyzzy bar)
"derived #<struct-type bar> #<struct-type xyzzy>"
#<struct-type xyzzy>
The derived hook is inherited (like any other static slot), so it fires in bar also. The function can distinguish which class is being derived by the super argument. @"
here strings in PS are fine for this purpose and
even allows whitespace anywhere
but because of the latter you can't indent it
with your other code
"@ -split "`r`n" | % {'<SOL>{0}<EOL>' -f $_ }
<SOL> here strings in PS are fine for this purpose and <EOL>
<SOL> even allows whitespace anywhere <EOL>
<SOL> but because of the latter you can't indent it <EOL>
<SOL> with your other code <EOL>https://unix.stackexchange.com/questions/76481/cant-indent-h...
long %s stringnicely breaks upwith indentation and all"
? In my experience, this always gets ugly when you want to insert spaces (= about always). Do you put them at the end or at the start of each string (apart from the first or last string)I think scala’s mkString (https://superruzafa.github.io/visual-scala-reference/mkStrin...) is the best solution, visually, for such things, but unfortunately, it would require hackers in the parser to do the concatenation at compile time, where possible.
Scala’s multiline strings look nice, too, if you want to insert newlines, except for the stripMargin thing (https://docs.scala-lang.org/overviews/scala-book/two-notes-a...)
"foo" "bar" // error
"foo " "bar" // OK
"foo" " bar" // OK
"foo" "" "bar" // OK: "" doesn't start with non-whitespace, since it's empty
"foo" " " "bar" // OK
The nice thing about this is that it's perfectly comatible with existing C.All we have to do is to implement a compiler warning which detects when the rule is violated.
Users who implement it have to fix situations like "foo" "bar" into "foo" "" "bar".
Probably the rules should be smarter. Some kind of tokenization concept could be at play so that gluing together two letters or digits is bad, or two punctuation tokens, but letter/number and punctuation is okay.
"foo" "1" // error? OK?
"foo" "bar" // error
"foo" ".bar" // OK
"1." "2" ".3" // OK
"1." ".2" // error: punct-punctThe alternative is what exactly? Have the entire standard library exposed at once? Make all modules create non-conflicting names for exported objects, so that the json parse function has to be called json_parse and the csv parse function has to be called csv_parse?
Seems less than ideal to me.
If these things are classes in a plain old single-dispatch oop system, you can havec a json-parser and csv-parser which have parse methods.
There could be packages/namespaces. So csv:parse and json:parse. These packages are standard and so they just exist; nothing to import.
In Python, you cannot use anything without an import! The top-level modules (which serve as de facto namespaces) themselves are not visible.
Say there is a csv module with a parse. You cannot just do:
csv.parse(...)
you have to first say import csv
This jaw-droppingly moronic.I will not Go lang has the same feature carried forward from C. It helps a lot in the reading code side of the code lifecycle. And Go compiler makes you keep the imports up to date, which is good.
It lets you debug Python problems which the system created in the first place.
> If they have some weirdly screwed up paths with multiple pythons installed and multiple copies of the modules etc., same.
Doesn't happen in a sane language. Or, even not a sanely defined language/implementation.
I can easily have multiple different GCC copies (possibly for different processor targets) on the same machine. Each one knows where its own files are; an #include <stdio.h> compiled with your /path/to/arm-linux-eabi-gcc will positively not use your /usr/include/stdio.h, unless you explicitly do stupid things, like -I/usr/include on the command line.
It can be slightly inconvenient but doesn’t feel moronic to me. It means that except for the built-in functions, everything can be traced to either a definition or an import. Makes tracking code much easier.
from python import def # now you can def
That should be even easier to track things; now you don't have to deal with the difficulty of def not being defined anywhere in your code. It's traced to an import, which is telling you that def comes from python, liberating you from having to know that and remember it.And backslash doesn’t let you have the literal obey the proper indenting. Might as well use “””
I don't want to be finding definitions of things that the language provides in the code.
Languages that don't work this way have IDE's, editor plug-ins or other tools for easily finding the definitions of things that are in the language, without hunting for them through intermediate definition steps in the same file.
"I've spent all my life in and out of jails, so I expect bars on doors and windows ..."