Google's Python style guide
google-styleguide.googlecode.com
google-styleguide.googlecode.com
for eg.
from sound.effects import echo
echo.EchoFilter(input, output)
What happens is that you then end up importing all of these methods and very quickly you start getting name conflicts. Lets say you want to support a third-party echo function: from sound.effects import echo
from vendor.soundutil.effects import echo as soundutil_echo
You see this all the time in SDK and web API packages. Dozens of modules called 'auth' (which auth? twitter? facebook?) or 'oauth' or 'request'.Lets say you have a user page that integrates with social networks, would you rather:
from facebook.api.auth import auth
from twitter.api.auth import twitter_auth
etc. etc. or import facebook.api
import twitter.api
..
facebook.api.auth()
twitter.api.auth()
You end up either doing 'import as' hacks and a lot of renaming. Code is a lot clearer to read when you see full method names such as facebook.api.auth rather than just 'auth' and 'echo' everywhere. You also don't lose documentation paths.My general rule of thumb is to use 'from' infrequently, never do import *, to retain the part of the path that still keeps namespacing sane and clear to the developer and as the doc says to never do relative imports.
It means you can scan any part of the code and understand what is going on without going back up to the top of the file. Also makes search/replace easier (rather than s/echo/echo_new s/sound.effects.echo/mynewpackage.echo)
The other one I didn't see mentioned is nesting levels and method lengths. Python isn't well suited to deep-nested and long methods. Especially if your coding style is to comment out blocks of code during development as you test things, you always end up commenting out parts and then having to re-indent the rest of it.
The same usually applies if you have long 'and' 'or' clauses in ifs that span multiple lines and make it harder to understand the code. I usually wrap those tests into separate methods (if you are using them once, you will probably use them again)
but for nesting, I try to stick to 2 levels max. If you go beyond that it is usually a hint that you can refactor the codepath and perhaps even separate out into another method.
I just happen to be doing this a few hours ago while writing an option and argument parser for a command line utility that has sub-commands. a quick re-factor made the code and all the different options and which options apply to which sub-commands a lot easier to understand
Edit: just further on breaking up code and bounds checking into methods, it makes life easier for other developers and for your future self. there is nothing more exhausting than trying to debug a module and finding a 3-page long method called 'run', which you end up having to break down yourself anyway. separate all the bounds checking into one or two line methods, break everything else up, document it, write some tests for it and then forget about it - that is done and it works. get on with important things.
checking nesting levels and method length is almost something I would want to put in a linter
What I don't understand is why they prefer:
from thepackage.subpackage import amodule
over import thepackage.subpackage.amodule as amodule
I always found the second to be more clear, since it doesn't mix the idea of package-resolution with the idea of picking-things-out-of-a-module. On the down side, I end up typing "amodule" twice.Really, Google? I find the following far more convenient to read:
result = [(x, y) for x in range(10) for y in range(5) if x * y > 10]
Than the alternative: result = []
for x in range(10):
for y in range(5):
if x * y > 10:
result.append((x, y))Also, i think part of the difference can be found when you're dealing with large amounts of code in maintenance. I'd much rather read something dead simple (if somewhat verbose) than something that makes me think at all (another example here would be (in the ruby world) use of !unless).
Like anything that has to do with taste, to each their own (i prefer a vinegary bbq while you might like a smokier bbq)
result = [(x, y) for x in range(10)
for y in range(5)
if x * y > 10]
Anyway, I'd call it a simple case.Also, it has nice semantics - the second example has a clear execution order. This one doesn't - it can all happen at once, from the program's perspective and unless you make assumptions about order in which the items are calculated (as opposed to returned), they don't even need to be all ready before you start iterating on them. If you assume it can happen at once, the list comprehensions can be neatly mapped to parallel computations.
So, it shouldn't be that hard to optimize something like:
data = [sin(x) for x in arange(0, pi, pi/20)]
to run on a GPU.edit: small clarifications
I'm not sure what you're getting at here. With Python list comprehension semantics, those are both the same, including execution order. I don't see any ambiguity or need for "assumptions about the order of results". Am I missing something?
Optimization of Python list comprehensions to run on a GPU would take some serious mojo to ensure independence of each clause. Not an impossible amount, but certainly not trivial, especially as you move beyond calling 'sin'.
When you use a list comprehension, you may (or not) care about the order of the resulting items, but your program is completely shielded from the order in which the resulting list is calculated - unless your function is affecting a global state while the LC is being evaluated and each evaluation depends on the state changed by the last one. You can't insert a print in the outer loop, for instance, unless you explicitly nest the LCs.
With the map builtin deprecated because of LCs, wouldn't it make sense to exploit concurrent-like behavior with the nicer LC syntax? It makes a lot of sense to execute strictly in order for generators, but it doesn't make that much with LCs and I never saw code whose correctness depends on strictly ordered causation of side-effects.
Obviously, being that much multiprocessor-friendly makes little sense under a GIL, but CPython is not the only implementation of Python.
Python is and probably ever shall be my favorite language of the OO-imperative 20th century/first decade of the 21st century style, but it will not be making the leap to the next generation of languages. And I sort of hope it doesn't even try; better to be the best of breed imperative-OO than a half-assed hybrid that does nothing well.
If I care that (3, 6) comes before (4, 4), I find it much easier to ascertain that from the form that uses indentation for nesting loops.
[(x, y) for x in range(10)
for y in range(5)...] # TODO(qznc) check Unicode handling
This never occured to me, but it provides a good pointer whom to ask for details. On the other hand git-blame should be able to provide the same info.Challenge accepted:
for i in $(git grep -l TODO); do git blame -f $i |grep TODO |grep "$NAME"; done
Granted, it can be long on a big codebase, but does the trick.In short: metadata which is not updated automatically is most likely not up to date, and may be not correct in general.
#TODO #4001 needs more bananas
Where 4001 references something in our bug tracker (task, feature, issue, bug, etc.)It could be far worse.
One exception is some projects are strict about lining up code properly in multi-line statements, and spaces are more consistent in that respect.
I prefer tabs, but most Python code I've seen has been 4 space indents.
Community consensus is 4-spaces and that's it.
I used to prefer tabs as well, and I still stand by it.
But you can always configure your editor to use 4 spaces. And 4 spaces looks like 'less wasted space'.
I'll keep using tabs for some things, and 4 spaces to most professional projects.
About 2 spaces I'll just say one thing: NO
1 - yes, screens are huge, but it doesn't mean people can/will use small fonts or will scroll the screen
With today's big/wide screens it's more useful to have code side by side.
2 - Abuse. The 80 character limit is a pretty good indicator that you should be doing something else instead of having your code go over 80 characters.
Long lines are confusing, and you most likely can split the logic in several lines, facilitating maintenance.
I think having a soft and a hard limit makes more sense, if anything, I would make 80 characters the soft limit and perhaps 100 a hard limit, although I'd prefer them to be 100 and 120.
However, I've actually found it to be a good thing - I'm sure it aids readability.
That will definitely break stuff. Why not #!/usr/bin/env python?
their goal isn't to write portable code, it is to write fast code that runs on google servers
http://git.savannah.gnu.org/cgit/coreutils.git/tree/src/env....
Could you please explain what am I missing?
One significant difference: when running python from a shell there's a fork() and an exec(). env(1) doesn't fork: there are not two processes. (In other words: the shell does not exit when you run a command, env vanishes).
There is obviously still some overhead to using env (and I just learnt from the source that env has argument processing). I tried to replicate AncientPC's test but on my machine both invocations take around 0.015s. (Perhaps their username is an indication as to why they see a (2%!) difference...).
But okay. What I really came here to say is: security. You can make all the efforts in the world to ensure that all programs get called with full pathnames but then one env shebang and you're suddenly open to running whatever's first in the user's $PATH and happens to call itself "python".
EDIT: eg http://portaudit.freebsd.org/d42e5b66-6ea0-11df-9c8d-00e0815...
I notice the difference anecdotally without measuring it, but I don't know if that is a conception bias because I know env should be slower.
The best way would be an autoconf script in your package and an install run that finds and verifies the local framework.
I have to admit that I have never done this though. I have a few Python scripts with decent distribution and just rely on the direct path (and a batch file for win32)
I've been using this handle since '93 out of inertia. My tests were done on an idle i7-620M, sequentially.
If your invocation takes about 0.015s, that means you're not looping enough for differences to appear. A 2% increase from 0.015s is 0.0153s, invisible due to significant digits cut off.
There is ~2.1% overhead when using env.
This doesn't matter in most cases, but in a Google-sized company with standardized environments it's worth it.
#!python
Not to use the user's PATH, but some system ordained python (via some kernel work).
env is a hack.
Kernels shouldn't be aware of the search path shells use.
http://google-styleguide.googlecode.com/svn/trunk/pyguide.ht...
I commented about this on another post. There's a nice explanation here as to why using mutable objects as default values in function/method definitions is bad:
http://effbot.org/zone/default-values.htm
In short, it can be bad to set default arguments to mutable objects because the function keeps using the same object in each call.
What kind of SyntaxErrors are cought by the except: handler? Not all, I presume:
try:
a b
except:
pass
This fails with SyntaxError on Python 2.7.2 on my machine.For example,
try:
eval(".")
except:
passhttp://google-styleguide.googlecode.com/svn/trunk/pyguide.ht...
Oh, not different map/reduce? ;)
OTOH, map() does otherwise have a little overhead compared to list comprehension.