Python Libraries you should know about
doda.co
doda.co
> I specifically excluded awesome libs like requests, SQLAlchemy, Flask, fabric etc. because I think they're already pretty "main-stream". If you know what you're trying to do, it's almost guaranteed that you'll stumble over the aforementioned. This is a list of libraries that in my opinion should be better known, but aren't.
https://gist.github.com/4061368
All it does is grab paragraphs from python.org's html a couple thousand times.
==== Total trials: 100000 =====
bs4 total time: 31.6
pq total time: 9.3
lxml (cssselect) total time: 5.4
lxml (xpath) total time: 4.3
regex total time: 8.9 (doesn't find all p)
What does it mean? Unless you're running thousands of queries for parsing, it doesn't matter which library you choose. My computer old and slow. Pick which one is the easiest, that you'll fight the least with. Don't put energy into unnecessary optimization. Using a good library is like choking someone, they'll fight for a little while until they pass out. (you'll remember this analogy next time you want to switch libraries, do you really need to choke someone to get your job done?) After they pass out, it's smooth sailing and you don't have to worry. Don't rock the boat unless you have to.- Python.org is quite a simple web app, it would be interesting to run this against something a little more complex like a long wikipedia article or http://www.nytimes.com/ or even the Alexa Top 500
- It would also be interesting to split the times between parsing and selecting as I feel that's where the difference between pq and lxml comes in.
All in all it seems BS4 is quite a bit faster than I gave it credit for (especially factoring in parsing+selecting instead of just the former)
==== Total trials: 100000 =====
bs4 parsing time: 0.6
bs4 selecting time: 146.3
pq parsing time: 0.0
pq selecting: 15.7
lxml parsing time: 0.0
lxml (cssselect) selecting time: 12.4
lxml parsing time: 0.0
lxml (xpath) selecting: 11.5That one is going into my permanent memory file.
Could probably do a top 7 just by Reitz:
From: https://github.com/kennethreitz
1. Requests
2. CLINT - Easy CLI tools inc cross platform colour
3. Envoy - "subprocess for humans"
4. Tablib - csv, excel and plenty others, tabular data
5. python-guide - A work in progress book
6. dynamo - Amazon Dynamo as a python dict
7. gistapi.pyI specifically excluded awesome libs like requests, SQLAlchemy, Flask, fabric etc. because I thought them too "main-stream". If you know what you're trying to do, it's almost guaranteed that you'll stumble over the aforementioned. I tried to compile a little bit of a list of libraries that SHOULD be better known, but aren't.
Thanks a million.
It makes it very simple and intuitive to build command line apps.
For example for:
define_opt('server', 'daemon', type=bool, cmd_name='daemon', cmd_short_name='d')
Can I just write define_opt('daemon') ?
server - this is the section/module. I could grab this from the name of the current module, but that means you must use unique module names, and not change them. Otherwise, your config files would stop working.
daemon - required, obviously, since this is the name of the option.
type - I can't assume a type for you. By default it is a unicode instance, so you can omit this parameter if that's your use case.
cmd_name/cmd_short_name - I can try to assume that cmd_name is the same as the second positional argument, but once again, there could be a conflict. For example, if you want to have options.db.filename and options.log.filename, you can't use cmd_name = 'filename' for both. cmd_short_name is even worse, since here you may want a specific letter to be used (such as an upper case D instead of d). Note that these parameters are optional, since most of your values will likely go in the config file, not on the command line.
An alternative API might look like this:
define_opt_int('server', 'shutdown_timeout')
define_opt_bool('server', 'deameon')
define_opt_cmd_unicode('server', 'pid_file', cmd_name='pid', cmd_short_name='p')
With a fallback to: define_opt('log', 'level', type=lambda x: x if x in LOG_LEVELS else 'DEBUG', cmd_name='log-level', cmd_short_name='L')
Any feedback is greatly appreciated.Interesting module nonetheless!
pq_page = pquery(url=PAGE_URL)
Note that PyQuery has some encoding issues too (or rather the sites I were scraping were too bad, showing two different encodings in meta tag!), here are two different things I have done to workaround: page = requests.get(PAGE_URL)
pq_page = pquery(page.text)
If that doesn't cut it (because requests detects it wrong too), try forcing the encoding in requests: page = requests.get(PAGE_URL)
page.encoding = 'utf-8'
pq_page = pquery(page.text)[0] https://github.com/dsc/pyquery/blob/master/pyquery/pyquery.p...
I know matplotlib comes with Python(x,y) but that's a pretty awesome one too.
EDIT: I added it below the `parse` example. Thanks again!
I'd never heard of pattern before, and while it looks like it's a nice bundle of features, I'm concerned by the fact it references pyWordNet by name even though it hasn't been an independent project since 2006 (http://osteele.com/projects/pywordnet/). Has anyone actually used it?
pyquery equivalent: Nokogiri (http://nokogiri.org/). Lets you select elements with jQuery-like selectors. Uses libxml2 as its parser.
watchdog equivalent: watchr (https://github.com/mynyml/watchr). Run code when the filesystem changes.
path.py equivalent: rush (http://rush.heroku.com/). Provides a far better API to the filesystem than the standard library.
I also found this equivalent to fuzzywuzzy, but I’ve never used it: amatch (http://flori.github.com/amatch/)
Python Imaging Library - Today's web is full of images and PIL makes it easy for image manipulation. Although, it's not extremely performance efficient at very large scale.
http://code.google.com/p/parsedatetime/
it seems to accept syntax similar to the 'at' command does (and obviates the need for my python C module to do that parsing based on the scheduler parser for 'at'). examples include "1 day ago", "ten hours from now" and the like. very useful.
The docs got my hopes high, but they fail to mention sh isn't supported on Windows -- I saw that by looking at its source on GitHub. I'll take a look at pbs, but this kinda bummed me out.
Of course, I could simply fork it, remove the "not supported on Windows" check and use it, but that feels lazy and dirty ;)
From a quick look at the code, it's not supported on windows because of its reliance on terminal utilities(pty, termios). Doesn't the banner state to use the previous version(pbs) for windows support?
https://github.com/amoffat/sh/blob/master/sh.py#L33
And if for some reason it still doesn't work out, you can always use subprocess(which this library is using anyway) or https://github.com/kennethreitz/envoy
Thanks! Like I said, I don't have lots of experience with Python, so this is very valuable info for me :)
Doesn't the banner state to use the previous version(pbs) for windows support?
Yes, but the PyPi page for pbs 0.110 states: "PBS will no longer be supported." I'm not comfortable with basing my code on a library which won't be supported in the future.
And if for some reason it still doesn't work out, you can always use subprocess(which this library is using anyway) or https://github.com/kennethreitz/envoy
I didn't know about envoy either. Thanks again! :)
EDIT: Upon closer examination, the whole OProc class depends on os.fork(), which is available only on Unix, according to Python docs.
for example, see this bug on so - http://stackoverflow.com/questions/10575919/strange-date-par...
in short: it's ok for non-critical cases where you just need "some date" from input. but don't use it if you would rather have an error than an incorrect interpretation.
>>> path('a') / 'b' / 'c' path('a/b/c')
It would be fun to have that in Ruby!
Path.py may already do it (docs are limited so I'm not sure) but something like below would be better:
my $foo = dir('a', 'b', 'c');
This is how Path::Class works in Perl (https://metacpan.org/module/Path::Class). This way the path is a filesystem agnostic directory object.NB. Twisting operator overloading isn't all bad though. For eg. IO:All nearly gets it all right and being a little mad I will sometimes use it :) https://metacpan.org/module/IO::All
>>> print SimpleFS('a').b.c
a/b/c
The bigger thing (and original purpose) was uses like this:
>>> root = SimpleFS("/")
>>> sda_size = root.sys.block.sda.size
>>> for fd in root.proc.self.fd:
... os.close(int(fd))
>>> root.var.log[sys.argv[0] + ".log"] += "Size of disk: %d" % sda_size