Whoosh – Fast, full-text indexing and searching library in Python
bitbucket.org
bitbucket.org
For example, you could use Whoosh for development environment and Elasticsearch for production if you need something more robust.
Whoosh is fine for very small datasets (megabytes) and low loads, just watch out for file permission issues.
http://elasticsearch.com/product/ ("petabytes of data with ease")
You won't convince them, they've already jumped off the cliff. It's all perfectly normal to them. If you want to do development, find a company that hasn't been contaminated yet.
Please educate me, what am I missing out on?
I have two responses to that, which may seem contradictory but are both heartfelt:
1) Your company uses Maven? You're lucky!
2) Maven? The prosecution rests.
> Please educate me, what am I missing out on?
You're happy. I'm obviously just another annoying, delusional ops guy who doesn't know what he's talking about. Why would you think you're missing out on anything? Go about your business. Everything is just fine.
But it sounds like your problem isn't Java/JVM, but rather really shitty code. If it breaks when you change minor versions of the JVM, the code is broken.
At some point you reach the conclusion that the problem with an ecosystem is more fundamental than a few lazy developers. But even if not, your statement isn't actionable. If all the code in a certain set happens to be really shitty, the only rational choice is to not use that code. Any other argument ends up being one only of semantics.
You'll have all seen the pattern. The sort of people who use language X tend to be the same people who use VCS Y tend to be the same people who think doing Z is a good idea tend to be the same people who don't see any problem with IDE W and protocol V.
Culture (by definition) is something that propagates amongst a community, and sometimes it's just poisonous. Or at least seems poisonous from an outsider's point of view. Except if it's java, in which case there's no two ways about it.
As others have posted, this is indicative of bad code. Java can certainly leak memory in the form of retained live references.
> Of course you shouldn't be using version 1.2.3.4.5, our code only works with 1.2.3.4.4, everybody knows that.
That's completely independent of the language or runtime being used, and is purely a project management issue (which breaking changes go into which version).
> that only runs on an obscure 6-year-old JVM
If there is code that imports classes specific to a JVM (com.sun, etc), then that code is doing something pretty much universally agreed on in the Java community to be a 'bad thing'. Otherwise, bytecode from Java 1.0 still runs on the latest JVM without issue.
You can use any language or runtime badly. There is a lot of code out there written in Java, and a lot of 'commodity' Java developers writing it. Of course there is going to be a greater volume of bad code, that's not an indictment of the platform itself.
That passage was about the runtime. At one point, I had to deal with four different JVMs in the same server farm.
> You can use any language or runtime badly.
No other ecosystem has so consistently fucked me over. Not even PHP. I actually prefer Java to PHP for just writing code, but PHP to absolutely anything running on the JVM for operations.
Know when I stopped regularly getting woken up in the middle of the night? 2011, when I stopped supporting anything on a JVM. Ironically, it would have been two years earlier, but I foolishly self-inflicted Cassandra on myself. Never again.
With Python this is usually very easy.
I've maintained my fair share of legacy code and more often than not one of the biggest issues is tracking down ancient libraries which work with my code. There are very few packages I trust to maintain fairly stable interfaces and be available years into the future (things like Apache HTTPd, PostgreSQL and memcached); everything else lives in the project lib folder. I can "hg clone" my app at any time and get the exact version of the third party packages I need.
Every so often I update the packages and test the integration.
Looking now, it seems that pip can finally do this with pip-compile.
edit: I don't mean to say that it is a selling feature for everyone (i.e. people who aren't already using Python) but for those who are already using Python it definitely is.
I think that would be a very difficult argument to make.
Also, a pure-python solution allows for an in-process index/search, rather than adding another external process dependency which needs to be monitored and maintained.
It's probably the best first iteration search app that I've come across, and you can always slip in solr or something with more umph when you need it.
I think this makes it pretty clear. Not sure how much more explanation you need.
When document counts get into the hundreds you see orders of magnitude faster queries with SOLR (etc), not to mention much more sophisticated querying options.
The documentation indicates there is a multiprocessing option. (Also configurable memory limits, if you didn't see that.)
FWIW I'm using Flask, Flask-SQLAlchemy, Whoosh and Flask-WhooshAlchemy [0]. Quick to get going with but the mix-and-match approach lets me easily rip pieces out as the project grows.
To put it in context, the psycopg2 calls to PostgreSQL take about 100ms to retrieve the associated data once I've found my search results with Whoosh.
(Most of my response time sits with SQLAlchemy ORM, building up matrices of data, which is why in production I'm caching the more complex queries with memcached).
Overall, for a project of this size (I can't imagine having to index more than a few hundred thousand objects) I'm very happy with Whoosh. If I get to the point of indexing millions of objects, I'll optimize it then.
It seems to be designed for indexing manpages. That is, a medium-sized semi-static database with a few different dimensions per document.