Greplin opensources Lucene Utils and Bloom Filters
tech.blog.greplin.com
tech.blog.greplin.com
This is really useful and painfully lacking from Lucene. Great stuff.
[1] https://github.com/Greplin/greplin-lucene-utils/blob/master/...
Anyway, for prefix matching, you want to take careful note of the size and shape of your data, because it matters for which algorithm is fastest. If your data all fits in RAM (or better yet, all fits in L2 cache), then I've had very good results with binary search (O(log N)) to find the first matching result, and then linear scan to find all possible suffixes. This is a lot more cache-friendly than radix trees, which have better theoretical performance but often touch memory that's all over the place.
Yes, thank you.
2 things: Your account is exactly "1337" days old today, happy leet day :). Also, are you the same zem that posted on newslily a while ago? If so, you posted some really awesome stuff, thanks :)
(Sorry if I've recognized you and asked this question before, I have a terrible memory)
I wouldn't really say it's still going strong, haha. Unfortunately, we never really got the traction on that that we needed to allow it to keep running by itself (Cody and I were submitting a lot of the content, which was fine, but it would have been really awesome if we didn't have to).
<sarcasm>Big surprise there, though, we were trying to compete against HN and reddit</sarcasm>
Back in November, we both started working a new project: http://thingist.com, which has been taking a lot of my time lately, so I haven't been submitting as much.
Anyway, good to see you around here, man :)
We open source projects for two primary reasons:
1) Because we use lots of open source projects internally. It makes sense to reciprocate where we can.
2) There's no better place to meet great engineers than in GitHub pull requests.
https://github.com/jaybaird/python-bloomfilter - offers scalable bloom filters
https://github.com/axiak/pybloomfiltermmap - uses mmap
I'm asking, because we're getting great mileage out of mmap in Clojure, albeit for read-only mappings. And that's in a search engine :-) Using mmap for large data is great, because you avoid enlarging your heap and the garbage collector doesn't even have to care about your data.
Eventually, we needed more flexibility than Solr easily offered though. For example, we've added far more efficient sharding, document modifications (updates and deletions), flushing, and near real time search than either Lucene or Solr support out of the box (and they were much easier to add to Lucene than Solr, since Lucene makes fewer assumptions about your dataset/use).
I think Solr is a great tool if your needs happen to fit into their model - but if they diverge a lot, it sometimes makes more sense to build your own custom framework on top of Lucene.
We didn't use them outright since we have fairly different requirements/constraints (our data has some pleasant properties that makes it easier to shard and facet than the general case) and we wanted something a bit simpler.
Shoot me an email sometime though (email in profile)! I'd love to buy you lunch and pick your brain ;-) You guys clearly know what you're doing!
Publishing 2 github repositories, each with several classes which are mostly trivial, is not "opensourcing".
Much like the fact that a weekend project is not a "startup".
I'd imagine this falls in to that category. The takeaway for me is that the code is useful, regardless of the size, and they took the time to let other folks enjoy it.
But I've seen much larger open source patches - that are no less important - that have received much less publicity that this.
Not to disparage the Greplin team's work -- I think it's totally awesome that they keep open sourcing pieces of their infrastructure. We should all be doing more of this.
Err, yes it is. If you are going to try and be pedantic then at least be accurate.
A more accurate version of your complaint would be "They shouldn't get so much attention for open sourcing a few classes". That argument has some merit, but actually the bloom filter library is something that Java has been missing for a while (I know, because I've written one in Java myself).