Riot – Full-text search engine in Go
github.com
github.com
I'm in bioinformatics and the first time I implemented a wavelet tree to reduce the size of genomes in memory.. It was just breathtaking.
I need quick fuzzy search on a low-end embedded device that has limited storage(both RAM and HDD), was thinking about putting the index on a server with plenty RAM then do websocket or RPC for that.
Now to go with Wavelet trees you may or may not need to know about suffix arrays and optimal suffix array construction. Take a look at this: https://en.wikipedia.org/wiki/Suffix_array This is what's going to give you space efficiency in combination with a wavelet tree. And the wavelet tree also gives you good rank/select efficiency.
Edit: Here's a suffix array construction algorithm implementation I did (not sure if it's fully correct) https://github.com/ethanwillis/comp7295_final/blob/master/sa... It is based on this paper: https://local.ugene.unipro.ru/tracker/secure/attachment/1214...
The existing stuff isn't slow.
To be clear, I've never used the language. I actually dislike programming, though I've decided to get back into it because I have a couple of projects I want to poke at. I'm now deciding between Java and Python.
Maybe I should do an 'ask HN' submission.
I've narrowed it down to Python or Java. I've done C, C++, BASIC, QBASIC, Perl, PHP, and even some COBOL. I've played with a few others.
Have you thought of learning a new programming paradigm? Why not check out Elixir or Rust?
I learn best from people with exacting standards.
If you're going to eliminate languages for that, you're basically out of options.
I think Python is the much better choice to start out, but Java is pretty great too. It's hard to find languages that are actually bad choices... maybe COBOL.
I hired professionals in 1995. I was done doing any of the coding by 2000. I sold and retired in 2007.
I took only one course in C. Everything else was learned on my own, informally.
Most of the times all your questions will be answered by a simple google search, as there are mountains of good questions and answers on stackoverflow.
IMHO for you, getting back into programming will be as simple as opening the interpreter and starting to type.
I don't think programming has changed, the basic mentality is still the same. And nothing beats good old experience.
The new niche "trends" such as asyncronous programming, actor-based programming etc. could be easily learned by lingering on HN for a while :).
Best of luck getting back on the keyboard!
I'm in the process of building my own search engine (as a learning exercise, but also because it's related to my day job). I've learned that it's one thing to write a full-text search engine, like this one, and it's quite another to do field-specific searches with faceting support and so on, like Algolia and Lucene-based search engines do.
That said, this is clean and simple. I like it. I can definitely learn from this.
Tracking (document, field) values can be used for query by range or by geolocation primitives (that's what Lucene does, where it will index that data into a special tree-like structure, and for each query, it will build a custom 'iterator' and use it along with other iterators to match documents), and for static ranking of matched documents.
BTW, Lucene and Algolia are vastly different in terms of the underlying architecture.
As Mark mentions on his summary page, the best place for that kind of information is our CTO's "Inside the Engine" series (8 parts).
https://blog.algolia.com/inside-the-algolia-engine-part-1-in...
If so, then Sphinx could be a deal breaker to some.
You can connect to Sphinx using a MySQL client, use it as a MySQL storage engine or using MySQL as a data source. But it's not specifically tied to MySQL.
Beyond being a user of both I don't have affiliations with either of them.
Possible reasons why the original project stagnated: [0].
Mention of the manticoresearch as a fork project was removed from the sphinx forum [1] - so I can guess that developers who moved to the new project did not part on good terms with Andrew - the original author.
that's the last reply in that thread from person involved in Manticore, posted on the 23rd of Oct 2017:
" aditirex just replied to 'Sphinx search fork':
===cut=== > But your the people who are already using sphinx, why we should change?
The open-source version of Sphinx received 5 code commits since November last year, from which 3 are related to building stuff. Last release was 12 months ago. There are also a lot of unresolved reported bugs (many of them are crashes) in the bug tracker. Andrew said a while ago that the open-source version would only receive fixes (which doesn't seem to happen either). No one wanted to do the fork, it was the only way several big users saw it in order to continue using Sphinx. Don't ask me how we got into this situation, I'm not the right person to answer to that.
> What are the main benefits rather than using Sphinx?
Manticore is pretty much continuing Sphinx. Last year we had 4 developers + Andrew working on the code, 3 of them are working now on Manticore.
If Sphinx just works for you there is no reason to switch. But we're adding new features, fix existing bugs, the software is tested by some big users before getting released, you get a software that has support from it's developers.
> Is Foolz\SphinxQL\SphinxQL working?
Everything works as before, it's a fork, not a total new software. "
[0] http://sphinxsearch.com/blog/2017/07/24/sphinx-2017/ [1] http://webcache.googleusercontent.com/search?q=cache%3Ahttp%...
I'm looking for a light weight elastic search alternative.
> RESTful search server written in Python, powered by SQLite.
Which I'm pretty sure can be embedded in Solr, has plugins for Elasticsearch and others.