Very cool project. Which algorithm do you use for indexing/retrieval?
Right now it is not very complicated. An inverted index of the text that is sent over is created and added to the database. The DB is just a big hash map. Searching right now is just an O(1) lookup in this hash map of the text being searched in this hash map.
I have plans to improve this of course. There is a lot of room for improvement. I am also planning on adding scoring using TF-IDF, but it's not done yet.
Any help with this would be more than appreciated!