Of course, performance won't be stellar if it has to go to disk to find every result, but even with that you should be able to get sub-second results with some tweaking.
I guess I assumed the entry-point is a 2 to 4 node cluster ? But this may be my own wrong assumptions.. they may be really ram efficient.
I confess I have a somewhat irrational dislike of what I call big-java .. I really wanted to like Cassandra+Spark for instance, but the install is just such a heavy download of code, I just feel an aversion to installing that.
I have worked with Java, but not recently.. there may be incredibly efficient and small codebases that belie the big-java stereotype.
Iv'e been enjoying the node.js / npm ecosystem myself - but I know smart people who love java-land. We can tolerate each others rants and still share a beer :]
The Lucene-based search engines are also very well optimized. I'd tend to think of a 30M row dataset as a single machine size unless you need failover or the records are quite large.
Wow, that is definitely the opposite of true.