Lucene: The Good Parts
blog.parsely.com
blog.parsely.com
Doug Cutting worked on VTwin at Apple long before he wrote Lucene. This was a codename for Apple's search technology they were building into the OS. I knew about it because at the time I was building computer labs for my college and bought a lot of Apple hardware (about a million dollars worth) and my rep added me to the private beta.
Many years later, when I saw the Lucene project being worked on and saw that Doug was behind it, I immediately reached out to him and asked him if he was interested in joining the Jakarta group, as I thought this would be a great addition to our growing community.
Needless to say, the rest is history, but I feel like if nobody had reached out to Doug, he may not have gotten as much exposure and may not have been motivated to start the rest of the amazing projects that he's worked on, including Hadoop.
=)
What does that even mean? SQL is a wire protocol for relational queries. Document/blob/key-value stores can be made to speak SQL. Most of them just choose not to implement SQL support for some weird reason.
I'll posit that most non-relational stores resist supporting any SQL because it would so quickly make plain how handicapped are their query and set-based capabilities.
As an example, imagine a single database where you want to store actual desktop documents, such as the formats supported by Apache Tika: https://tika.apache.org/1.8/formats.html -- If you try to model this using a SQL schema, you'll likely be in for a world of pain. From a UX standpoint, a user just wants to "search across all documents", but you have hundreds of heterogeneous types with varying degrees of field-level compatibility.
I have read that it doesn't scale out as nicely as document stores do, and knowing SQL Server I don't have trouble believing that. Personally I'm a very long way away from needing to worry about scale out[1] in the applications I use a DBMS for, though, so that's never really kept me up at night.
Word I've heard on the street is that the story's similar for PostgreSQL.
In discussions of "NoSQL" technologies, it's common to use SQL as a shorthand for traditional, relational databases which generally support SQL. We're talking about databases like MySQL, Oracle, Postgres, MS SQL Server, etc.
Sure, Andrew could have been more precise with this language, but from the context, do you honestly think he's using SQL to mean specifically Structured Query Language, the 4G language developed in the 70's?
My comment was specifically in response to your approach, not your overall point: > What does that even mean? SQL is a wire protocol for relational queries.
First, if we're being pedantic, I disagree with your usage of "wire protocol" to describe the SQL language, but I get what you're saying. What I was trying to comment on was your feigned confusion. It was perfectly clear to me what the author meant by "SQL" in the context of the article, and I think you understood too.
If you had said something like, "SQL is a language and bad term to describe Relational Database Management Systems (RDBMS), which is what the author really meant," I wouldn't have commented. I probably would have upvoted actually, because I agree.
Anyway, not a big deal, I don't think we're in conflict with the substance of your point, just how you chose to express it in this one example. Have a great day. :)
What use case do you have in mind?
1.store the field or document so you can get back the value (ex: when querying the document)
2.index the field for filtering
3.separately index the field for aggregation(doc_values)
While in rdbms you only need(1).