254 karma · joined March 14, 2012
I was just commenting how row-limitations were a very NoSQL way of enforcing data limits and inappropriate for an RDBMS.
post.categories (varchar) = "1,51,78,84,100"
and other wonderful non-normalized approaches.No, no at all.
Co-Founder: $33k Salary, 45% Equity
Co-Founder: $33k Salary, 45% Equity
You: $33k Salary, 10% Equity
I'm hoping the problem is self-evident.In 2008/2009, another engineer and myself built an ad-platform that received around 500M impressions per day, 5M clicks per day. And it wasn't just recording a tweet or publishing out to followers. We took the user input query, had to do some keyword/relevancy targeting, geofiltering, matching to advertisers and deliver back a large result set of adverts. All within 100ms.
Our platform was also apache, mod_php, memcached, mysql and rabbitmq. So definitely not the most optimal of platforms by any means. We had two colos with ~20 servers (dell r410s) at each facility.
Twitter just recently announced 400M tweets/day. I'm not trying to brag about my experiences, because looking back now we made numerous amateur mistakes, but just showing that Twitter's "scale" is a joke compared to everyday challenges at any large internet ad network.
Can we stop labeling the set of "not a rdbms" data storage mechanisms with the stupid fucking "NoSQL" moniker.
I suppose that might work pretty well.
I dunno I can't envision Solr being more efficient than a properly designed RDBMS for these situations. If you were integrating a full-text search I'd absolutely believe that to be the case but...
Until them, I'm quite content with one of their competitors.
Does anyone have particular insight to share on this? Last I checked, Solr's geospatial searching methods are rather inefficient -- haversine across all documents, bounding boxes that rely on haversine and Solr4 was looking into geohashes (better but have some serious edge-case problems where they fall apart).
Meanwhile PostgreSQL offers r-tree indexing for spatial queries and is blazing fast.
Am I missing some hidden power about Solr's geospatial lookups that make it faster/better than an r-tree implementation?
Obviously languages can serve multiple purposes but the intended uses of these have almost no overlap.
Java itself can be a beautifully terse language.