Stonebraker Explains Oracle's Obsolescence, Facebook's Challenge
blogs.barrons.com
blogs.barrons.com
This.
So I'm dealing with this problem. There is nothing out there.
Neo4J doesn't really do in-graph processing[1]. BlazeFB/OrientDB/RDF Stores all are similar to Neo4J
Pregel/GraphX/Giraph are graph processing engines, but lack property stores.
I want a single system that does both. I want to run PageRank (etc) and query-by-property on the same system.
Titan was promising, but they stopped working on it when they were bought.
I'm surprised no one is fixing this.
[1] http://neo4j.com/blog/categorical-pagerank-using-neo4j-apach... note this bit: "I can scale each Apache Spark node to perform parallel PageRank jobs on independent and isolated processes all consuming a Hadoop HDFS file system where the Neo4j subgraphs are exported to." (ie, Spark runs against HDFS, not Neo4j)
Was selcted by wikidata when titan was bought.
Sure Facebook might have preferred a better solution, but most of us aren't Facebook, and the solutions that would work for Facebook might not be ideal for small companies.
In other words, just because our data sets are exponentially larger doesn't mean we still don't need to be able to perform complex joins or other relational operations. The data sets are more complex, but the queries we're running on them are even more complex. While some other data stores may be optimal for niche use cases, SQL offers the best mix of speed and flexibility while still meeting the business requirements.
I remember them from ages ago as basically an XML focused database. I'd see them at trade shows with a small booth. They had interesting technology but had very engineer-y marketing and not a lot of customers. My company had been on the lookout for such a technology and IIR our engineering team checked them out for a bit, but they weren't a great match for our product and we ended up using dtSearch instead. [2]
In the last year or two, I've started hearing MarkLogic show up again all over the place and wondered what was going on. Turns out they got new leadership (Gary Bloom) and have been making a big push to grow. It's funny how that happens, I wonder how many other serviceable companies with decent tech are hiding out just waiting for the right CEO to come along and push them into the spotlight (I'd also add that 'XML' is no longer one of their marketing keywords).
1 - http://blogs.barrons.com/techtraderdaily/2015/02/13/oracles-...
https://gigaom.com/2011/07/07/facebook-trapped-in-mysql-fate...
That aside Facebook having a buy versus build decision to make seems pretty strange. What would they even buy given that Cassandra and HBase which they created are two of the most scalable databases right now. Strange observation.
Hadoop is not what the article is talking about though, nor is Cassandra. The issue is not about whether it scales out or not, it's about whether it is memory-based or disk-based. If queries have to pass through a block API at any point, they can never perform as well as a database where all queries run out of main memory, like Hana or VoltDB. On block-based systems you can't run some categories of analytical queries interactively, they always end up being too slow.
He's saying oracle has a problem because they don't have a good in-memory story ready. Right now they're aiming at this problem from two directions: Oracle NoSQL (a distributed KV store comparable to Riak), and Oracle 12c In-Memory (in-memory engine bolted onto the oracle db). Neither are particularly convincing to me, and definitely not a match for Hana or VoltDB.
My current employer is making the switch from Row based to Hadoop - which I feel is because of hype & not justified technically, given the size of our cluster. The goal is to reduce speed of data delivery to clients, but I believe a column-based DBMS with an optimized ETL, would be the way to go.
Wonder how it'll look like in 5 years & if others companies' IT are buying into hype too.
Context from the interview: Sooner or later, the business intelligence world will move to the data science world, using things like regression analysis, Bayesian analysis — these are lots of big words, but all of these techniques, if you look at them, it’s an array-based, not a table-based calculation.
But the article at the link you suggest doesn't even contain the word "array", so I'm still no closer to understanding how this concept differs from traditional tables.
I don't know if anyone has implemented a database optimized for array processing, but the benefit for predictive modeling should be obvious. I'm sure optimization for GPU processing wouldn't hurt either.