Do GPU-optimized databases threaten Oracle, Splunk and Hadoop?
diginomica.com
diginomica.com
Instead of putting 20 rusty bicycles together and claim to be a revolutionary fuel- and cost efficient rocket ship, why don't we build a rocket ship from the beginning that actually flies quite well?
Hardware and chips optimized for DB engines, queries and huge amount of streaming data will be welcomed.
It makes great sense to me as a techie, my users will be happy. My staff will be happy. I will sleep at night. There will be no surprises, everything will just work. We will not have to explain that "the cluster just doesn't do that well when it's 90% full" or "it's been unstable since we lost that node a month ago, we think that the new capacity will let it get to a better state when we put that in" or "there's a network problem, we don't know why, somethings interacting with something else and long queries are getting killed".
However, my CFO will never give me the money to build a rocket ship. NEVER. I have looked into the eyes of his successor and all his acolytes and they will never give me the money either. To make this happen the current corporate hegemons of HR & Finance must be thwarted and destroyed. We must take their significant others, ride their quadrupedal beasts and burn their smartphones. Only when the laments of their personal digital assistants swirl in my ears like a Bach requiem will I take a case to the Operating Committee to do things right.
Until then it would be suicide.
Do you mean the rate of change is slowing down, or that they're actually getting slower?
I'm not getting excited until Nvidia really ships something that keeps up with their promises. For the past few years, their products were underwhelming compared to marketing materials before launch.
There's more ways to improve than just faster cores.
Performance pr. core has been pretty stagnant, but performance pr watt has been improving. With the current consumer focus being so heavily on various battery powered 'computers', this seems like a reasonable trade off.
DBs are CPU-limited and, yes, single-core performance has been stagnating. But that's unlikely to be a vast conspiracy.
What are you basing this on? Database engines are only CPU-limited only if you have unlimited I/O bandwidth. Also, database engines hardly ever have only one query to process. Even if queries were not parallellizable (which they certainly are), single-core performance has little impact on overall performance.
Oracle has been doing this for years. The only problem is the price.
>SQL in Silicon: Adds co-processors to all 32 cores of the SPARC M7 that offload and accelerate important data functions, dramatically improving efficiency and performance of database applications. Critical functions accelerated by these new co-processors include memory de-compression, memory scan, range scan, filtering, and join assist. Offloading these functions to co-processors greatly increases the efficiency of each CPU core, lowers memory utilization, and enables up to 10x better database query performance. Oracle Database 12c In-Memory option fully supports this new capability in the current release. In addition, this new functionality is slated to be available to advanced developers to build the next generation of big data analytics platforms.
https://www.oracle.com/corporate/pressrelease/sparc-m7-10261...
Computers are not bicycles. The differences are important: the rate of change of the whole ecosystem is much higher, the ways economies of scale play out are different, load-balancing approaches have no analogue in the physical world, refactoring is zillions of times cheaper than it is with physical objects. Intuitions from physical engineering are often simply wrong when applied to computing.
Maybe you are not wrong for some smaller scale N, but after you need to make globally available datastores the speed of light is your enemy, not interconnect between two processors or some special NUMA configuration.
From that page it seems once enabled, there aren't any special requirements to get a GPU accelerating a query. As a result, I'd be surprised as a result if "GPU optimized" databases overtake regular-db-with-gpu-acceleration-addins.
Disclosure: I am the CEO of MapD.
You have to remember that many of our customers use us with our visual analytics frontend - they don't care about what happens behind the scenes, just that they can interactively explore billions of records with near-zero lag.
I'm a huge fan of PostgreSQL and certainly if you need an ACID transactional database with the latest SQL support and extensions you can't go wrong. But if you need the fastest analytics performance there might be more-suitable platforms. Different tools for different jobs.
Delete/Update in-flight.
This answered my question. So now for the follow-up. Are they limited to one core because their goal is to be ACID? As much as people like to blame performance on the tenets of ACID, it's still pretty important for a great deal of things.
I think bigger thread is cheap memory and raise of in-memory computing. Today you can have a workstation with half TB RAM for fairly reasonable price. Hadoop is already being crushed by Spark.
Also Spark is not a solution, it is an engine.
We run Spark on HDFS using all the paraphernalia of Hadoop to maintain some sort of sanity around it, how do you manage access? Encryption? Scheduling?
Oracles problem, customers who are keen on the cloud may well go to GCE, MS or AWS. Customers who are not so keen seek to save the money with open source and commodity implementations.
If you are sitting on a huge installed base the very last thing you want to see is disruptive technology shaking out your customer base. Oracle has seen three waves in five years - Hadoop, Cloud and now GPU/SCM. It's a tribute to the software and strategy of Oracle that they aren't bleeding rivers of red ink.
In the case of Hadoop, it might be possible to transparently translate the scripts to a GPU backend.
Here's a recent article by NVIDIA on using GPUs for graph computation which is somewhat related: https://devblogs.nvidia.com/parallelforall/gpus-graph-predic...