Announcing Apache Spark 1.1
databricks.com
databricks.com
https://issues.apache.org/jira/browse/SPARK-2360 https://github.com/apache/spark/pull/1351
Re: databricks cloud - shoot me an e-mail and I'll see if I can help. Right now demand exceeds supply for us on accounts, but I can try!
Online graph algorithms aren't there yet (probably what you mean). We just started adding online MLlib algorithms, so this is the main focus for now.
> This release adds significant internal changes to Spark focused on improving performance for large scale workloads.
We looked at Spark Streaming briefly when choosing which CEP engine to use. We ended up not using it as its performance wasn't on par with other offerings.
I hope the performance improvements they've done carry over to the spark streaming product. http://spark.apache.org/streaming/
We ended up using esper and a proprietary engine. The biggest problem with some other streaming utilities is
1) performance
2) you can't step time when doing back testing. ie when replaying a day of trading you'll often have signals that say something like be passive for the next 10 seconds and then go to the bid for 10 seconds and then mid point and if you still aren't done then cross the spread after an additional 10 seconds.
When backtesting you obviously don't play back in real time or it would take 6.5 hours to re simulate the day, you play back as fast as you can so you need your CEP system to step time as it goes so it properly fires your time based triggers.
I'm checking out memSql right now to see how it lets you step time whit queries.
Startup idea.... It would be nice to have one unified way of doing real time and batch processing of data. That way your real time trading engine can be back tested in the same way it runs in production. I think/hope Spark is on the way to solving this.
If anyone has any input on the best way to unify streaming vs batch event processing please let me know!!!
When everything runs at wire speed and all data is always fully online, there is no meaningful distinction between "batch" and "streaming".
That said, the reason you generally do not see anything in open source that does it is that it requires a more sophisticated internal design and architecture than you typically see in any of the open source databases or big data platforms. Even a system designed to only run on a single machine (as opposed to a parallel cluster) is probably 50kLoC of dense, complex, low-level C++ just to get a basic kernel off the ground.
The main reasons you do not see a lot of startups doing this: The design of these types of storage/execution engines is rare knowledge. You have to design and implement most of your own algorithms and data structures -- few things can be farmed out to template libraries or system call uses. Relatively few developers are sufficiently skilled in C++ to successfully implement these kinds of systems with these kinds of performance envelopes. The code base for an MVP, excluding the expansive test infrastructure, is pretty huge, so you need a lot of man-hours with the above skills and expertise.
The high initial cost of building these types of systems and the difficulty of finding the necessary talent make them unattractive to investors as startups. You often have to spend $10-20M before you are even going to know if there is traction. That is a pricy bet.