2016 Spark Summit East Keynote
slideshare.net
slideshare.net
https://spark-summit.org/east-2016/events/graphframes-graph-...
[1] https://issues.apache.org/jira/browse/SPARK-8360 [2] https://spark-summit.org/east-2016/events/keynote-day-3/
> CPU speeds have not kept up with I/O in the past 5 years.
I presume he means the other way around?
Also, what does he mean by native memory management? Does he mean off-heap allocation?
And what's he referring to regarding code generation?
Native refers to (I think) the following issues: https://issues.apache.org/jira/browse/SPARK-12785 https://issues.apache.org/jira/browse/SPARK-8641
Code generation is enabled by SPARK-8641, but not sure exactly what it entails. I think it is related to some of the RDD transformation/action merging they do to optimize runtime operations in 2.0.
Your thoughts?
Basically you take the general, composable API functions and generate specific equivalent code that avoids the overhead of dealing with abstract interfaces like Iterators.
That's pretty cool. Similar to what Rust already does when the compiler elides allocations for chained iterators. I wonder if value type support in the JVM would avoid the need to do this type of code gen.
How does Spark guarantee data memory contiguity(or do they at all)? Do they use misc.sun.unsafe or some form of memory management?
This is the thesis of those who have been watching the explosion in solid state disks. The claim is that bulk I/O is becoming so damn fast that the pendulum is swinging toward computing power being the new limiting factor.
Either way CPU processing power has never really been a constraining factor(since most people don't want to write data aware code), it's always been bus-bound be it DRAM, IO or other.
Datastax evangelized people to use Spark to run queries over Cassandra but it looks so awkward and time consuming to copy jars around to the master, basically you need a dev ops team to this and even more scriptology for production.
val myDF = sqlContext.read.format("com.databricks.spark.csv") //allows you to read a csv file (for simplicity)
.option("header", "true")
.option("delimiter", "\\t")
.option("mode", "PERMISSIVE")
.option("inferSchema", "true")
.load(filename)
.where(filter_query)
.cache()
myDF.registerTempTable("myDF")
sqlContext.sql("SELECT COUNT(*) FROM myDF").show()
sqlContext.sql("SELECT COUNT(*) FROM myDF where filter_query").show()Why not just use Presto? It gives you basically full SQL capability for Cassandra with minimal effort.
We tested it thoroughly and came to the conclusion that is wasn't a mature enough solution.
Matei was actually talking about the existing Spark Kafka direct stream implementation, which has been available since Spark 1.3
The video of the talk is available here: http://livestream.com/fourstream/sparksummiteast2016-tracka/...
But you can always download and build the source!