IBM Invests to Help Apache Spark
bits.blogs.nytimes.com
bits.blogs.nytimes.com
You dispatch n jobs, where n is quite large, and you want to know; have all n jobs have completed, or has less than n jobs have completed. How to do so with a small fixed number of bytes with very high probability?
Give each job a random 128-bit ID number. XOR each ID number together as you start each job, and XOR into the same value as each job completes. If all the jobs have completed, the result is 0. The chance of zero turning up randomly if all jobs are not complete is negligible.
The technique is mentioned here under 'Lineage Tracking': https://highlyscalable.wordpress.com/2013/08/20/in-stream-bi... but there's a better blog post I remembered reading but can't find at the moment...
As the article mentioned, IBM certainly did validate the Linux "market." When people would ask me what was great about Linux I used to just say that IBM was investing billions in Linux, and that was an acceptable answer for people.
As far as I know, virtually any MapReduce job can be rather trivially translated to Spark's .map() and .reduce() operations. And the downsides are: it's model hasn't yet been proven at the largest scale's MapReduce has been used, and possibly the use of Scala (although Java / Python bindings are obviously available). Were there any other major factors in your hesitance?
Another issue is that I am sort of retired now. I still accept small consulting jobs and do a lot of writing but my technology choices have shifted to fun things like Pharo Smalltalk, Haskell, etc.
From what I understand of it - it's an implementation of apache storm in Haskell.
That is impressive. I wonder how that will be split among core contributors, consultants, etc.
"At the core of this commitment, IBM plans to embed Spark into its industry-leading Analytics and Commerce platforms, and to offer Spark as a service on IBM Cloud. IBM will also put more than 3,500 IBM researchers and developers to work on Spark-related projects at more than a dozen labs worldwide; donate its breakthrough IBM SystemML machine learning technology to the Spark open source ecosystem; and educate more than one million data scientists and data engineers on Spark."
By deciding to sponsor Spark, I think IBM is becoming practically it's owner, without having to do anything prior to this move. Does it mean it is possible today to "acquire" technology a project by naming your own price?
Hmm, why do you think this makes IBM the technology's "owner"? What do you mean by owner exactly?
This is a really interesting phenomenon - existing enterprises taking lead roles in open source projects in such a way that they almost look like they "own it". Not the first time, I mean, RedHat did exactly this with Linux in some ways, for example.
Teradata recently made a similar move with PrestoDB - which had Facebook as the owner but no hardcore platform technology development supporter. And MapR did the same with Apache Drill.
I have been telling everyone I know my opinion for a long time - open source is not just a movement - it is a strategy. It is amazing the potential for change and/or disruption that open source causes.
And, that potential may not always be good - just look at Hadoop fragmentation as an example (although there is a lot of good in that fragmentation as well).
Declarative large-scale machine learning (ML) in SystemML aims at flexible specification of ML algorithms and automatic generation of hybrid runtime plans ranging from single node, in-memory computations to distributed computations on MapReduce or Spark. ML algorithms are expressed in an R-like syntax, that includes linear algebra primitives, statistical functions, and ML-specific constructs.
http://researcher.watson.ibm.com/researcher/view_group.php?i...
[1] http://www.theregister.co.uk/2015/06/15/ibm_backs_apache_spa...
http://venturebeat.com/2015/06/14/ibm-spark/
The spark ecosystem itself has a lot of players now.
Horton/Cloudera/MapR in their hadoop distros Typesafe: https://www.typesafe.com/community/other-projects/apache-spa...
There's a second, more advanced course too: http://bigdatauniversity.com/bdu-wp/bdu-course/spark-fundame...