Microsoft announces major commitment to Apache Spark
blogs.technet.microsoft.com
blogs.technet.microsoft.com
* Scala
* Python
* Java
* R
All these languages are equal, but Scala tends to be more equal than others in some areas of the API. I also believe R is mostly restricted to the DataFrame API.Third-party language support:
* Clojure [0]
To develop on Spark in a new, non-JVM language, you'd need a bridge to Java. That's how PySpark works [1], and I believe R follows a similar pattern.[0] https://github.com/yieldbot/flambo
[1] https://cwiki.apache.org/confluence/display/SPARK/PySpark+In...
https://github.com/henridf/apache-spark-node https://github.com/EclairJS/eclairjs-node
Scala -- native API
Java -- java wrappers for Scala API, slightly clunky but same implementation
Python+R -- janky process involving forking a process and feeding strings back and forth between JVM and python/R via pipes
For example, if I want to convert an RDD of JSON strings into Python dictionaries:
import json
rdd_dict = rdd.map(lambda x: json.loads(x))
Same goes for any external Python libraries I install on the cluster and want to use in my Spark job. You can even run your Python code on PyPy [4]!For me, working in Python generally feels like a first class experience on Spark. There are areas -- like GraphX [0], certain niche features [1] -- where Scala is definitely easier to work with, but with time that is becoming less [2] and less [3] true thanks to the DataFrame API.
[0] https://spark.apache.org/graphx/
[1] http://stackoverflow.com/q/23995040/877069
[2] https://github.com/graphframes/graphframes
That said, from my experience the Python integration is still first-class. I think it might lag the Scala APIs a bit, but that isn't a huge deal unless you really need to be on the bleeding edge at all times. The only really big downside in my book is that there's a small performance hit you take by choosing to use the Python API instead of the Scala one, and that might compound into something of practical significance if you're using it for really big data tasks.
I assume you'd have similar problems using it with R or a .NET language.
[1] http://www.businessinsider.com/microsoft-azure-vs-aws-revenu...
Servers and tools did $4.5B of revenue in FY12, Q3 (https://www.microsoft.com/Investor/EarningsAndFinancials/Ear...). This was before Azure really had any footprint.
You do the math.
https://en.wikipedia.org/wiki/Embrace,_extend_and_extinguish
However, even if they don't, I haven't seen any indication that they've repented of their past scummy tendencies. The Windows 10 rollout is perhaps the most obvious case in point.
I'm curious if you've seen something that makes you more optimistic than I am.
It's mind boggling to consider how many years their shenanigans have probably set back computing as a field.
As for trusting them now, I don't believe an organization like that can truly change in such a short time.
Hence, the total apparent lack of desire to do anything new and innovative ("Let's copy AWS! Genius!").
We were in a meeting talking about infrastructure and testing. I said that we used EC2 for some of our infrastructure, and we were thinking of moving some of our local dev infra to AWS. Someone in the room asked me what that was. I responded with "Amazon web services" and they countered "you mean the company that sells books and stuff?" -- they had never heard of AWS, nor EC2
This comes from someone who had been with the organization for some time. He was very well respected by a lot of folks. He had more awards in his office than I could count. I still think he's a great guy.
The Microsoft monoculture hurt quite a bit.
That said, I'm not entirely certain why they decided to implement TFVC (TFS's default - and until recently, only - version control system) instead of just basing it on SVN. Aside from that this was at the height of the period when Microsoft had a compulsive need to sink resources into building their own proprietary clone of every damn thing.
Although TFS may seem terrible, it works for super large organizations, and big code bases in a way that SVN and Git just didn't.
Teams that had issues grasping the fundamentals of VCS in general still had issues. And in return you got the lesser compatibility and greater number of code warts that came with proprietary enterprise software over open source.
But I'm curious, what worked specifically for you in TFS that wasn't in the SVN ecosystem? It's probably the client was unaware of a lot of features.
If I want to do a deploy of my lambda software, I edit some Python in IntelliJ, I then run a build and upload it to lambda using some bash scripts, and test it using the GUI in Chrome.
If I'm editing C code, I edit some code in Netbeans, and edit it over Netbeans SSH integration in VB. I then go compile it using make on the remote machine, and load it. I have a bunch of bash scripts to do testing on that remote machine that are executed via SSH.
----- I have several more duct taped integrations. With VS, this is all in one system.
I've seen HN light up with joy at news of all the open sourcing and enhancements MS are undertaking in recent years. I don't see the disorganised mess but perhaps I'm missing something?