* Scala
* Python
* Java
* R
All these languages are equal, but Scala tends to be more equal than others in some areas of the API. I also believe R is mostly restricted to the DataFrame API.Third-party language support:
* Clojure [0]
To develop on Spark in a new, non-JVM language, you'd need a bridge to Java. That's how PySpark works [1], and I believe R follows a similar pattern.[0] https://github.com/yieldbot/flambo
[1] https://cwiki.apache.org/confluence/display/SPARK/PySpark+In...
https://github.com/henridf/apache-spark-node https://github.com/EclairJS/eclairjs-node
Scala -- native API
Java -- java wrappers for Scala API, slightly clunky but same implementation
Python+R -- janky process involving forking a process and feeding strings back and forth between JVM and python/R via pipes
For example, if I want to convert an RDD of JSON strings into Python dictionaries:
import json
rdd_dict = rdd.map(lambda x: json.loads(x))
Same goes for any external Python libraries I install on the cluster and want to use in my Spark job. You can even run your Python code on PyPy [4]!For me, working in Python generally feels like a first class experience on Spark. There are areas -- like GraphX [0], certain niche features [1] -- where Scala is definitely easier to work with, but with time that is becoming less [2] and less [3] true thanks to the DataFrame API.
[0] https://spark.apache.org/graphx/
[1] http://stackoverflow.com/q/23995040/877069
[2] https://github.com/graphframes/graphframes
That said, from my experience the Python integration is still first-class. I think it might lag the Scala APIs a bit, but that isn't a huge deal unless you really need to be on the bleeding edge at all times. The only really big downside in my book is that there's a small performance hit you take by choosing to use the Python API instead of the Scala one, and that might compound into something of practical significance if you're using it for really big data tasks.
I assume you'd have similar problems using it with R or a .NET language.