Your comment reminded me of a post by Bruce Schneier with a similar theme, called "Data Is a Toxic Asset" [0].
[0] https://www.schneier.com/blog/archives/2016/03/data_is_a_tox...
254 karma · joined May 6, 2014
Your comment reminded me of a post by Bruce Schneier with a similar theme, called "Data Is a Toxic Asset" [0].
[0] https://www.schneier.com/blog/archives/2016/03/data_is_a_tox...
I don't know what "strictly typed Python" is, but the type checking discussed in this post is 1) completely optional and 2) static. You don't have to do it if you don't want to, and even if you do do it, it doesn't affect in any way Python's runtime behavior or flexibility. You can have a codebase that works fine yet at the same time fails mypy's checks.
Also, no, we are not moving towards "strictly typed Python", even if that were a thing. PEP 484, which laid out the standard type hints, included this line in bold under "Non-goals" [0]:
> It should also be emphasized that Python will remain a dynamically typed language, and the authors have no desire to ever make type hints mandatory, even by convention.
As for the DenseVectors, that looks strange and is perhaps worthy of a report on the project tracker.
1 file per partition is exactly what you want when the output is large, so that multiple executors can share the work of writing out the output.
> The expected (undocumented) solution is to compress into one partition, but then non-primitive data structures get output incorrectly.
By "compress" do you mean coalesce? Yes, coalescing the RDD/DataFrame into one partition is the commonly accepted solution for when you want to force 1 output file. The cost of doing so is that you lose parallelism, since only 1 task will be able to write out that file.
And what do you mean by "non-primitive data structure" in this case?
That said, I think `flintrock launch my-cluster` is almost as easy as doing `pip install ...`.
You do need an AWS account and you do need to set your preferences like region and key name in a config file, but I don't see how you can get out of doing even that without subscribing to some managed service like Databricks that abstracts everything away and replaces it with a nice Web UI.
If you're on AWS, it's already quite easy to set up a cluster today, no?
There's EMR of course, and there are tools like spark-ec2 [0] and Flintrock [1].
There are a few more tools listed on Spark Packages that target different cloud providers [2], too.
Disclaimer: I am the primary author of Flintrock and am a contributor to spark-ec2.
[0] https://github.com/amplab/spark-ec2
[0] https://whispersystems.org/blog/the-ecosystem-is-moving/
For example, if I want to convert an RDD of JSON strings into Python dictionaries:
import json
rdd_dict = rdd.map(lambda x: json.loads(x))
Same goes for any external Python libraries I install on the cluster and want to use in my Spark job. You can even run your Python code on PyPy [4]!For me, working in Python generally feels like a first class experience on Spark. There are areas -- like GraphX [0], certain niche features [1] -- where Scala is definitely easier to work with, but with time that is becoming less [2] and less [3] true thanks to the DataFrame API.
[0] https://spark.apache.org/graphx/
[1] http://stackoverflow.com/q/23995040/877069
[2] https://github.com/graphframes/graphframes
* Scala
* Python
* Java
* R
All these languages are equal, but Scala tends to be more equal than others in some areas of the API. I also believe R is mostly restricted to the DataFrame API.Third-party language support:
* Clojure [0]
To develop on Spark in a new, non-JVM language, you'd need a bridge to Java. That's how PySpark works [1], and I believe R follows a similar pattern.[0] https://github.com/yieldbot/flambo
[1] https://cwiki.apache.org/confluence/display/SPARK/PySpark+In...
The library is already pretty feature-complete, and just last week the latest release added ssh-agent client support. [1]
[0] https://www.reddit.com/r/Python/comments/41zycz/asyncssh_asy...
[1] https://github.com/ronf/asyncssh/blob/master/docs/changes.rs...
Why would it be a "bad mark" if everyone understood the case was difficult? Given a difficult case, wouldn't a surgeon be graded badly only if they did poorer on average than other surgeons tackling similar cases?
As long as a case's "difficulty" is measured in a consistent way, I don't see how surgeons would be incentivized to avoid difficult cases.
Quote [2]:
> But nevertheless i think tip4commit should be opt-in instead,
> because i consider the current tactic to be a very intrusive form of growth-hacking.
[0] https://github.com/tip4commit/tip4commit/issues/127
[1] https://news.ycombinator.com/item?id=8542969
[2] https://github.com/tip4commit/tip4commit/issues/127#issuecom...
[1] http://www.amazon.com/Fixing-Broken-Windows-Restoring-Commun...
Remember that Python has supported function annotations since 3.0 [0]. Python 3.5 will only add the option to type check those same annotations.
Any type annotations that can be checked by Python 3.5 can also be parsed by earlier versions of Python, though without being checked.
The only wart in this master plan is specifically the one where you want to type annotate variables (as opposed to functions). Comments are the only way they could do this in a backwards compatible way.
The Python devs are well aware of the ugliness of this compromise, which is why they end the section you linked to with this:
> If type hinting proves useful in general, a syntax for typing variables may be provided in a future Python version.
I can't comment on Go, but I wouldn't call Python's choice here unreasonable.
So I guess something just didn't translate correctly to Thomas's GitHub mirror.
Right now people just paste links as-is, or use this numbering syntax [0] which works but could definitely be replaced with something better.
After that, supporting lists and fenced code blocks with syntax highlighting would be handy, though those are probably more prone to abuse.
[0] Where the link goes here.
Basically, it looks like the developers don't see any benefit to migrating that outweighs the cost of a migration. [1]
There is a GitHub mirror [2] maintained by, I think, one of the developers, but they don't accept pull requests.
It's interesting to note that, for such a popular project like tmux, there are so few contributors. [3] Five, to be exact.
Maybe that's an artifact of how the mirroring to GitHub works (perhaps the commit authors aren't translated correctly?), but I suspect it's more an artifact of an outdated contribution process and tooling.
[0] http://sourceforge.net/p/tmux/mailman/search/?q=github
[1] http://sourceforge.net/p/tmux/mailman/message/33826395/
That's a demo of Databricks Cloud [0], which is Databricks's product offering.
Among the many things it offers is an interactive notebook that looks like open source alternatives like Zeppelin.
PEP 492, from what I understood, is more about streamlining the syntax, and the major differences between asyncio (pre this proposed syntax) and green threads in terms of programming model still apply.
It lets me write stuff on Gmail/Inbox or a number of a other sites in Markdown.
Then I hit Control + Option + M (on OS X) and bam, it coverts the Markdown to neat HTML.
It's really great for writing technical emails with code snippets and stuff. You even get syntax highlighting for those code snippets, too.
I think you mean a Docker-centric Vagrant, no? Packer is for building images, Vagrant is for launching environments.
You can use Packer to build Docker images akin to `docker build`, except Packer will let you use a variety of provisioners to build your Docker image like Chef, Ansible, or just simple Bash scripts. It's `docker build` without the Dockerfiles. [1]
Docker Machine seems geared towards launching a resource that is ready to run Docker, whether it be a VM running locally or a remote instance in some cloud. It's a single-purpose Vagrant, to really stretch the analogy. :)
Once you have a resource that runs Docker, I think that's where Docker Compose [2] steps in. Docker Compose is more like the full Vagrant in that it spins up a full, functioning environment with running services and whatnot.
[1] https://www.packer.io/docs/builders/docker.html (See the last section, titled "Dockerfiles".)