HNHacker News
TopNewBestAskShowJobs

nchammas

254 karma · joined May 6, 2014

submissionscomments
nchammas··on The Red Cross Blood Service: Australia's largest ever leak of personal data
> This is exactly why large stores of data should be considered a liability and NOT an asset.

Your comment reminded me of a post by Bruce Schneier with a similar theme, called "Data Is a Toxic Asset" [0].

[0] https://www.schneier.com/blog/archives/2016/03/data_is_a_tox...

nchammas··on Static types in Python
> we are moving towards strictly typed Python

I don't know what "strictly typed Python" is, but the type checking discussed in this post is 1) completely optional and 2) static. You don't have to do it if you don't want to, and even if you do do it, it doesn't affect in any way Python's runtime behavior or flexibility. You can have a codebase that works fine yet at the same time fails mypy's checks.

Also, no, we are not moving towards "strictly typed Python", even if that were a thing. PEP 484, which laid out the standard type hints, included this line in bold under "Non-goals" [0]:

> It should also be emphasized that Python will remain a dynamically typed language, and the authors have no desire to ever make type hints mandatory, even by convention.

[0] https://www.python.org/dev/peps/pep-0484/#non-goals

nchammas··on How Many Die from Medical Mistakes in U.S. Hospitals? (2013)
It was probably about the work of Peter Pronovost, or something closely related:

https://en.wikipedia.org/wiki/Peter_Pronovost

nchammas··on Implementation of various string similarity and distance algorithms in Java
Similar library for Python: https://github.com/jamesturk/jellyfish
nchammas··on Spark 2.0.0 Released
I think it would be dangerous to have the default be "coalesce to 1 partition before writing", but I agree this should be better documented since it takes many people by surprise.

As for the DenseVectors, that looks strange and is perhaps worthy of a report on the project tracker.

nchammas··on Spark 2.0.0 Released
> What happens is that each Spark partition outputs its own CSV, which is never what anyone would want!

1 file per partition is exactly what you want when the output is large, so that multiple executors can share the work of writing out the output.

> The expected (undocumented) solution is to compress into one partition, but then non-primitive data structures get output incorrectly.

By "compress" do you mean coalesce? Yes, coalescing the RDD/DataFrame into one partition is the commonly accepted solution for when you want to force 1 output file. The cost of doing so is that you lose parallelism, since only 1 task will be able to write out that file.

And what do you mean by "non-primitive data structure" in this case?

nchammas··on Generating Recommendations at Amazon Scale with Apache Spark and Amazon DSSTNE
Well, the absolute easiest way to run Spark is to do it locally (e.g. you can brew install it on a Mac and just go) or to pay for a proprietary service like Databricks, which makes setting up a cluster take a few clicks.

That said, I think `flintrock launch my-cluster` is almost as easy as doing `pip install ...`.

You do need an AWS account and you do need to set your preferences like region and key name in a config file, but I don't see how you can get out of doing even that without subscribing to some managed service like Databricks that abstracts everything away and replaces it with a nice Web UI.

nchammas··on Generating Recommendations at Amazon Scale with Apache Spark and Amazon DSSTNE
> as long as setting up a cluster is a bit easier!

If you're on AWS, it's already quite easy to set up a cluster today, no?

There's EMR of course, and there are tools like spark-ec2 [0] and Flintrock [1].

There are a few more tools listed on Spark Packages that target different cloud providers [2], too.

Disclaimer: I am the primary author of Flintrock and am a contributor to spark-ec2.

[0] https://github.com/amplab/spark-ec2

[1] https://github.com/nchammas/flintrock

[2] https://spark-packages.org/?q=tags%3Adeployment

nchammas··on New Cities
Relevant reading:

* https://en.wikipedia.org/wiki/Induced_demand

* https://en.wikipedia.org/wiki/Demand_destruction

* https://en.wikipedia.org/wiki/Disappearing_traffic

nchammas··on License update
I think this blog post [0] is more of a direct response to the discussion in that thread.

[0] https://whispersystems.org/blog/the-ecosystem-is-moving/

nchammas··on Microsoft announces major commitment to Apache Spark
You're right about Python and R having to pass data back and forth to the JVM for certain operations, but also keep in mind that native code still runs in the native interpreter. That means you have access to the full ecosystem of the native language.

For example, if I want to convert an RDD of JSON strings into Python dictionaries:

    import json
    rdd_dict = rdd.map(lambda x: json.loads(x))
Same goes for any external Python libraries I install on the cluster and want to use in my Spark job. You can even run your Python code on PyPy [4]!

For me, working in Python generally feels like a first class experience on Spark. There are areas -- like GraphX [0], certain niche features [1] -- where Scala is definitely easier to work with, but with time that is becoming less [2] and less [3] true thanks to the DataFrame API.

[0] https://spark.apache.org/graphx/

[1] http://stackoverflow.com/q/23995040/877069

[2] https://github.com/graphframes/graphframes

[3] http://stackoverflow.com/a/37150604/877069

[4] https://github.com/apache/spark/pull/2144

nchammas··on Microsoft announces major commitment to Apache Spark
First class language support in Apache Spark:

  * Scala
  * Python
  * Java
  * R
All these languages are equal, but Scala tends to be more equal than others in some areas of the API. I also believe R is mostly restricted to the DataFrame API.

Third-party language support:

  * Clojure [0]
To develop on Spark in a new, non-JVM language, you'd need a bridge to Java. That's how PySpark works [1], and I believe R follows a similar pattern.

[0] https://github.com/yieldbot/flambo

[1] https://cwiki.apache.org/confluence/display/SPARK/PySpark+In...

nchammas··on Asyncssh – Python Asyncio Client/Server Implementation of SSHv2 Protocol
As Jeff and I noted on Reddit [0], the author of AsyncSSH, Ron, is very responsive and his work is high quality.

The library is already pretty feature-complete, and just last week the latest release added ssh-agent client support. [1]

[0] https://www.reddit.com/r/Python/comments/41zycz/asyncssh_asy...

[1] https://github.com/ronf/asyncssh/blob/master/docs/changes.rs...

nchammas··on Asyncssh – Python Asyncio Client/Server Implementation of SSHv2 Protocol
`await` is Python 3.5+. AsyncSSH is compatible with Python 3.4+.
nchammas··on Can Good Doctors Be Bad for Your Health?
Good question. Perhaps that is one of the root issues here. As a layperson, I wouldn't know how to answer it, but I'm guessing the answer probably involves more rigorously detailing the various aspects of each case so that comparisons can be fairly made across surgeons treating similar cases--apples to apples and all that.
nchammas··on Can Good Doctors Be Bad for Your Health?
> For example, if surgeons knew their success rate were public, they would be incentivized to take easier cases. Who would take a difficult case if they knew it would constitute a bad mark on their record almost for sure?

Why would it be a "bad mark" if everyone understood the case was difficult? Given a difficult case, wouldn't a surgeon be graded badly only if they did poorer on average than other surgeons tackling similar cases?

As long as a case's "difficulty" is measured in a consistent way, I don't see how surgeons would be incentivized to avoid difficult cases.

nchammas··on What the Intercept’s New Audience Measurement System Means for Reader Privacy
"Together with Parse.ly, we’ve arrived at a system whereby readers of The Intercept will not directly ping Parse.ly. Instead, they will continue to send web requests to our own servers, which will, in turn, forward some of those requests on to Parse.ly, after stripping out readers’ internet protocol, or IP, addresses. Parse.ly will use these requests to track our readers via random unique identifiers that we generate. It will not be possible for Parse.ly to correlate readers’ visits to The Intercept with their visits to other Parse.ly-enabled sites."
nchammas··on YC Research
Is there a talk or discussion somewhere that touches on some of these insights into creating a great research organization?
nchammas··on Everyone you know will be able to rate you on the terrifying ‘Yelp for people’
For some reason this reminds me of the debacle [0][1] where some site was accepting "tips" from users who wanted to support any GitHub project of their choosing. They did it without explicit opt-in of the project owners (Armin Ronacher in this case).

Quote [2]:

> But nevertheless i think tip4commit should be opt-in instead,

> because i consider the current tactic to be a very intrusive form of growth-hacking.

[0] https://github.com/tip4commit/tip4commit/issues/127

[1] https://news.ycombinator.com/item?id=8542969

[2] https://github.com/tip4commit/tip4commit/issues/127#issuecom...

nchammas··on In Iraq, I raided insurgents. In Virginia, the police raided me
The author's recommendation that the police build up community relationships to increase trust and reduce unnecessary confrontation reminded me of a book, Fixing Broken Windows [1], that made a similar recommendation for what it called "community policing".

[1] http://www.amazon.com/Fixing-Broken-Windows-Restoring-Commun...

nchammas··on Go Lang: Comments Are Not Directives
The decision to put type annotations for variables in comments was made specifically for backwards compatibility reasons.

Remember that Python has supported function annotations since 3.0 [0]. Python 3.5 will only add the option to type check those same annotations.

Any type annotations that can be checked by Python 3.5 can also be parsed by earlier versions of Python, though without being checked.

The only wart in this master plan is specifically the one where you want to type annotate variables (as opposed to functions). Comments are the only way they could do this in a backwards compatible way.

The Python devs are well aware of the ugliness of this compromise, which is why they end the section you linked to with this:

> If type hinting proves useful in general, a syntax for typing variables may be provided in a future Python version.

I can't comment on Go, but I wouldn't call Python's choice here unreasonable.

[0] https://www.python.org/dev/peps/pep-3107/

nchammas··on Python 3 in Science: the great migration has begun
On a related note, Apache Spark recently landed Python 3 support in master, which will be released as part of Spark 1.4 in the next month or so. [0]

[0] https://issues.apache.org/jira/browse/SPARK-4897

nchammas··on Tmux 2.0 released
Hi Nicholas! That's good to hear, and definitely makes more sense given how popular and old tmux is as a project.

So I guess something just didn't translate correctly to Thomas's GitHub mirror.

nchammas··on Ask HN: Should HN support Markdown comments?
The feature I miss most is being able to link text.

Right now people just paste links as-is, or use this numbering syntax [0] which works but could definitely be replaced with something better.

After that, supporting lists and fenced code blocks with syntax highlighting would be handy, though those are probably more prone to abuse.

[0] Where the link goes here.

nchammas··on Tmux 2.0 released
Searching the tmux mailing list [0], it looks like there were some recent discussions about migrating to GitHub.

Basically, it looks like the developers don't see any benefit to migrating that outweighs the cost of a migration. [1]

There is a GitHub mirror [2] maintained by, I think, one of the developers, but they don't accept pull requests.

It's interesting to note that, for such a popular project like tmux, there are so few contributors. [3] Five, to be exact.

Maybe that's an artifact of how the mirroring to GitHub works (perhaps the commit authors aren't translated correctly?), but I suspect it's more an artifact of an outdated contribution process and tooling.

[0] http://sourceforge.net/p/tmux/mailman/search/?q=github

[1] http://sourceforge.net/p/tmux/mailman/message/33826395/

[2] https://github.com/ThomasAdam/tmux

[3] https://github.com/ThomasAdam/tmux/graphs/contributors

nchammas··on Apache Zeppelin – Interactive analytics on Spark
Just to be clear, though that video shows an example of what a notebook is in this context, that isn't a demo of Zeppelin. (It wasn't clear which you were referencing in your comment.)

That's a demo of Databricks Cloud [0], which is Databricks's product offering.

Among the many things it offers is an interactive notebook that looks like open source alternatives like Zeppelin.

[0] https://databricks.com/product/databricks-cloud

nchammas··on PEP 492 – Coroutines with async and await syntax
I'm not sure from what perspective you were looking to compare asyncio to green threads (performance, readability, etc.) but there is an interesting blog post [0] by the lead architect of Twisted about the fundamental difference in programming model between explicit yielding (e.g. asyncio) and implicit yielding (e.g. green threads, regular threads). It may give you some answers.

PEP 492, from what I understood, is more about streamlining the syntax, and the major differences between asyncio (pre this proposed syntax) and green threads in terms of programming model still apply.

[0] https://glyph.twistedmatrix.com/2014/02/unyielding.html

nchammas··on Ask HN: What are the most useful Chrome extensions?
Markdown Here [1]

It lets me write stuff on Gmail/Inbox or a number of a other sites in Markdown.

Then I hit Control + Option + M (on OS X) and bam, it coverts the Markdown to neat HTML.

It's really great for writing technical emails with code snippets and stuff. You even get syntax highlighting for those code snippets, too.

[1] http://markdown-here.com/

nchammas··on Announcing Docker Machine Beta
> It looks interesting but seems like a very Docker-centric Packer

I think you mean a Docker-centric Vagrant, no? Packer is for building images, Vagrant is for launching environments.

You can use Packer to build Docker images akin to `docker build`, except Packer will let you use a variety of provisioners to build your Docker image like Chef, Ansible, or just simple Bash scripts. It's `docker build` without the Dockerfiles. [1]

Docker Machine seems geared towards launching a resource that is ready to run Docker, whether it be a VM running locally or a remote instance in some cloud. It's a single-purpose Vagrant, to really stretch the analogy. :)

Once you have a resource that runs Docker, I think that's where Docker Compose [2] steps in. Docker Compose is more like the full Vagrant in that it spins up a full, functioning environment with running services and whatnot.

[1] https://www.packer.io/docs/builders/docker.html (See the last section, titled "Dockerfiles".)

[2] http://docs.docker.com/compose/

nchammas··on Introducing DataFrames in Spark for Large Scale Data Science
Btw, there is an new PR by Josh Rosen to fix SPARK-4454 here: https://github.com/apache/spark/pull/4660
← PreviousPage 2 of 3Next →