HNHacker News
TopNewBestAskShowJobs

ims

798 karma · joined September 12, 2011

submissionscomments
ims··on Killer Rabbits in Medieval Manuscripts
My great aunt wrote some of the original scholarship on this type of marginalia: https://www.goodreads.com/book/show/5064937-images-in-the-ma...

It's a shame that she is not mentioned. In answer to one of the other posters wondering about Monty Python, the answer is yes. She told me that Eric Idle or possibly Terry Jones (?) at one point had a fascination with these illustrations and I believe they corresponded. In any case, it wasn't just the rabbit sketch - some of the interstitial animations between skits are taken directly from tropes in these drawings.

ims··on Massachusetts Court Blocks Warrantless Access to Real-Time Location Data
Check the post history for this person. There's a weirdly specific and persistent anti-Massachusetts vendetta.
ims··on Why was it so hard to take a picture of a black hole?
Do you think the fact that the CT scanner at your local hospital needs to be calibrated and computationally reconstructed from X-ray intensities mean it does not result in an "image"?

When we use side-scan sonar to create representations of the ocean floor (e.g. https://commons.wikimedia.org/wiki/File:Laevavrakk_"Aid".png), they are computationally reconstructed from the raw data which are not intrinsically recognized as pixels without reconstruction. Are these not "images"?

What is your actual contention here? Is it that any representation which is not the result of a traditional visible-light camera doesn't count as an "image"?

If so it's an irrelevant distinction to make. If not, you need to articulate in a specific and informed way why the way they reconstructed the image was wrong or could be improved.

It seems from your blog that you don't really understand what a "prior" is and why it might be useful for this kind of signal processing.

ims··on Ask HN: Who is hiring? (April 2019)
DrivenData Labs | Data Scientist and Lead Data Scientist | Berkeley, CA or Boston, MA | Full-time | ONSITE

We run online machine learning challenges with social impact, and we work directly with mission-driven organizations to drive change through data science and engineering. Since 2014 we’ve worked with more than 35 organizations across 50+ projects in areas like international development, health, education, research and conservation, and public services.

We pride ourselves on being a great place to work and to learn. We take the development of our team members very seriously and we value the priorities that we each have in our lives at work and outside of work. We like to tackle problems that matter as a team. We help each other develop clean, well-organized, well-documented code in service of correct and reproducible data science. We believe the work we do should positively impact people’s lives.

Ultimately, we're a team of smart, passionate data scientists and engineers interested in doing good work for good reason. We're looking for more great people in Boston or the Bay area. We're excited to hear from you!

Positions: https://drivendata.workable.com/

ims··on Office Depot computer scans gave fake results
The actual complaint: https://www.ftc.gov/system/files/documents/cases/office_depo...
ims··on What you need may be “pipeline +Unix commands” only
Sometimes there's a middle ground: make your "map" and "reduce" steps separate scripts.

If you want to do the parsing in Python instead of awk, just make a tiny script that reads from stdin and writes to stdout - that way you can put it between xargs or parallel and whatever else is in the pipeline.

The parallelization is a separate concern, so it doesn't need to be mixed in with the parsing (or whatever) concern. The downloading is a separate concern; use wget or requests in a Python script or whatever, it doesn't need to be mingled with the parsing.

ims··on Intellectual Denial of Service Attacks
This is an interesting plot point in Vernor Vinge's Zones of Thought science fiction series. The "net of a million lies" has been astroturfed and boobytrapped by living beings and AIs for generations. Programmer archaeologists sometimes try to extract useful information at great risk to their entire societies.
ims··on Kyoto Tycoon in 2019
I'd love to hear more about this - email is in my profile if you have a minute.
ims··on Kyoto Tycoon in 2019
> When I find myself in the latter kind of work environment and I need to quickly get sortable and/or indexable datastructures of any kind, then a key-value store is the way to go

Interesting, can you expand on this?

ims··on Ask HN: Who is hiring? (December 2018)
DrivenData Labs | Software Engineer (Python) With Focus On Data Applications | Berkeley, CA | FULL-TIME, ONSITE | http://drivendata.co

-- OVERVIEW --

DrivenData brings the transformative power of data science to organizations tackling the world's biggest challenges. We run online machine learning challenges with social impact, and we work directly with mission-driven organizations to drive change through machine intelligence and analytics.

We are looking for a talented software engineer who is interested in using their job to take on tough social challenges, while growing their data acumen and building real-world applications. As a core member of a small team your role will include managing code development, brainstorming approaches to engineering problems, working closely with data science and machine learning developers, and taking an open and constructive mindset to getting things done across multiple projects. You'll work directly with data scientists that started their careers as software engineers, bringing an experienced understanding of software processes alongside opportunities to learn new quant skills, tools, and ways of approaching data applications.

-- ROLE --

This is a software developer role ideal for a Python engineer with 3-5 years of professional experience who is interested in data (possibly looking to transition into data engineering or data science). Advanced proficiency in Python and comfort with Linux a necessity. Good opportunity to learn the quant skills necessary to work in the data space. No need to have a background in math or a CS degree, but the job will involve a lot of quantitative thinking so the applicant should not be afraid of math.

Working on a small team means doing a little bit of a lot of things. We're looking for somebody who can ask the right questions to figure out what is important, iterate between brainstorming together and working independently, and exercise sound engineering judgment to make reasonable decisions under conditions of ambiguity.

Doing client-facing work involves turning uncertainty into a reasonable path forward. As a team, we value arguments for how to proceed based on evidence, and we want somebody who will present their opinion and engage in a discussion around the best way forward.

Here are some of the things you'd be doing on an ordinary day:

- Internal software development: Maintain our Django codebase for drivendata.org, fix bugs, add features, safely refactor and maintain test coverage. Develop new internal tooling and improve on existing apps.

- Client-facing software development. Build a variety of applications, generally small green-field proofs of concept. Quickly learn and adopt new technologies on demand based on client needs; a typical engagement may include at least one data technology we haven't all worked with before (e.g. Elasticsearch, Apache Storm, Cassandra).

- Light DevOps Tasks

-- REQUIREMENTS --

Critical skills: Python (advanced), Linux (advanced), SQL (intermediate to advanced). Must be able to learn quickly by reading appropriate documentation in order to write clean, idiomatic code.

Nice to have: Experience using IaaS like Amazon AWS or PaaS like Heroku. Experience using Docker. Exposure to big data tools like Spark or Hadoop, or familiarity with the underlying ideas like MapReduce.

Apply directly here: https://drivendata-labs.workable.com/jobs/839418

ims··on Experiment that ended in 1767 still linked to higher incomes, education levels
Andrew Gelman calls the larger problem the "Garden of Forking Paths" - http://www.stat.columbia.edu/~gelman/research/unpublished/p_...
ims··on Ask HN: Who is hiring? (November 2018)
DrivenData Labs | Software Engineer (Python) With Focus On Data Applications | Berkeley, CA | FULL-TIME, ONSITE | http://drivendata.co

-- OVERVIEW --

DrivenData brings the transformative power of data science to organizations tackling the world's biggest challenges. We run online machine learning challenges with social impact, and we work directly with mission-driven organizations to drive change through machine intelligence and analytics.

We are looking for a talented software engineer who is interested in using their job to take on tough social challenges, while growing their data acumen and building real-world applications. As a core member of a small team your role will include managing code development, brainstorming approaches to engineering problems, working closely with data science and machine learning developers, and taking an open and constructive mindset to getting things done across multiple projects. You'll work directly with data scientists that started their careers as software engineers, bringing an experienced understanding of software processes alongside opportunities to learn new quant skills, tools, and ways of approaching data applications.

-- ROLE --

This is a software developer role ideal for a Python engineer with 3-5 years of professional experience who is interested in data (possibly looking to transition into data engineering or data science). Advanced proficiency in Python and comfort with Linux a necessity. Good opportunity to learn the quant skills necessary to work in the data space. No need to have a background in math or a CS degree, but the job will involve a lot of quantitative thinking so the applicant should not be afraid of math.

Working on a small team means doing a little bit of a lot of things. We're looking for somebody who can ask the right questions to figure out what is important, iterate between brainstorming together and working independently, and exercise sound engineering judgment to make reasonable decisions under conditions of ambiguity.

Doing client-facing work involves turning uncertainty into a reasonable path forward. As a team, we value arguments for how to proceed based on evidence, and we want somebody who will present their opinion and engage in a discussion around the best way forward.

Here are some of the things you'd be doing on an ordinary day:

- Internal software development: Maintain our Django codebase for drivendata.org, fix bugs, add features, safely refactor and maintain test coverage. Develop new internal tooling and improve on existing apps.

- Client-facing software development. Build a variety of applications, generally small green-field proofs of concept. Quickly learn and adopt new technologies on demand based on client needs; a typical engagement may include at least one data technology we haven't all worked with before (e.g. Elasticsearch, Apache Storm, Cassandra).

- Light DevOps Tasks

-- REQUIREMENTS --

Critical skills: Python (advanced), Linux (advanced), SQL (intermediate to advanced). Must be able to learn quickly by reading appropriate documentation in order to write clean, idiomatic code.

Nice to have: Experience using IaaS like Amazon AWS or PaaS like Heroku. Experience using Docker. Exposure to big data tools like Spark or Hadoop, or familiarity with the underlying ideas like MapReduce.

Apply directly here: https://drivendata-labs.workable.com/jobs/839418

ims··on Boston-area startups are on pace to overtake NYC venture totals
> Lumping part of New Hampshire in with Boston is like lumping Northern Ireland in with London. One side considers it a horrible insult and the other doesn't see why you wouldn't want to be associated with them.

For those reading who are not from the area, this is fun to imagine but not actually true.

The tension between Northern and Southern New England is less serious even than the Boston/New York "rivalry" which is notional and mostly confined to sports.

ims··on Data Science Is America’s Hottest Job
I'd guess that historically the Mathematical Optimization department has comprised "legacy" operations research people (spreadsheet modeling and decision analysis, linear and nonlinear modeling with AMPL-type tools, SAS statistics, etc) and the Data Science group is newer and more software-ish.

Am I right?

ims··on #deletefacebook
Just a quick correction, he was president of the Harvard Law Review - not the school.
ims··on For AI to thrive, it must explain itself
Just as with self-driving cars, you have to look at the counterfactual. We don't have an interpretable audit trail for the decisions made by human employees either.

So under the status quo, what happens in the case of an aberrant or problematic decision and how does society cope with this lack of interpretability?

We try to piece together a plausible, retrospective narrative by looking at fact patterns and taking into account the education, experience, actions, explanations, and rationalizations of other humans. (Trusting these accounts is further complicated by human traits such as self-interest, emotion, and cognitive biases.)

There exist entire professional specialties devoted to this problem (litigation, internal investigations, police detectives, accident boards) who spend a lot of time on this. Even so, much of the time we still don't really know why exactly people made the decisions they did.

This inconvenient fact does not stop us from employing humans, and society has not collapsed under the weight of the liability issues.

ims··on Ask HN: Who is hiring? (January 2018)
DrivenData Labs | Software engineer (Python) w/ focus on data applications | Berkeley, CA | ONSITE | FULLTIME

DrivenData brings the transformative power of data science to organizations tackling the world’s biggest challenges. We run online machine learning challenges with social impact (drivendata.org), and we work directly with mission-driven organizations to drive change with statistical modeling, data engineering, and tool building (drivendata.co).

We are looking for a talented software engineer who is interested in data — possibly looking to transition into data engineering or data science — and in using their job to take on tough social challenges. As a core member of a small team your role will include managing code development, brainstorming approaches to engineering problems, working closely with data science and machine learning developers, and taking an open and constructive mindset to getting things done across multiple projects. You’ll work directly with data scientists that started their careers as software engineers, bringing an experienced understanding of software processes alongside opportunities to learn new quant skills, tools, and ways of approaching data applications. This is a full time position in Berkeley, CA (SF/Bay Area).

Doing client-facing work involves turning uncertainty into a reasonable path forward. As a team, we value unemotional arguments for how to proceed based on evidence, and we want somebody who can be assertive enough to get the point across but dispassionate enough to plow through even if their favored course of action doesn't happen this time. We're looking for somebody who can ask the right questions to figure out what is important, iterate between brainstorming together and working independently, and exercise sound engineering judgment to make reasonable decisions under conditions of ambiguity. Duties and responsibilities: internal software development, maintain our Python codebase for drivendata.org, fix bugs, add features, safely refactor and maintain test coverage. Develop new internal tooling and improve on existing apps. Client-facing software development; build a variety of applications, generally small green-field apps. Light DevOps Tasks (spinning up EC2 instances, logging into a servers for diagnosing issues, setting up databases both locally and in the cloud). Requirements: Advanced proficiency in Python, practical experience with writing solid and well-tested code, working knowledge of SQL, and comfort with Linux a necessity. No need to have a background in math or a CS degree, but the job will involve a lot of quantitative thinking so the applicant should not be afraid of math Working on a small team means doing a little bit of a lot of things. Able to quickly learn and adopt new technologies based on client needs; a typical engagement may include at least one data technology we haven't all worked with before. Must be able to read appropriate documentation in order to write clean, idiomatic code.

Nice-to-have experience: IaaS like Amazon AWS or PaaS like Heroku, Docker, big data tools like Spark and Hadoop, tools design for data-intensive applications e.g. Cassandra, Storm, Elasticsearch, etc.

If interested, send a resume and links to things you'd like us to see (e.g. Github, personal site, blog or projects) to isaac [at] drivendata.org with "HN" in the subject line.

ims··on Ask HN: Who is hiring? (November 2017)
DrivenData Labs | Software engineer (Python) w/ focus on data applications | Berkeley, CA | ONSITE | FULLTIME

DrivenData brings the transformative power of data science to organizations tackling the world’s biggest challenges. We run online machine learning challenges with social impact (drivendata.org), and we work directly with mission-driven organizations to drive change with statistical modeling, data engineering, and tool building (drivendata.co).

We are looking for a talented software engineer who is interested in data — possibly looking to transition into data engineering or data science — and in using their job to take on tough social challenges. As a core member of a small team your role will include managing code development, brainstorming approaches to engineering problems, working closely with data science and machine learning developers, and taking an open and constructive mindset to getting things done across multiple projects. You’ll work directly with data scientists that started their careers as software engineers, bringing an experienced understanding of software processes alongside opportunities to learn new quant skills, tools, and ways of approaching data applications. This is a full time position in Berkeley, CA (SF/Bay Area).

Doing client-facing work involves turning uncertainty into a reasonable path forward. As a team, we value unemotional arguments for how to proceed based on evidence, and we want somebody who can be assertive enough to get the point across but dispassionate enough to plow through even if their favored course of action doesn't happen this time. We're looking for somebody who can ask the right questions to figure out what is important, iterate between brainstorming together and working independently, and exercise sound engineering judgment to make reasonable decisions under conditions of ambiguity.

Duties and responsibilities: internal software development, maintain our Python codebase for drivendata.org, fix bugs, add features, safely refactor and maintain test coverage. Develop new internal tooling and improve on existing apps. Client-facing software development; build a variety of applications, generally small green-field apps. Light DevOps Tasks (spinning up EC2 instances, logging into a servers for diagnosing issues, setting up databases both locally and in the cloud).

Requirements: Advanced proficiency in Python, practical experience with writing solid and well-tested code, working knowledge of SQL, and comfort with Linux a necessity. No need to have a background in math or a CS degree, but the job will involve a lot of quantitative thinking so the applicant should not be afraid of math Working on a small team means doing a little bit of a lot of things. Able to quickly learn and adopt new technologies based on client needs; a typical engagement may include at least one data technology we haven't all worked with before. Must be able to read appropriate documentation in order to write clean, idiomatic code.

Nice-to-have experience: IaaS like Amazon AWS or PaaS like Heroku, Docker, big data tools like Spark and Hadoop, tools design for data-intensive applications e.g. Cassandra, Storm, Elasticsearch, etc.

If interested, send a resume and links to things you'd like us to see (e.g. Github, personal site, blog or projects) to isaac [at] drivendata.org with "HN" in the subject line.

ims··on MathJax CDN shutting down on April 30, 2017
Github code search for 'script src "cdn.mathjax.org"':

> 1,811,496 code results.

Woof.

ims··on Identifying HTTPS-Protected Netflix Videos in Real Time [pdf]
The scraping and automated viewing in question pretty clearly violate Netflix's terms of use. As junior officers in the U.S. Army, the authors are more vulnerable than most to trivial but "correct" accusations of illegal activity, so I wonder if they were at all concerned about the government's sweeping interpretation of the Computer Fraud and Abuse Act.

> In order to generate these fingerprints, we first mapped every available video on Netflix. We took advantage of Netflix’s search feature to do this mapping by conducting iterative search queries to enumerate all of Netflix’s videos. This enumeration was done by visiting https://www.netflix.com/search/<value> where <value> was ‘a’, then ‘b’, etc. and then parsing the returned HTML into a list of videos with matching URLs.

This is not the same as but still in the same class of "unauthorized" use that Weev was charged with carrying out on AT&T endpoints. No privacy concern here, and in theory you are authorized to view this Netflix content but not to "use any robot, spider, scraper or other automated means to access the Netflix service; decompile, reverse engineer or disassemble any software or other products or processes accessible through the Netflix service; insert any code or product or manipulate the content of the Netflix service in any way; or use any data mining, data gathering or extraction method." Though Weev's conviction was vacated on appeal, that was only based on a venue problem so the prosecution's legal theory about violating terms of use still seems to be in play.

Not concern trolling here, I do this sort of scraping all the time and there's no reason to believe the authors are at any risk. It's just an interesting juxtaposition that illustrates how overly broad the DOJ's interpretation of CFAA is, and how selectively it can be pursued. As the EFF notes, one of the major impacts is that is puts security researchers in a legal gray area (https://www.eff.org/issues/cfaa).

ims··on You are the WM
I like minimalism, but IMO the closest you can get to this idea without wasting time reinventing the wheel is a tiling windows manager. I have been using i3 for about 3 years now and can't imagine going back.
ims··on Computational Statistics in Python
Cam Davidson-Pilon's "Bayesian Methods For Hackers" is a great resource. It's actually written as a group of IPython notebooks so you can actually download them and play with the cells to really understand what's going on.

Link: https://github.com/CamDavidsonPilon/Probabilistic-Programmin...

ims··on Why Racket? Why Lisp? (2014)
> (x + (if is_true(): 1 else: 2)) [is invalid in Python] be­cause the if–else con­di­tional is a state­ment, and can only be used in cer­tain positions.

Point taken, but troll-mode pedantry: (x + (1 if is_true() else 2)) would be valid :)

ims··on Warning: Wi-Fi Blocking Is Prohibited
If you want to know more, check out "Three Felonies a Day" by Harvey Silvergate[1]. (I don't necessarily agree with his take on everything, but it's still eye opening and worth a read.)

[1] http://www.amazon.com/Three-Felonies-Day-Target-Innocent/dp/...

ims··on Everything you ever wanted to know about hexagonal grids
That's funny, I brought this exact puzzle up last time this link was posted. Twins! Nice solution.

https://news.ycombinator.com/item?id=5809724

ims··on “Nothing you can do impresses me”
> † A concept I reject, but understand.

I would like to hear more about that.

ims··on Checklist of Rationality Habits
You're talking about whether the MBTI framework is correct, or verifiable, or scientific. That's interesting because the context for this whole discussion is around a rationality checklist which may not be correct, or verifiable, or scientific.

It seems like the more appropriate question when discussing these frameworks might be whether they are useful. Many people find the MBTI framework a useful way to think about themselves. Just like many people find GTD useful for being organized, or Paleo useful for choosing what to eat, or Agile useful for coordinating software development, or [system/lens/framework here] useful for [thing that people do], even though they are not demonstrably "correct" (or even superior to competing systems).

Correct and useful aren't necessarily the same thing, especially when we're discussing systems that are more of a descriptive worldview than an actual set of predictions.

ims··on Python idioms I wish I'd learned earlier
I think the example in #4 misses the point of using a Counter. He could have done the very same for-loop business if mycounter was a defaultdict(int).

The nice thing about a Counter is that it will take a collection of things and... count them:

    >>> from random import randrange
    >>> from collections import Counter
    >>> mycounter = Counter(randrange(10) for _ in range(100))
    >>> mycounter
    Counter({1: 15, 5: 14, 3: 11, 4: 11, 6: 11, 7: 11, 9: 8, 8: 7, 0: 6, 2: 6})
Docs: https://docs.python.org/2/library/collections.html#counter-o...
ims··on Aboard a Cargo Colossus: Maersk’s New Container Ships
Sure, but many of the ships out there today were not built so recently.
ims··on Aboard a Cargo Colossus: Maersk’s New Container Ships
It just means that the waterways are currently too narrow or too shallow for such increasingly massive ships, either naturally or because the deepening simply hasn't caught up to the increasing size of cargo vessels. The U.S. Army Corps of Engineers is involved with many port dredging projects to make the navigable waterways deeper and wider.[1] There has always been an iterative relationship between size/draft of ships and width/depth of waterways -- in fact, the size of many ships the world over is constrained for practical purposes by the size of the Panama Canal.[2]

The US government has been improving ports since its early days. In fact, the expenditure of federal funds on "internal improvements" such as port improvements was quite the contentious issue in the early republic.[3] But most of the world's trade moves by sea, and the US is and always has been a maritime nation.[4] The ROI for port improvements is laughably high, so it has always been a no-brainer.

[1] http://dqm.usace.army.mil/Education/Index.aspx

[2] http://en.wikipedia.org/wiki/Panamax

[3] http://en.wikipedia.org/wiki/Internal_improvements

[4] http://www.stratfor.com/analysis/geopolitics-united-states-p...

← PreviousPage 2 of 5Next →