HNHacker News
TopNewBestAskShowJobs

benhamner

1,488 karma · joined August 4, 2011

twitter.com/benhamner
submissionscomments
benhamner··on Large-Scale Machine Learning for Drug Discovery
We hosted a smaller scale version of this on Kaggle with Merck several years ago: https://www.kaggle.com/c/MerckActivity

It's great to see that competition's result both supported and extended!

benhamner··on Uber Expectations as We Grow
As Uber grows in a city, the average wait time should also decrease pretty dramatically. There are more cars on the road, so a pickup from any given location/time should arrive faster on average than it would two years ago.

The precise methodology is opaque in the post, but there's a chance that these results are driven more by increasing liquidity in the marketplace instead of increasing expectations of Uber over time.

benhamner··on Facebook was down
$400 per second is about $13 billion per year
benhamner··on Ask HN: The purest skills needed to become a programmer?
Communication. Excellent and precise communication.
benhamner··on Ask HN: How do you stay motivated to work on side projects?
Set a goal of making a bit of progress every day. Even 10-15 minutes a day adds up over time, and it keeps you in the habit of making ongoing improvements.
benhamner··on Female Founders Conference 2015 applications are open
This is phenomenal & a huge thanks to YC for promoting this. The massive gender disparity in tech is painfully clear to anyone walking the halls of a SF startup, attending a team meeting at a large corporation, or observing . It is a multifaceted and highly complex challenge, but hopefully strong leadership on this issue from central players like YC will help things move in a positive trajectory.

To those saying "what about an Latino/African American/etc. founders conference?": you can't be everything to everyone, and you have to start somewhere. Looking at this as a binary world view where "you can only have nonexclusive events" or "you need to have an exclusive event for every potentially underrepresented group" is counterproductive and gets us nowhere. Supporting one underrepresented group will hopefully have positive downstream effects across the board. If things like this are successful in moving one underrepresented group on a better trajectory, then that model can be more easily replicated.

To those saying "why not a YC conference just for male founders?" or "this isn't necessary" or the like: you're being petty and this isn't for you. The tech industry has plenty of open opportunities for people to connect across the board, and stronger connections tend to be forged within smaller sub-communities. If this and events like this encourage more females to become founders, then that's a great outcome.

benhamner··on Mattermark (YC S12) raises a $6.5M Series A to keep VCs in the know
Congratulations! Y'all have had a wild ride so far, for being among the coolest and most practical companies in the data space
benhamner··on Do We Need Hundreds of Classifiers to Solve Real World Classification Problems? [pdf]
This is consistent with our experience running hundreds of Kaggle competitions: for most classification problems, some variation on ensembled decision trees (random forests, gradient boosted machines, etc.) performs the best. This is typically in conjunction with clever data processing, feature selection, and internal validation.

One key exception is where the data is richly and hierarchically structured. Text, speech, and visual data falls under this category. In many cases here, variations of neural networks (deep neural nets/CNN's/RNN's/etc.) provide very dramatic improvements.

This study does have a couple limitations. The datasets used are very small & form a very biased selection of real-world applications of machine learning. It doesn't consider ensembles of different model types (which I'd expect to provide a consistent but marginal improvement over the results here).

benhamner··on Ask HN: Founders whose startups failed, do you experience bias in interviews?
You mean lie on the resume / hide their level of responsibility?

That doesn't make any sense whatsoever (beyond being dishonest, it would likely be surfaced in reference checks and stop a hire).

I'd see the startup experience as very positive, assuming they could articulate why the startup failed and learned from the experience. There's many startup CEO's who would make (and are) great sales, strategy, or product leaders, and CTO's who would be great fits for engineers, architects, or engineering managers.

benhamner··on Ask HN: Who is using the .NET stack for their startup?
We use C# and F# extensively at Kaggle: kaggle.com is an ASP.NET MVC app on Azure, and a good portion of our data processing is in C# or F#.
benhamner··on Hacker News API
Suggestions for improving the API, to make it more valuable for data mining and analytics. This assumes more historical data is available.

1. Provide a way to bulk download the data (that's a click, instead of scraping the API)

2. Add a field for the maximum position a story reached on the front page

3. Add the numerical score for the comment (at least on comments that are N days old, which won't interfere with the reason to hide the scores on the main comments)

Some other changes that would be awesome (but are less realistic) include:

4. A historical event log of votes (even better would be relating those votes back to users, but I imagine that's not going to happen for privacy reasons. An intermediate possibility would be a vote log connected back to anonymized user ids, assuming the anonymous id -> real user id mapping is difficult)

5. A historical event log of display position changes for stories & comments

6. An event log of pageviews with as much metadata as possible to release without infringing on privacy

benhamner··on Ask HN: Horror co-founder stories.
Make sure you do background checks (references + backchannel).
benhamner··on Apple Close to Buying Beats Electronics for $3.2 Billion
If this is true, who benefits from leaking it & in what way do they benefit?
benhamner··on Employee Equity
There's several factors that make this irrelevant.

1) YC's incentives look very different from VC's: they get common shares, not preferred. This means their incentives are more aligned with the founders than with future investors (for example, if a VC has a controlling stake in the company & wants to fire the founder and dilute the common shares to basically nothing, then YC gets similarly diluted).

2) The dilution effect of the option pool on YC's shares is trumped by the dilution effect of future investments on YC's shares. If expanding the option pool has a marginal dilution affect but dramatically increases the likelihood of success, then that's a no-brainer for YC to push for.

3) YC's business model is dominated by the extreme outlier successes (e.g. Dropbox, AirBnb). Thus, YC does better by doing these three things better:

A. Increasing the likelihood that future successes are funded by YC (i.e. the founding team chooses YC early on)

B. Increasing the probability that a startup will become an outlier success.

C. Given that a startup is becoming an outlier success, multiply that success to the extent possible.

This piece hits nicely at each of those points. For A, YC takes a leadership role in how to structure a cap table, making founders look more to YC. Also, YC startup employees (ie future YC founders) think better of YC. For B and C, once a company grows beyond the founders, each employee makes very meaningful decisions on a daily basis that impact both the company's likelihood of success and magnitude of success. Aligning these employee's motivations with the company's further helps make these decisions better for the company.

benhamner··on How Can Yahoo Be Worth Less Than Zero?
"They could reassign some of their brightest engineers to write trading algorithms."

Unfortunately for that strategy, their last CEO effectively got rid of all their good machine learning guys.

benhamner··on Ethereum: A Turing-Complete Cryptocurrency
Text content: http://pastebin.com/NCGRv74u

(Saw comments that the site went down and still had it loaded on my machine)

benhamner··on Google’s $179 Moto G puts every single cheap Android phone to shame
>don't understand why anyone would choose top end phones from the brands like Samsung or Apple, if half of the cost of a phone goes to profit their shareholders.

When consumers spend money to buy products, where the money goes (whether it's costs of production or shareholder profit) almost never factors into the purchase decision.

benhamner··on Build an algorithm to predict friendships, then actually use it to meet people
Ping me (b@kaggle.com) if you're interested in running this competition more formally on https://kaggle.com.

We've run hundreds of machine learning competitions & offer a real-time leaderboard to encourage competitive participation, a very active community of data scientists, and many other features that simplify running this type of challenge.

benhamner··on 2013 NIPS Proceedings – Advances in Neural Information Processing Systems
Ha, potentially. The dangers of drawing conclusions from N=2.
benhamner··on 2013 NIPS Proceedings – Advances in Neural Information Processing Systems
This is an interesting case study in how the submitted headline affects the page's rank on HN.

I'd originally submitted this earlier today, with the headline "Neural Information Processing Systems (NIPS) 2013 Proceedings": https://news.ycombinator.com/item?id=6815771

It only got one other upvote and never made it to the front page. Meanwhile, another article submitted at almost exactly the same time with far less interesting content but a more provocative headline (on a user getting banned from Uber for API abuse) got ~15 votes, pushing it onto the front page.

I re-submitted this page with a slightly more descriptive (and buzzwordy) headline ("State of the art Machine Learning papers: NIPS 2013"), and it almost immediately ended up on the front page. Since then the headline has been reverted to the page's headline, again slowing the rate at which it's received votes.

benhamner··on 2013 NIPS Proceedings – Advances in Neural Information Processing Systems
I made a nicer version of this to browse for the ICML 2013 papers, based on a similar thing Andrej Karpathy did for the NIPS 2012 papers: http://benhamner.com/icml2013preview/#/

Don't believe anyone's done this for NIPS 2013 yet.

benhamner··on Show HN: Game where you write Python robots to fight other players
Brandon - what methods are you using to run untrusted Python code in a secure sandbox?
benhamner··on George Orwell: Politics and the English Language (1946)
One of my favorite essays. I re-read this every six months as a reminder to be as clear and precise as possible in my communication.
benhamner··on AI Startup Says It Has Defeated Captchas
This is cool, but there's no indication from the article that it's novel, or that it's better than existing methods.

The article linked to 28 other different systems that have claimed to beat / demonstrated beating captchas at some point: http://www.karlgroves.com/2013/02/09/list-of-resources-break....

Without a performance comparison to existing methodologies on a benchmark dataset and precise details on the models, this is a neat marketing demo and nothing more.

benhamner··on LinkedIn Intro: Doing the Impossible on iOS
The privacy outrage around this is nonsensical.

Over 500 million people trust Google with complete and indefinite access to their email. The leap from trusting no external email providers to trusting Gmail is much greater than this incremental step of trusting LinkedIn as well. The risk is similar to trusting an established company to automatically backup your emails, and smaller than trusting startups like Greplin (which rebranded and got acquired) to safeguard a dump of all your emails.

This is not to say the privacy and uptime risks are non-existent: the attack surface area is marginally increased and there is another system that could break.

Claiming LinkedIn's doing a "MITM attack on your email" is on the same level as saying "Google is Big Brother." Both statements capture an element of reality, but with an extremely alarmist bent.

benhamner··on Make $377,000 trading Apple in one day
"According to his study, in one day (May 9), playing one stock (Apple (AAPL)), Hendershott walked away with almost $377,000 in theoretical profits by picking off quotes on various exchanges that were fractions of a second out of date."

This is a massive red flag. There's so many things that could be done incorrectly that the result is meaningless if the profits were only theoretical instead of observed.

benhamner··on BlinkDB: Queries with Bounded Errors and Response Times on Very Large Data
I'm thrilled that AMPLab and CSAIL are building this.

For the vast majority of analytics problems and projects I've worked on, approximate numbers are just as good as exact results. One of the biggest productivity blockers can be queries and analytics that take hours instead days to run, instead of seconds to minutes, as these dramatically decrease the number of iterations you can execute and ideas you can test.

We commonly work on sub-sampled versions of datasets to enable interactive queries and analytics - it's really great to see someone formalizing this process and handling the details in a simple and principled manner.

benhamner··on A Look Inside Our 210TB 2012 Web Corpus
The massive advantages that Google has include over a decade of data on the pages that people actually visited in response to a specific query as well as having an in-memory index of the public web, parts of which are updated on the order of seconds to minutes.

I wonder if there is a viable business in maintaining an in-memory & up-to-date index of the public web & selling access to it, with a pricing model that scales according to the amount of computation you are doing on it.

benhamner··on Infochimps Acquired By CSC
https://twitter.com/nickducoff/status/364861776363388929
benhamner··on Basic Neural Network on Python
You may be interested in this ICML 2006 paper, which empirically compared many standard algorithms across a combination of metrics and UCI datasets - http://www.cs.cornell.edu/~caruana/ctp/ct.papers/caruana.icm...
← PreviousPage 3 of 4Next →