HNHacker News
TopNewBestAskShowJobs

benhamner

1,488 karma · joined August 4, 2011

twitter.com/benhamner
submissionscomments
benhamner··on Basic Neural Network on Python
Both datasets you used (iris and digits) are way too simple for neural networks to shine.

Neural networks / deep neural networks work best in domains where the underlying data has a very rich, complex, and hierarchical structure (such as computer vision and speech recognition). Currently, training these models is both computationally expensive and fickle. Most state of the art research in this area is performed on GPU's and there are many tuneable parameters.

For most typical applied machine learning problems, especially on simpler datasets that fit in RAM, variants of ensembled decision trees (such as Random Forests) to perform at least as well as neural networks with less parameter tuning and far shorter training times.

benhamner··on Harry Potter and the Methods of Rationality just got several fascinating updates
If you've not read this, strap yourself in - you're in for a ride. Eliezer, an AI researcher, has created this novel from a simple yet fascinating standpoint - what if Harry was a brilliant rationalist who engaged the magical world from a scientific standpoint? Don't dismiss this because it's a fanfic or incomplete - it introduces rationality, breaks many fantasy and geek tropes, and builds off Rowling's universe in a highly entertaining and thought-provoking page turner (I guess a modernization of this idiom would be "iPad flicker").
benhamner··on Physicists To Test If Universe Is A Computer Simulation
Are they using a pseudo-random number generator with a set seed, or is there a better source of randomness? ;)

I treat free will as an assumption:

- If we are, in fact, predestined, assuming I have free will doesn't hurt or matter.

- If we do have free will, I'm applying that free will to make decisions based on a correct assumption

- If we do have free will and I had assumed we were predestined, this could have worse results. At the very least, it is unlikely to improve outcomes

benhamner··on Physicists To Test If Universe Is A Computer Simulation
If you're interested in going a bit deeper on this, Nick Bostrom (http://nickbostrom.com/) has pulled together an intriguing case: http://www.simulation-argument.com/simulation.html

A provoking question: if we found evidence that convinced us, beyond a reasonable doubt, that we were living in a computer simulation, would you change your behavior?

benhamner··on Reasons Not to Stretch
Studies like this make you question how much of our folk knowledge and wisdom is false. I stretch before and after my daily runs under the assumption that this has been reducing my injury risk. From these results, I should stop wasting time on pre-run stretches.

Hopefully we'll see more studies built by attempting to support or reject the folk wisdom we've assumed throughout our lives that haven't been subjected to rigorous scientific scrutiny through controlled experimentation. This could be the foundation of very impactful scientific careers.

One challenge would be communicating and disseminating results of such studies to the public. Many times (especially where diet and lifestyle choices are concerned) the results would run counter to large and entrenched commercial interests with enormous marketing budgets and expertise at influencing consumer choices. Countering this influence as an individual researcher or cash-strapped government agency isn't trivial, to say the very least.

benhamner··on Tableau files for IPO
Whoops! Unfortunate typo. Hopefully a mod will fix
benhamner··on Tableau files for IPO
Interesting stats from their S-1 (http://edgar.sec.gov/Archives/edgar/data/1303652/00011931251...):

- $127.7m 2012 revenue

- 749 full time employees at end of 2012

- 70% of revenue from product licenses; 30% from services

- 11,000 customer accounts

- No customer represented more than 5% of total revenue

benhamner··on Ask PG: Do you ever get overwhelmed?
A different way to frame the question: "what keeps you up at night?" (beyond the very uncommon event of HN being down)

Massive quantity of emails? YC startups that aren't the rare runaway success? YC startups that are the rare runaway success? Meta issues around YC? Something unrelated to YC? Pressure to live up to the pg legend irl? Not being disciplined about avoiding overworking? Something in your private life that's not appropriate to share? Anything?

benhamner··on Every recorded meteorite strike on Earth since 2,300 BCE mapped
More accurately described as "where people were located since 2,300 BCE that were capable of recording meteor strikes and recorded them in a form that survived until present day and made its way into a database of meteor strikes." Latitudinal variations in the density of meteor strikes wouldn't surprise me; longitudinal ones should be entirely explained by the observational effect. Hence we see a ton of meteor strikes in the continental US and virtually none in rural China, though both are on similar latitudes.
benhamner··on Show HN: Analytics.js – The analytics API you've always wanted
Web analytics just got meta
benhamner··on IPython gets $1.15M funding
Thrilled to see this. IPython Notebook has become my go-to tool for data munging and analytics. I look forward to seeing the IPython team take it to the next level.
benhamner··on Introducing the Predictive Interface
Very curious to see how well this works in practice.

An excellent example of a usable predictive interface is DataWrangler from the Stanford Vis group - http://vis.stanford.edu/wrangler/app/. Bit of a learning curve, but well worth it if you're working with semi-structured text data.

The paper describing their philosophy and the predictive interface is here: http://vis.stanford.edu/files/2011-Wrangler-CHI.pdf

benhamner··on Why becoming a data scientist is not easier than you think
Well written, but I believe you missed the point of the original article.

No one ever claimed that taking one class made someone an expert "data scientist." Instead, that single class wetted Luis, Jure, and Xavier's (the three competition winners) appetites, and pushed them more to learn more about machine learning and natural language processing. They then went on to dive much deeper, and excelled specifically in one area of applied NLP.

However, without that first class, there's a good chance none of them would have ever focused on (or heard of) machine learning. Their story is growing increasingly common. Like the Netflix Prize, Andrew Ng's first Coursera class did its part in shining a spotlight onto our dark little corner of the universe.

I'd be very cautious about a long checklist of items that are necessary to be a successful data scientist (which is a pretty ill-defined and encompassing term at this point). That is a decent summary of many useful tools of the trade, but they are by no means useful for all problem domains. For example, I could spend years working on machine learning for EEG brain-computer interfaces without a good reason to use databases or "big data" NoSQL technologies. I especially enjoyed MSR's take on the matter in "Nobody ever got fired for using Hadoop on a cluster" http://research.microsoft.com/pubs/163083/hotcbp12%20final.p...

When we're hiring data scientists or seeking successful ones, we've found focusing on demonstrated excellence in one relevant area plus general quantitative competencies and the curiosity and tenacity to learn new tools and techniques works far better than a laundry list of skills and experiences.

benhamner··on Becoming a data scientist might be easier than you think
When we host competitions on Kaggle, a lot of work goes into asking the right question and structuring the problem. The domain expertise is incorporated in this step, as well as in putting the competition results to use in production.

This splits the "domain expertise" and "predictive modeling" components into two separate chunks. While domain expertise is crucial for asking the right questions, we've found that it isn't as necessary for the predictive modeling component. For example, in the essay scoring contest we hosted, none of the winners had touched natural language processing prior to the contest. However, they beat out many experts with decades of experience in NLP.

For an internal data science team, the "domain expertise" component is at least as important, as they are charged with asking the right questions as well. However, this does not mean competition winners cannot develop and learn this - they have already demonstrated their creativity and tenacity in one domain (applied machine learning), and this carries over nicely to other domains from our experience.

benhamner··on The rise of LinkedIn’s news feed
How are you evaluating feed relevance?
benhamner··on Stop Applying to Startups
Shoot me a message if you're interested in Kaggle - ben@kaggle
benhamner··on Which machine learning classifiers are fast enough for medium-sized data?
Interesting comparison, but only using generated mixtures of gaussians for training data severely limits any conclusions that can be drawn from this. Naturally the method with the same assumptions as the generating process had the best performance.

It is important to note that both the performance of the machine learning algorithm (in terms of the error metric) and its runtime are very dependent on the source data in most cases.

benhamner··on Machine Learning in Python Has Never Been Easier
What do you need?
benhamner··on Netflix never used its $1 million algorithm due to engineering costs
In the Netflix prize, did all teams merge into one group and stop trying? Nor would we expect so with K groups, especially with exponentially decreasing prize amounts. If this is even slightly a risk, you can compensate by making the maximum number of prizes awarded a function of the number of participating teams. For example, you can award 10 prizes if there are >100 teams, and if not award floor(num_teams/10) prizes, with the prize pool redistributed among the top 10%.

It was pure luck that the threshold for the million-dollar prize was crossed. If it had been arbitrarily set at 11% (as opposed to 10%), then there's a good chance the million dollar prize would have never been paid out.

The advantage of paying out K top prizes is that other teams that don't win outright may have developed additional useful models or insights (or used more computationally efficient algorithms), and you may access to these in this manner.

benhamner··on Netflix never used its $1 million algorithm due to engineering costs
This is why you should award prizes for the top K performers, as opposed to the first one to cross a benchmark.
benhamner··on Netflix never used its $1 million algorithm due to engineering costs
As Eliezer said - they implemented the two most important algorithms from the contest, which advanced the state of the art in recommendation systems & gave the majority of the benefits, and didn't implement the long tail of algorithms that each only gave a very slight marginal benefit & would have been costly to re-train and maintain.

This is one of the advantages of running shorter competitions: normally it takes 1-3 months to approximately hit the asymptotic level of performance on a dataset given the inherent noise in it and the state of the art in machine learning. The shorter competitions are focused on finding the low-hanging fruit that generate large improvements (such as SVD & RBM's in Netflix's case) and exploring the space of possible model structures, as opposed to optimally ensembling across a large number of models to eek out the last 0.01% of performance.

Exploring the space of useful features & possible models enables you to trade off computational efficiency & maintainability vs. model performance in production as well. The $1 million dollars Netflix put to the prize leveraged >> $1 million in human effort to explore the possible models, from which they found and applied the two best suited for their production implementation.

(disclaimer - I work with Kaggle)

benhamner··on Ask HN: Help a Hacker Out
"I'm particularly interested in working on challenging problems dealing with algorithms and big data"

Give one of our competitions a shot - http://kaggle.com/. They're a good way to experiment with algorithms and machine learning on well-defined problems, and collaborate with people who have similar interests. Also, we love hiring people who win competitions.

benhamner··on Machine Learning using Quantum Algorithms
This is an old article (from 2009). Hartmut Neven provided an update at ICML 2011 in the latter part of his keynote talk - http://techtalks.tv/talks/54457/
benhamner··on Official enrollment for Stanford's online AI class has begun
I imagine they want demographic data on who's taking the course. If you were running a large-scale education experiment, wouldn't you want to be able to measure how things are going and account for variables such as gender, age, and education level? I bet they are going to look at location as well, via IP addresses.
benhamner··on Calculize: A Mathematical Scripting Language
Nice. A good next step would be to add an interactive window, like the main window in Matlab.
← PreviousPage 4 of 4