HNHacker News
TopNewBestAskShowJobs

benhamner

1,488 karma · joined August 4, 2011

twitter.com/benhamner
submissionscomments
benhamner··on Kaggle Datasets – Discover and analyze open data
Thanks and great point! Added this to our list
benhamner··on Kaggle Datasets – Discover and analyze open data
Not all the datasets are ML specific, but hopefully this helps:

- https://www.kaggle.com/tags

- https://www.kaggle.com/tags/linguistics

- https://www.kaggle.com/tags/multiclass-classification

- https://www.kaggle.com/tags/text-data

benhamner··on Kaggle Datasets – Discover and analyze open data
Thanks for the feedback. This is likely a "not quite yet" vs. "never".

Definitely understand the motivations from a user standpoint for not needing to login to download.

There's some non-obvious benefits we get as a small team by requiring login, in addition to new user growth. Bandwidth for hosting data can be large, and it's easier to reason about and prevent abuse in the context of authenticated users.

We do enable previewing the dataset while logged out, and the preview functionality will become more full-featured.

benhamner··on Kaggle Datasets – Discover and analyze open data
Our goal with Kaggle Datasets is to provide the best place to publish, collaborate on, and consume public data.

As a data publisher, you have an easy way to publish data online, see how it's used, and interact with the users of the data. You can create the dataset via a simple web interface, and update it through the interface or an API. We automatically version these updates under the hood.

As a data consumer, you can browse the data online and download it (through the web or an API). You can see the code and insights others have generated on the data through Kaggle Kernels (hosted, versioned IPython notebooks that run in Docker containers). You can fork their code to get started on the data, or start coding from scratch on your own analysis. If you find improvements that could be made to the metadata (dataset/file/column-level descriptions), you can make those directly.

We're rapidly iterating on this product and expanding it's functionality, and would love any feedback and suggestions.

benhamner··on AWS Public Datasets
Over 13,000 community-uploaded public datasets: https://www.kaggle.com/datasets

(I work at Kaggle)

benhamner··on The State of Data Science and Machine Learning
All Kaggle problems aren't created equal. Some look like a train matrix, a single target, and a test matrix.

Others are far more complex and start with much messier data and/or complex formulations.

Examples:

- www.kaggle.com/c/nips-2017-non-targeted-adversarial-attack/ - www.kaggle.com/c/the-allen-ai-science-challenge

benhamner··on The State of Data Science and Machine Learning
Completely agree. Data quality issues was a big part of our motivation with Kaggle Datasets (an open data platform where the quality of the dataset improves as more people use it) and Kaggle Kernels (a reproducible data science workbench that combines versioned data, code, and compute environments to create reproducible results).

Two examples of this: Kaggle Datasets supports wiki-like editing of metadata (file and column descriptions) and makes it easy to see, fork, and build on all the analytics created on the data so far.

We're just getting started with each of these products: we want Kaggle Datasets to support a fully collaborative model around working with all your data in the future, and Kaggle Kernels to support every analytics and machine learning usecase.

benhamner··on Artificial Intelligence Software Is Booming, But Why Now?
Ready access to high-quality usecases and training data, along with shared knowledge of the methods that work well on these helps: https://www.kaggle.com/competitions?sortBy=numberOfTeams&gro...

(disclaimer: I work at Kaggle)

benhamner··on R for Data Science
Great book!

We also have almost 10,000 forkable & executable R examples on Kaggle (https://www.kaggle.com/kernels - select R from languages). Almost all of these use at least one of Hadley's libraries

benhamner··on Y Combinator Companies
I scraped a CSV version of the data & made a graph of the verticals by batch: https://www.kaggle.com/benhamner/y-combinator-companies
benhamner··on How to Become a Data Scientist, Part 2
We also are growing https://www.kaggle.com/datasets, which won't necessarily have clean data, clear problem statements, and a well-defined task.
benhamner··on Data Science Competitions 101: Anatomy and Approach
The edge is almost never the choice of model architecture. It's very common knowledge now that deep neural networks and gradient boosted machines are incredibly effective at different problem classes, and that ensembling almost always marginally boosts performance.

The edge winning teams have varies from competition to competition. It includes

- robust cross-validation strategies - robust feature selection strategies - creative feature engineering - finding uniquely valuable external datasets that improve performance - robustly controlling for distributional differences between train and test sets - (and unfortunately on occasion) information leakage

benhamner··on The Startup Zeitgeist
Are you in a position to share a raw, granular form of the data (potentially with more anonymization)?
benhamner··on OpenAI Gym: Toolkit for developing, comparing reinforcement learning algorithms
Agree! Down the road, Kaggle will be the Kaggle for RL algorithms ;) This provides a really cool open source environment to build on
benhamner··on Google launches Public Datasets program
In a similar vein, Kaggle datasets enables you to run Python, R, Julia, and SQL on many public datasets https://www.kaggle.com/datasets
benhamner··on TensorFlow Implementation of Deep Convolutional Generative Adversarial Networks
Can you send an introduction and a sample of the data? I'm interested in publishing it on Kaggle (b at kaggle . com)
benhamner··on I Am Sam Altman, President of Y Combinator. AMA
How much traffic per day does hacker news typically get?
benhamner··on Investigate switching away from GitHub
It's surprising just how silent GitHub is on these issues. Most of their engineers probably read Hacker News, yet none of them respond.
benhamner··on In lawsuit over hacking, Uber probes IP address assigned to Lyft exec
One interesting piece of this: it looks like Github turned over IP traffic logs for who accessed that page of Uber's Github repo to Uber.
benhamner··on R beats Python, R beats Julia, Anyone else wanna challenge R? (2014)
Have you used R?

As both a heavy R user and a software engineer, I can promise you that one of the quintessential aspects of R is it's actually "written by statisticians, for statisticians."

You can't accuse R of having great code and language design. Or good code and language design. Or even mediocre code and language design.

Imagine what you would get if you got a million monkeys drunk, put them on a roller coaster with laptops, and had them bang keys while they were upside down on loops. And then the result suddenly, miraculously runs and produces output. Now you understand R's software design.

benhamner··on Report: U.S. average daily mobile game time drops over 30% in a year
This could potentially be explained by the amount of time non-gamers spend on mobile devices increasing dramatically relative to gamers.

This explanation could make sense: gamers tend to be early adopters of new technologies (including mobile), and then make up a smaller fraction of the total users as technologies become mainstream.

Can someone with access to the data support or reject this explanation?

benhamner··on U.S. Paychecks Grow at Record-Slow Pace
The surprising thing here is that the NYTimes left out a graph and a direct link to source data in this article.
benhamner··on Palantir Technologies Raises $450M
Looking at Palantir only as a software company is incredibly misleading. They sell high-ticket custom solutions to massive organizations that they leverage many of their internal software tools and libraries to build. This involves large amounts of custom product and engineering development, ingesting complex data sources, integrating with operational processes, and communicating results and insights back to the customer in a way that changes the customers behavior and adds value.
benhamner··on Facebook’s Piracy Problem
A more accurate title: Youtube's Facebook Piracy Problem
benhamner··on Escher – Build beautiful interactive Web UIs in Julia
Beautiful! Appears inspired by RStudio's Shiny http://shiny.rstudio.com/
benhamner··on Airbnb Engineering
I don't work for Airbnb, but most top notch engineering organizations I know have far more unsolved problems than solved problems.

Each additional magnitude of scale also brings a new wave of engineering challenges and opportunities.

I'd be shocked if this didn't apply to Airbnb. This is just as much a teaser of what's to come as it is a shrine to what 's been built.

(I personally have don't have the SF requirement nor any comment on that - Kaggle fully supports and promotes a distributed workforce)

benhamner··on Do We Need Hundreds of Classifiers to Solve Real World Classification Problems? [pdf]
Random Forests are great on many tasks, but this analysis is incredibly biased: it only includes the incredibly small and simple datasets in the UCI repository. Many real world tasks are far more complex than that, especially those involving text, speech, images, video, and large scale web data.
benhamner··on Amazon Machine Learning – Make Data-Driven Decisions at Scale
You hit the nail on the head. Completely agrees with all my experience at Kaggle and applying machine learning across a broad number of industries
benhamner··on Programming competitions correlate negatively with being good on the job
This is flat out wrong.

It disagrees with every other data point I have, so I'm very skeptical with both the methodology (which is opaque) and the conclusion.

From all of my experience at Kaggle (we run machine learning competitions), with our community, and from being close to programming competition sites & understanding their communities, doing great at competitions is an unambiguously positive signal.

(It's worth noting that doing great at competitions is only a positive signal - the lack of competitions is by no means a negative signal).

Many of our customers have found that their best hires have come from competitions. In a lot of cases, this surfaces candidates that would normally be completely overlooked because they don't fit the "top tier CS school" mold that recruiters commonly overfit to.

Several companies have had a successful recruiting strategy built on poaching our top users (https://www.kaggle.com/users).

Peter Norvig's criticism that "programming contest winners are used to cranking solutions out fast and that you performed better at the job if you were more reflective and went slowly and made sure things were right" is specific to programming competitions with very short time durations (vs. the machine learning competitions that I'm used to running, which typically last months and incentivize solutions that generalize well).

However, we've seen that many programming competition winners also do well on machine learning competitions, and the same qualities that aid in competitive programming (creativity, efficiency, tenacity, fluidity with tools, and the ability to build something that works) help win machine learning competitions.

benhamner··on Google Maps Pacman
NOW WE KNOW how the self-driving cars work!
← PreviousPage 2 of 4Next →