1,488 karma · joined August 4, 2011
- https://www.kaggle.com/tags/linguistics
Definitely understand the motivations from a user standpoint for not needing to login to download.
There's some non-obvious benefits we get as a small team by requiring login, in addition to new user growth. Bandwidth for hosting data can be large, and it's easier to reason about and prevent abuse in the context of authenticated users.
We do enable previewing the dataset while logged out, and the preview functionality will become more full-featured.
As a data publisher, you have an easy way to publish data online, see how it's used, and interact with the users of the data. You can create the dataset via a simple web interface, and update it through the interface or an API. We automatically version these updates under the hood.
As a data consumer, you can browse the data online and download it (through the web or an API). You can see the code and insights others have generated on the data through Kaggle Kernels (hosted, versioned IPython notebooks that run in Docker containers). You can fork their code to get started on the data, or start coding from scratch on your own analysis. If you find improvements that could be made to the metadata (dataset/file/column-level descriptions), you can make those directly.
We're rapidly iterating on this product and expanding it's functionality, and would love any feedback and suggestions.
(I work at Kaggle)
Others are far more complex and start with much messier data and/or complex formulations.
Examples:
- www.kaggle.com/c/nips-2017-non-targeted-adversarial-attack/ - www.kaggle.com/c/the-allen-ai-science-challenge
Two examples of this: Kaggle Datasets supports wiki-like editing of metadata (file and column descriptions) and makes it easy to see, fork, and build on all the analytics created on the data so far.
We're just getting started with each of these products: we want Kaggle Datasets to support a fully collaborative model around working with all your data in the future, and Kaggle Kernels to support every analytics and machine learning usecase.
(disclaimer: I work at Kaggle)
We also have almost 10,000 forkable & executable R examples on Kaggle (https://www.kaggle.com/kernels - select R from languages). Almost all of these use at least one of Hadley's libraries
The edge winning teams have varies from competition to competition. It includes
- robust cross-validation strategies - robust feature selection strategies - creative feature engineering - finding uniquely valuable external datasets that improve performance - robustly controlling for distributional differences between train and test sets - (and unfortunately on occasion) information leakage
As both a heavy R user and a software engineer, I can promise you that one of the quintessential aspects of R is it's actually "written by statisticians, for statisticians."
You can't accuse R of having great code and language design. Or good code and language design. Or even mediocre code and language design.
Imagine what you would get if you got a million monkeys drunk, put them on a roller coaster with laptops, and had them bang keys while they were upside down on loops. And then the result suddenly, miraculously runs and produces output. Now you understand R's software design.
This explanation could make sense: gamers tend to be early adopters of new technologies (including mobile), and then make up a smaller fraction of the total users as technologies become mainstream.
Can someone with access to the data support or reject this explanation?
Each additional magnitude of scale also brings a new wave of engineering challenges and opportunities.
I'd be shocked if this didn't apply to Airbnb. This is just as much a teaser of what's to come as it is a shrine to what 's been built.
(I personally have don't have the SF requirement nor any comment on that - Kaggle fully supports and promotes a distributed workforce)
It disagrees with every other data point I have, so I'm very skeptical with both the methodology (which is opaque) and the conclusion.
From all of my experience at Kaggle (we run machine learning competitions), with our community, and from being close to programming competition sites & understanding their communities, doing great at competitions is an unambiguously positive signal.
(It's worth noting that doing great at competitions is only a positive signal - the lack of competitions is by no means a negative signal).
Many of our customers have found that their best hires have come from competitions. In a lot of cases, this surfaces candidates that would normally be completely overlooked because they don't fit the "top tier CS school" mold that recruiters commonly overfit to.
Several companies have had a successful recruiting strategy built on poaching our top users (https://www.kaggle.com/users).
Peter Norvig's criticism that "programming contest winners are used to cranking solutions out fast and that you performed better at the job if you were more reflective and went slowly and made sure things were right" is specific to programming competitions with very short time durations (vs. the machine learning competitions that I'm used to running, which typically last months and incentivize solutions that generalize well).
However, we've seen that many programming competition winners also do well on machine learning competitions, and the same qualities that aid in competitive programming (creativity, efficiency, tenacity, fluidity with tools, and the ability to build something that works) help win machine learning competitions.