EvalAI: An Open-Source Alternative to Kaggle
eval.ai
eval.ai
"To my mind, the crucial but unappreciated methodology driving predictive modeling’s success is what computational linguist Mark Liberman (Liberman 2010) has called the Common Task Framework (CTF). An instance of the CTF has these ingredients:
(a) A publicly available training dataset involving, for each observation, a list of (possibly many) feature measure- ments, and a class label for that observation.
(b) A set of enrolled competitors whose common task is to infer a class prediction rule from the training data.
(c) A scoring referee,to which competitors can submit their prediction rule. The referee runs the prediction rule against a testing dataset, which is sequestered behind a Chinese wall. The referee objectively and automatically reports the score (prediction accuracy) achieved by the submitted rule.
...
The general experience with CTF was summarized by Liberman as follows:
1. Error rates decline by a fixed percentage each year, to an asymptote depending on task and data quality.
2. Progress usually comes from many small improvements; a change of 1% can be a reason to break out the champagne.
3. Shared data plays a crucial role—and is reused in unex- pected ways.
...
The author believes that the Common Task Framework is the single idea from machine learning and data science that is most lacking attention in today’s statistical training.
https://www.tandfonline.com/doi/full/10.1080/10618600.2017.1...hm, hadn't heard this one before (https://en.wikipedia.org/wiki/Chinese_wall). Seems a bit anachronistic.
I might add that the comparison to Kaggle is added by the OP, I don't see it mentioned on the website anywhere.
It's on their github page,
That said, my point still stands. Additionally, the comparison they did is massively unfair. I'm pretty sure Kaggle competitions have most of the feature they claim it doesn't have (ex: multiple phases, custom metrics, evaluation in environments).
No sweat, we can't know everything, god save the interwebs!
> the comparison they did is massively unfair.
Agreed. This is all puff
Custom metrics Multiple phases/splits Remote evaluation Human evaluation Evaluation in Environments
The actual rankings are: #1 Kaggle #2 DrivenData
Honorable mention but poorly managed: AiCrowd
No one else has any level of funding to incent performance.
I guess that explains the access to computing power for all the notebooks.
The about page [1] for this project says,
> With EvalAI, we want to standardize the process of evaluating different methods on a dataset
Standardization often does come from openness, I'm just not sure juxtaposing against kaggle is the right move.
The lion's share of Kaggle's effort lies in their expertise in (1) making deals, and (2) making sure the challenge is legit (no bias in training data, not too easy, etc.). That's not easily replicable, so while an "Open-Source alternative" sounds good up front, the bulk of the work remains.
AIcrowd(https://www.aicrowd.com/) in that sense is interesting, because it hosts research challenges like the ones that EvalAI does and it also hosts more attractive Kaggle-like challenges. It's also the most modular platform to accommodate unorthodox evaluation and submission formats.