Show HN: Concept of a marketplace for machine learning datasets
datapie.in
datapie.in
As someone who likes making analyses from random datasets, I have a few issues with these types of services:
1) There is often no indication of the distribution rights of the data, or whether the data was obtained ethically from the source (i.e. following the ToS). I made this mistake when I used an OKCupid dataset released on an Open Data Repository; turns out it was scraped with a logged-in account and the dataset was taken down by DMCA
2) There is no indication of the quality of the data, and as a result, it may take an absurd amount of time cleaning the data for accuracy. Some datasets may not be salvageable.
3) Bandwidth. Good datasets have lots of data for better models, which these sites may not be able to support. (BigQuery public datasets solve this problem however)
The system will be bootstrapped with lots of interesting big datasets: imagenet, 10m images, youtube 8m, reddit comments, hn comments, etc. Our experience is that we need a central point for researchers to get easy access to open datasets that doesn't require a AWS or GCE account.
I recommand using true skill for noting user and dataset.
Keep up the good job.
[0] http://www.evanmiller.org/how-not-to-sort-by-average-rating....
If a person figured out a way to apply it usefully to some other area - which doesn't seem hard at all - (job skill ranking? :D), MS is the kind of company that would attempt to collect $$$ from it regardless. :(
We work with climate science researchers who have multi-TB datasets, and they have no efficient way to share them. Same goes for genomics researchers who routinely pay lots of money for Aspera licenses just to download datasets faster than TCP allows. We are using a Ledbat protocol tuned to give good bandwidth over high latency links, but only scavange available b/w as it is lower priority than TCP.
For the machine learning researcher: i'd like to test this RNN on the reddit comments dataset....3 days later after finding a poor quality torrent...oh, now i can do it. On our system, search, find, click to download. We will move towards downloading (random) samples of very large datasets (even to Kafka from where they can be processed as they are downloaded).
Sure distribution can have issues, but do you have any references for simple possession as training and test data?
But if the data/analysis is published, then the data source would need to be disclosed.
See more on the OKCupid case I mentioned above: http://www.vox.com/platform/amp/2016/5/12/11666116/70000-okc...
When it comes to data quality, the world is a messy place and the data that comes from it is messy too. Most professional data scientists spend an inordinate amount of time cleaning datasets, doing feature engineering etc… it’s part of their job description. We’re trying to eliminate some of that repetitive work by making sure that people can comment on, contribute to and give some signal back on the quality of the dataset. We also think a dataset is more than just the data... on data.world you can upload code, Notebooks, images, etc... anything that helps add context to the data.
Finally, when it comes to size... ML definitely needs it. However, there's a lot of interesting data out there thats still very complicated, very useful but not that big. Most datasets in the world are well under the terabyte size (or even 100s of GBs in size). We're rapidly expanding the size of datasets we support because we want that stuff too but we really want to help people understand all the data in the world!
My general sense on this though is that I'd like there to be more of an incentive for people to open up their datasets to the larger public. Maybe I'm being idealistic but a crowdsourcing type function where you pay for X dataset together with other users and then it's released under MIT, forever free etc.
As others have mentioned that'll probably bump against usage rights issues, a larger problem you'll have to deal with independent of your need to sell or distribute the datasets in question.
You would have to distribute all these datasets with a license and such a thing would restrict any users much more than you'd realize - sharing one's results would become a trickier matter just for one. Deciding if the data is good would be another thing since being able to sell stuff produces an incentive for uploading garbage.
Just as much, huge datasets are effectively going to be the product of many people's data. If some entity decides it's going to profit from that, how do the profits get divided, especially if there's no protocols for that division.
For example, should someone be able to sell my chest X-rays if I haven't agreed to it? A lot of large companies with big data sets would face far more issues if they were to be selling those. And considerations go on and on and on.
The free-wheel, free-sharing quality of present day machine has been a driving force in its progress and it would be a shame for that to go away. As can be seen with the Internet, freely available data can be incredible force and hopefully data will remain freely available.
""" Datapie offers data analysis without downloading the data. This means you need not download the massive data. No need to have massive distributed systems to process it. """
Where is your company trying to fit into the market? Is this a http://zerotoonebook.com/ or are we commodotized in this space already?
IBM offers many of the same data sets, paired with your company's private data, both annotated by Watson. IBM also does not learn from, or improve their own models, w.r.t. your company's proprietary data.
As an example, see [1].
like the idea - minimaxir had some good thoughts.