Open Source Datasets
deepmind.com
deepmind.com
I don't think that CC licenses are a good fit for data collections in general. Let's see what other open data projects do:
Some open data projects were not satisfied with CC licenses, which is why ODbL was created:
https://opendatacommons.org/licenses/odbl/
Also, many open data projects choose to put their stuff into the public domain (or license it as "CC0" which means exactly the same). This is, for example, what Wikidata does:
https://www.wikidata.org/wiki/Wikidata:Main_Page
(Wikidata the structured sister project of Wikipedia.)
Just because something is open source doesn't mean it can't have an academia-only restriction. Data sets should cost money in for-profit uses, open or not.
"The license must not restrict anyone from making use of the program in a specific field of endeavor. For example, it may not restrict the program from being used in a business, or from being used for genetic research." https://opensource.org/osd
Outside of copyleft, Creative Commons 4.0 licenses (CC-BY 4.0, CC-BY-SA 4.0, CC0) from 2013 are good licenses for open data: https://theodi.org/blog/cc-40-and-open-data
I actually use their css when formatting my .md files
Most of the consumers so far have been neuroscience researchers and statisticians, but we do hope (and think) that there's value for a wide variety of interests.
There's a bunch of different data, but the highlights are fMRI scans of people watching and/or listening to the movie Forrest Gump, eye tracking, and detailed annotations of the movie. We are also about to begin acquiring simultaneous EEG and fMRI.
http://studyforrest.org/data.html
Accessing the data is easy, and, as great admirers of Joey Hess, we also have it available in a git annex repo. :-)
http://studyforrest.org/access.html
---Alex
[EDIT] Given that this thread is about open source datasets, it's probably worth mentioning that the license is PDDL.
Unfortunately, Open Source does not help here -- I do not see how OS can be used with data sets. The main OS leverage with software development is that if you use software X to build software Y, X is usually present in some way, shape or form in your deliverable Y. Not so with training data -- once algorithm development is done you can (and usually do) strip training data out and have a finished product that does not require X to run.
Even if one were to require open sourcing derived datasets it is usually easy to segregate the dataset with a tainted (open source) license as you build up your data so the new datasets are not formally "derived" and thus would not need open sourcing.
I would love a better way forward on this, or at least a cleaner explanation of options.
The benefits of OS data are the same as the benefits of OS software. The distinction between "Free" and "Open" is the same as well.
Edit 1: OS data sets are nothing new. The UCI Machine Learning Repository[1] has been around for years. There is also an entire Open Data Stack Exchange site [2], and an Open Data Subreddit [3].
Edit 2: OS data sets are essential for developing new algorithms because they can be used as benchmarks. Nobody should trust a model that's been developed on a proprietary data set for use on anything other than that one data set.
[1]: https://archive.ics.uci.edu/ml/
However, there is a whole bestiary of open source licenses that span the spectrum of "use any way you want" to much more restrictive. But they were mostly thought through for software and data is different; what may prevent proprietary abuse in software may not have any teeth for data.