List of high-quality open datasets in public domains
github.com
github.com
I couldn't find a usecase for these lists myself yet. There is no way to verify the quality of the product or the activity (stars for example? last commit date?).
In one case I searched for aws adapters for a language, clicked on all links inside awesome-{{language}} just to find that all of them are inactive or a few days young. I ended up using something I found on google instead.
GitHub offers a reasonable way to manage contributions to them, compared to many other solutions. It is easy to suggest/fix something for external contributors, but the owner can work as a gatekeeper. This is something many link aggregators or bookmarking sites lack.
One example that isn't perfect, but I found interesting was https://github.com/Kickball/awesome-selfhosted It has a sentence about each project, license information, and tries to purge unmaintained projects.
Also, quality is more complex than just stars or activity. You need to know what the advantages or limitations are and how that applies to your project.
This list would be much improved with descriptions for each dataset and indication of schema, as some of the datasets listed have very unfriendly schema. (e.g. the IMDB interfaces link)
Kaggle's recently-released Public Datasets feature (https://www.kaggle.com/datasets) provides an interesting approach to presenting data and qualifying datasets by giving good examples of data robustness.
How is it this community can debate some inconsequential nonsense, and there is no discussion here of how we get a consistent set of meta-data for these data sources.
There are researchers both in academia and in the commercial world who would thrive if there were such a list with good consistent meta-data on how to interact with it.
Disclosure: I work regularly with open datasets, and the effort it takes to work with each different set overshadows any effort on actual analysis.
I expect this is because there is no ability on HN to differentiate (on articles and comments) between bookmark and upvote, most likely the majority of votes are for the purposes of bookmarking. Very often I want to upvote someone for a good comment, but I do that very sparingly now because I try to keep my upvoted comments list minimal so when I try to find something noteworthy I don't have to wade through pages of "liked" comments.
Your saved comments: https://news.ycombinator.com/saved?id=vosper&comments=t
My saved stories: https://news.ycombinator.com/saved?id=cbd1984
Your saved stories: https://news.ycombinator.com/saved?id=vosper
I have no idea if you'll be able to see mine or if I'll be able to see yours. We'll just have to find out together.
Edited To Add: No. I can't, and you won't.
The datasets are not public domain, but licensed under the Open Government Licence (which allows you to use and adapt the data for commercial use).
There's also the Global Open Data Index: a website that ranks countries by how much Government data is available as open datasets based on certain criteria. The current top spot is taken by Taiwan
1. Taiwan
2. UK
3. Denmark
4. Colombia
5. Finland
5. Australia
7. Uruguay
8. USA
8. Netherlands
10. Norway
10. France
http://index.okfn.org/place/I don't quite understand the criteria for being included in the list since I think it's:
https://groups.google.com/forum/#!forum/awesomepublicdataset...
Also nycopendata.socrata.com
Perhaps a distributed registration system ala DNS?
Quandl (https://www.quandl.com/browse) is similar to Engima, except they got rid of all the fun datasets and added more finance/economic datasets. Hmrph.
the remaining 50% of datasets have a lot of gems
quality in the breadth of data is important
state-wide liquor license, corp reg. OSHA is great data. AMS shipping records. FDA adverse events. Oil and Gas well locations/production. Consolidated weather reports since 1800... im almost certainly forgetting some.