A list of the biggest datasets for machine learning
datasetlist.com
datasetlist.com
> Danbooru2018 is a large-scale anime image database with 3.33m+ images annotated with 99.7m+ tags
At present, the most advanced tagger for Danbooru, DeepDanbooru https://www.reddit.com/r/MachineLearning/comments/akbc11/p_t... , still isn't good enough to do annotation by itself but I think that's mostly because no one has really tried.
> The image boorus are longstanding web databases which host large numbers of images which can be ‘tagged’ or labeled with an arbitrary number of textual descriptions; they were developed for and are most popular among fans of anime, who provide detailed annotations.
i.e. it's pulling from a public website which crowdsources the tagging
Would have been nice to add some sort of discriptor indicating what type of dataset it is. For example, I personally have no clue off the top of my head what the “MURA” dataset is.
Edit: I now see there is a little icon on the left of all lines. Bit ambiguous, but it sorta gets the point across.
Also https://registry.opendata.aws/ from AWS has a lot of datasets that could either be included en masse, or even linked to for the page to be more comprehensive. I like their categorization system (tags/labels) as well. They also have usage examples which is excellent to get a sense of what the data is/useful for.
500K messages from 150 users.
https://aws.amazon.com/datasets/apache-software-foundation-p...
Alot of projects use it!
Especially annoying are european universities and other government funded institutions