Announcing YouTube-8M: A Large and Diverse Labeled Video Dataset for Research
research.googleblog.com
research.googleblog.com
I can post a URL to the output once it's finished running, if it'd be of any use to anyone. Oh, and be warned, there's a strong chance that it's buggy. It's certainly not optimised (no threads).
EDIT: The script has now run. I've scraped ~10,000,000 Video IDs, but only ~5.5m of these IDs are unique, so there's probably a bug in my script somewhere (but I need sleep now). Files containing IDs for various categories are listed here: https://redfern.me/public/yt8m/, some notes are here: https://redfern.me/public/yt8m/README.md, and .tar.gz'd archive is available here: https://redfern.me/public/yt8m/yt8m-ids-probably-incomplete.....
It seems like it'd be interesting to explore their tagging compared to what is in video transcripts.
Would I be violating any law, copyright if I formatted it and put it on my server for that kind of consumption or via JSON?
> The code and dataset are licensed by Google Inc. under license Apache 2.0.
You can actually download each shard (~300 Mb) separately. They haven't yet released the PCA matrix and quantization parameters used with inception model, but should release them soon.
I'd guess the reasoning is, because it's a list of public URLs, there's no expectation of privacy.
I am searching (thrashing) around for my next "big" project. i have been thinking of drones measuring roof / building quality and the CV/ML requirements are fairly high - getting my teeth stuck into these would really give me a better feel for training my own system.
The problem is, how do I feed my family while taking the six months to do it all?