Discovering Millions of Datasets on the Web
blog.google
blog.google
Some other indices of open data sets I've found:
https://en.m.wikipedia.org/wiki/List_of_datasets_for_machine...
https://austingwalters.com/parsing-uspto-patents-to-create-a...
I actually have found FOIA requests[1] and downloads from government websites to be the easiest & most effective way to get robust datasets.
[1] https://austingwalters.com/foia-requesting-100-universities/
+ For individual requests, come over to https://opendata.stackexchange.com/ and ask!
+ Wikidata has loads of structured data, but using SPARQL is often a barrier. But you can request help: https://www.wikidata.org/wiki/Wikidata:Request_a_query
And data is simple, it's parameters plus timestamps plus a lot of storage.
Realtime access is harder but it's a well-specified problem.
The issue is inference. No one does inference extremely well except in limited circumstances. It's one of our greatest bottlenecks as humans and our software is going to be limited by it as well insofar as our understanding of what to build is controlled by what kind of inference we want.
There are a very limited number of companies that have complete datasets of buildings in all cities throughout the country. The leader is CoStar, they have the most complete and accurate data. What I have noticed with other CRE data companies is they are focusing on leasing activity or 1 other part of the CRE process. CoStar became the leader by valuing their data over all things, if other companies want to compete they need to do the same.
[0] https://youtu.be/M9pFTApeo_8
Additional link to his data site: http://people.stern.nyu.edu/adamodar/New_Home_Page/data.html
https://www.sec.gov/edgar/searchedgar/accessing-edgar-data.h...
[1]https://datasetsearch.research.google.com/search?query=EDGAR...
Unfortunately, most every result for the word `philosophy' is borderline garbage imho. Keyword indexing of datasets may need improving?
Overall a bit underwhelming.