AWS Public Datasets
aws.amazon.com
aws.amazon.com
I think it would show a kind of gilded age maturity if all the cloud providers cooperated on their public datasets, because they are for the public good.
(Work at g)
Shame on me to think "Public Datasets" meant that I could just download them via rsync/http.
If you're interested in truly public datasets, see http://academictorrents.com/ which hosts datasets as torrents.
My criticism is not towards them not providing free transfers of data. My criticism is that they say "these datasets are freely hosted and accessible" and public, while only being available while logged in to Google, which I don't think classifies as public.
I would not have any problem if it's just called "Google-hosted Datasets" or "Google-only Public Datasets", I just think the current naming is misleading.
Edit: compare this to the AWS Public Datasets which are actually available without a AWS account, just go to https://landsat-pds.s3.amazonaws.com/c1/L8/139/045/LC08_L1TP... for example, which is the Landsat Dataset
the term public means its available to the public. Signing up google, library card, doesn't negate that fact that any member of the public can access the data with a reasonable and insignificant barrier.
And if that's your definition of public, isn't everything public, if you have the right amount of money, knows the right people and can get access to the right place?
Google give you 1TB on data to query each month. That is a lot of data, if you need more for free collect the data over time.
But you still need a library card and if the facility gets overused they ask for a small tax to cover the operating costs.
I fail to see how this isn't considered public. Public literally just means accessible to the general public. That all. A public event can still charge a fee for entry. As long as the fee doesn't create a barrier for most the public, you know accessible to the general public.
Also if a library sold their data would it no longer be considered public? I believe it would still be considered public. User data isn't sold typically by name or email, byt activity. If a library compiled a list of books checked out and their frequency, the amount of people entering everyday etc that be the equivalent of most user data being sold. Very rare for a company to sell your actual personal data, when they do they disassociate your personal information with it.
So if the above doesn't disqualify a library from being public, then neither would this dataset that is public. If you really disagree with that then you are just trying to be pedantic at that point.
Website: http://www.refine.bio/
I thought you meant derived data.
(I work at Kaggle)
`kik pull titanic`
I helped build the Terrain Tiles dataset as part of Mapzen, which recently shut down. The OpenStreetMap data exists on the AWS Public Datasets page because it's useful to Humanitarian OpenStreetMap Team. If you're able to convince your company to generate and work with a public dataset, consider reaching out to the AWS and Google public datasets teams to get it hosted and publicized.
The descriptions note “Educators, researchers and students can apply for free promotional credits to take advantage of Public Datasets on AWS.” which is not a good sign.
However, the idea is not that you download it all (there's probably cheaper ways to acquire Landsat data), the idea is if you want to do whatever analysis on AWS, they've already got it neatly ingested for you.
They finally solved this by using meeting notes from UN Assembly. Which were transcribed by best of the translators? that access to meeting transcription was (unfair ??) advantage Google had over other tools. Was it wrong ? I don’t think so. Should have been those meeting notes be public: Yes
I tried this sentence: "Hei äiti, puhun suomea". The expected translation would be "Hi mom, I speak Finnish".
Instead Google's result was: "July's mother, I speak English".
Obviously the engine had been trained on unvetted data sets where the word "English" occurred in translations in a position where the original had the word "Finnish", and no context was provided to avoid this kind of mistake.
The word "July" came about because "hei" is also used as an abbreviation for "heinäkuu" (July). It was sobering that a supposed world-class AI couldn't distinguish between these two usages. Machine learning needs a lot of old-fashioned handtuned human-made heuristics.
Cortana when??
[1] http://fcon_1000.projects.nitrc.org/ [2] https://openfmri.org/