Sending 1.2M Tweets
shkspr.mobi
shkspr.mobi
If each of the 1.2 million tweets includes a ~150 KB image, that’s 180 GB of images hosted on Twitter for free.
It looks like there might be about 250mb available for all to share across the internet using this system
PingFS is even more out there!
The internet is full of weird corners to exploit in fun ways
[1] https://help.twitter.com/en/managing-your-account/how-to-dow...
Should there be (or does there exist) a type of license for data—different from the ones typically used for software source code (MIT, GPL) and ones typically used for creative work (CC), encouraging innovation but giving something back to dataset creator or maintainer?
For my personal stuff, if you'd like a different license, I'm happy for you to pay me for a more restrictive one. But if you build an ML using my open data, I expect that model to be released under a similarly licence.
To (partially) answer myself, contrary to what I implied CC-BY does cover this base if (for example) the creator of the dataset accepts a note in product’s “About” documentation as sufficient attribution.
Another issue with this dataset is the overlay changing over time in text content, font, and colour. The algorithm might overfit and think e.g. yellow font presence means higher output simply because the output was higher during that period. You could strip away the text, but then you're introducing potential errors into the dataset yourself.