Does this mean we can now scrape e.g. YouTube videos, Amazon reviews, IMDB reviews, Facebook events ... ?
Does this mean we can now scrape e.g. YouTube videos, Amazon reviews, IMDB reviews, Facebook events ... ?
>hiQ argued that LinkedIn’s technical measures to block web scraping interfere with hiQ’s contracts with its own customers who rely on this data. In legal jargon, this is called” malicious interference with a contract”, which is prohibited by American law
Does this mean that Google's random recaptcha check is interference?
So prices on Amazon.com are facts. User reviews are creative so probably copyrighted.
Similarly the videos on YouTube are copyrighted. However the number of views and the number of likes are probably scrapable.
I wonder if the number of stars are copyrighted. It's not creative, but a fact.
Lets draw some pararells to real life. If I go to public space like town square - can't I take pictures, notes and records then go home and draw my analytics from it? What if I read something in a book I bought, can't I quote it?
Same thing should be with web resources even if they are creative - as long as I don't publish them I should be able to scrape whatever public resources I want and use them in my analytics, machine learning or whatever.
No. At the risk of just repeating the comment you didn't understand, creative works are not "just data" - they are copyrightable works that the owner has control over who can use them, not just for profit, but for any reason with few exceptions.
You don't just get to drop someone else's work product into your algorithm without their permission.
Why not?
I can read a book to post my impression on it somewhere right? I can read it and say "it was beautiful" on twitter.
I can then automate my "taste meter" through machine learning, it reads a given book character by character, and spits out what I'd think of it if I actually read it. Then posts it on twitter, says "it was beautiful".
Did I break copyright law? I don't think so.
Not a lawyer, but:
You can do all of that, but:
You cannot scan the book you bought, and put it on your website for sale or even free - unless it's copyright is up or you are given permission by the copyright holder.
You can not take a picture of someones painting in high detail, then sell prints of it - unless it's copyright is up or you are given permission by the copyright holder.
https://www.rd.com/advice/travel/eiffel-tower-illegal-photos... http://www.photographers-resource.co.uk/photography/Legal/Ac...
https://www.diyphotography.net/10-famous-landmarks-youre-all...
Downloading publicly available data should (by definition of public) not be a violation of someone's rights. However it's easy to see why it wouldn't be desirable for someone to republish creative works as their own, so it's reasonable to give the author control over how their work should be published.
And in the case of price data or similar you would be hard pressed to deem anyone the 'author' of it, hence it would be weird to enforce the author's rights.
Copyright does make _copying_ tortuous. Broad personal use exceptions in USA, for example, make this appear not to be true, but it is the act of copying - even without publication - that is protected in general.
Ripping a CD in UK, for example is copyright infringement without a general personal use exception (there are exceptions, under Fair Dealing, but whatever you're doing almost certainly doesn't fall into them).
See eg UK CDPA1988, Chapter II, section 16(1)(a); or USC17, Chapter 1, 106(1).
Think of Law around data as using dependent types. The legal protections depend on the type of the data, and the type depends on the content (among other things). You have to determine the type BEFORE you can tell what the law says about it, since the law only cares about the type. You could probably encode the law nicely with something like Idris, but any "code as law" type governance system without dependent types won't be able to express existing law.
See the South Park WWITB issue.
I believe South Park used a videoclip from youtube, and Youtube’s ContentID system removed the video South Park had used, because Youtube considered it a violation of South Park’s copyright.