The SOPA Debate and Congress's Understanding of Child Porn
danwin.com
danwin.com
If we want to frame SOPA in the context of these kinds of incidents, one optimistic view would be that, as has happened in the past, science - if we can agree that determining whether or not a copyrighted work is stolen - will prevail.
Congressman: "That's no excuse. I've seen CSI. You can just go 'enhance, enhance, enhance' and find out anything about it."
[Warning: speculation, I do not work for Google...] Google has developed a very effective algorithm for detecting porn, not child porn. They can use flesh tone analysis for detecting nudity but algorithmic analysis of a single image to determine the age of the person in the picture seems way harder. File names don't work either: a picture of a handbag with a file name of "16yr_old_porn.jpg" is not illegal.
Battling child pornography requires a human to look at stuff. Search results get flagged by normal users (and by paid teams I'd bet). These flags are analyzed and used as patterns to detect other similar images (similar colors, sizes, names, patterns in the image, source urls, watermarks) and this secondary analysis can be done with fancy algorithms but it still requires groups of people to flag things and create these patterns.
If I'm right, then a similar method could be used to help combat infringing material. But as danwin points out, there is a new factor: to figure out if a particular piece of content is licensed or infringing, which is hard, even for a human. But there are some people who can look at a video, picture, (listen to a) song, etc, and be pretty sure that it is infringing. Clips or full videos of some popular movies on youtube are obviously infringing. Most mp3s on the web are infringing.
There would be false positives and false negatives (just like with child porn), but you could build a similar type of pattern database that could be used to find similar files. This same type of flagging/reporting could be used to identify infringing sites. I know youtube has an existing algorithm for identifying patterns in audio tracks to some database, but I don't think these patterns are based on flags, I think they are hard-coded.
Users could still browse directly to a URL (or IP) hosting questionable content but it would definitely cut down on the easy access and spread of files.
NOTE: I don't like the idea of Google (or any organization) being mandated to implement such a process but it may be better than interfering with DNS mechanisms and lawsuits and court orders. I also don't like how reliant we are on a private organization holding such power over what the world can find in their results. Nearly everyone agrees on child porn but I think one of the main reasons SOPA is controversial is because we don't all agree on the importance of fighting infringement. (Counterfeit goods maybe, but not license violations)
By no means do I claim that SOPA is a good bill. I think doing nothing is better than implementing what is in SOPA.
EDIT: An assumption made here is that we consider the way Google is fighting child pornography is effective, sufficient, and ethical. I think that is a reasonable assumption.
EDIT2: One factor I didn't consider is the quantity and diversity of child pornography versus infringing material. I'm not sure how the algorithm would be affected but it'd probably be slightly less effective.
Or you know, fair use.
It was auto-flagged by their system and I had to jump through hoops to get it back online. I can't imagine this sort of thing would survive SOPA. And what other platform is there but the Internet for the sharing of multimedia commentary and criticism?
There's no sane reason we should lose the visibility, distribution and community relevance of social platforms simply in exchange for exercising fair use of something. That said, the point was less about Youtube being unfair and more about how these platforms are already heavily biased against the assumption of fair use.
And its better than letting Congress enact SOPA, which is my whole my point upward in this thread. I know that's not saying much, but SOPA supporters claim that we can't (or won't) offer any alternatives to SOPA. Since Congress is always in the mindset of "we have to do something", let's give them an alternative that they can support that shows "something is being done".
Fair use may be defined in specific terms (rather debatable... "most interpretations" indeed.), but determining if a particular use is fair use is not a science.
If you are embedding the video in the middle of a document, its clearly not on its own and has been added to. However, if you embed 10 minutes of a video in the middle of your document, now that might be too much to be considered fair use.
I think your point is that a weaknesses of the algorithm I proposed is that media considered "fair use" could be flagged as infringing. My response to that is that Google already has procedures for dealing with images/videos that were questionably reported as offensive.
Quotes like this make your assessment a bit bumpy - do you have any citation to back this up? From what I can see, there is tons and tons of mp3 material on the Internet that is perfectly legal. While infringing material is mostly either trafficked via p2p or server-based filesharing systems. Maybe I'm just understanding "the web" to be more limited than "The Internet", though.
I know there are some mp3 sites like beemp3 [https://www.google.com/search?q=related:beemp3.com]. I suppose I was thinking about the ease of a human to judge that, yes, beemp3 is nearly entirely made of infringing music. That humans can report these sites and algorithms can find similar sites (and similar files) and develop a pretty decent database of items they could omit from search results. My point is only that it does seem technically feasible.
For those who haven't taken journalism law classes, a newspaper can print falsehoods without fear of being sued, especially about public figures. The threshold is that they cannot do it with "actual malice" or "gross negligence."...this is not the case in many other countries, including Britain.
Why not make it a suable offense to print any kind of falsehood, which are most definitely detectable by human (and probably machine) readers? Because if a newspaper -- or any media source -- had to worry about being sued because of simple mistakes of fact or, more problematically, because a whistleblower misled them -- then most newspaper businesses would err so far on the side of caution that they would not print anything useful about any public figure.
[insert ridicule for how the corporate-controlled media doesn't do that anyway]
To go back to the SOPA case...a company like Google has to deal with something much more difficult to ascertain than errors of fact, and so it worries that it will have to adopt a "it's better to apologize than to have to ask" policy, which would be detrimental to its search property.
Meanwhile, in the physical world, a knock-off handbag can be easily identified as such by airport personnel. This is the world these pro-PIPA politicians live in. Child porn, online as offline, is inherent in the work. You don't need supplementary information to decide if something is child porn. Online, the difference between a copy and the original is in some people's minds, not anchored in reality.
Even if this were true, just because Google has the capability to do this doesn't mean every company does. If it costs Google $X million to prevent online piracy on their sites, is it reasonable to expect every web startup to do the same?
http://www.pittsburghlive.com/x/pittsburghtrib/opinion/colum...
---- "I can't be responsible for every undercapitalized entrepreneur in America," Mrs. Clinton said in 1993, responding to charges that her plan would bankrupt businesses and cut employment.
AFAIK, this is complete BS! I was quite up to date with research in this field till about 5yrs ago and the systems are still in their infancy. If you put an automated system to automatically detect pornography, the false alarm rates would be very high. And differentiating child porn from regular porn is also extremely hard.
Companies providing the actual blocking services don't want to even see the URLs on the list (going so far as to ensure that it is encrypted at all times until the URLs can be safely hashed). They certainly don't want to download the contents and have them sitting (even temporarily) on a server somewhere.
As for telling which is child porn and which isn't...given the ways that people usually distribute this material (not through public webpages), it's doubtful that Google really ahs to do that much proactive monitoring of it to keep it out of the usual Google-pointed-to places
I used to work for a company that did high-demand image hosting. Filtering porn was a big deal at the time. It was considered to be of such low accuracy as to be useless. I can't imagine how one comes to the conclusion that a sexually exploited child can be located in an image by a machine.
Please don't do that, but how hard would it be?
It's about as effective as our existing regulatory systems. For instance, a lot of commerce is taxed and regulated, but you could just walk up to somebody and make exchange. Never air-tight, but it is enforcement.
Is my suggestion wise? I definitely don't think so, but it is feasible. And maybe I'm not considering something correctly -- but, please, don't be rude about my technical prowess. I'm just trying to strengthen our argument.
Without that, there is no way to tell if a link is really legal or not. And that is what Google was saying is impossible to handle.
Doubtful. Scary, but doubtful.
Let's say we went from "everything has copyright" to "you must register with the copyright office to get a copyright". Now it does have a lot of drawbacks but for most things like this post on HN I could care less that it's copied elsewhere yet currently there is a copyright on this post.
Now in a system like that what would prevent the copyright office to make publicly available a database of hashes for copyrighted content like videos,music,text that would allow people from Google and elsewhere to make a resonable assumption about licensing? This DB could contain a list of urls that are allowed to host the said content or offer more open terms.
In a way google already started such a DB for youtube but it's prone to abuse and contesting copyright on there isn't exactly simple to do.(as seen from the megaupload thing)
What about "derived work"? It's relatively easy to make a "copy" that displays/plays the same but has a completely different hash. In fact, it's very difficult to get the same hash on a excerpt.
And then there's fair-use.
And then there's the matter of who has to bear the cost of creating and placing the filters, sites outside our jurisdiction that don't bother with those expensive filters, etc. You end up hobbling the US tech industry for no reason, putting everyone at a competitive disadvantage.
Right now, there's a lot less incentive to do all of that. But change to a system like this and people will just start trading in stolen accounts, much like they already trade in stolen credit cards.
Which is exactly the point of SOPA...since we can't stop the flow of copyrighted media on the internet, we can't have the internet.
Now suppose we have a magic system that's better than expensive copyright lawyers that can manage 99% accuracy (which they didn't) and that we operate at internet scale on the roughly 60 billion web pages Google indexes. Then assume that no less than 20% of the internet is pirated material. Does that sound too high? Good. Your numbers will look even worse if it's lower. Now we do a little math:
0.2 * 0.99 = 19.8% are pirate pages correctly blocked
0.8 * 0.99 = 79.2% are innocent pages correctly allowed
0.2 * 0.01 = 0.2% are pirate pages mistakenly allowed
0.8 * 0.01 = 0.8% are innocent pages mistakenly blocked
So we've just censored four innocent pages for every one containing pirated material and there would be 120,000,000 pirate pages (0.2% of 60 billion) out there that aren't blocked. Feel free to work out the math, but if you want to say that 20% is too high a proportion, you'll have even more false positives, so even more innocent pages are censored for every pirate blocked.And this assumes that we have something better than ~$300/hr copyright lawyers screening all 60 billion web pages. Any computer program we make won't even be this good.
No one is going to mistake a four-year-old for a grandmother.
True. But plenty of people will mistake a 14, 15, 16, or 17-year old for an 18-year old.
The Viacom v. YouTube case proved that even expensive lawyers spending hours doing research can't manage to get even 99% accuracy.
Again true, but the odds of a randomly selected uploaded full-length movie being illegally uploaded is, I bet, significantly larger than the odds of a randomly selected explicit movie featuring people of illegal age. So the prior coming from what fraction of available content out there is illegal according to the two standards would likely lead to a larger false positive rate for pornography.
We can verify authenticity. We can transfer authenticity. I think about BitCoin and the security of the network. Everyone on the network can verify the origin and authenticity of a transaction.
Could something similar be put in place for media? A distributed DRM essentially? While DRM in its current form is a nightmare, isn't the main drawback being tied down to certain platforms? If DRM were decentralized, I don't think I would have a problem with it.
Thoughts?
Your concept seems ridden with so many problems and possibilities or abuse that I'll let this as an exercise to the reader... (example: how do you "tag" existing files? existing CDs? whose interest would this be? would people bother? etc.)
To get people to do it, it would be a part of uploading media files. If SOPA were to pass, I imagine hosting sites would require only "signed" media. It wouldn't be perfect, but any distributed file would be signed at creation. Self signing would be simple enough if it were built into media programs. The key to all this is that nobody controls it, and the standard is open source. Once again looking at BitCoin as a role model.
Also, I don't think it's fair to tag a concept "ridden with problems and possibilities for abuse". It's a concept, not an implementation. And it was a mighty vague one at that.
This isn't as hard as the author is making it sound and I suspect Youtube already handles it with movie clips. With the aid of an external "copyright registry", Google could see if the fingerprint of an unknown song is close to one in the registry. Obviously, there would be many false positives and true negatives, but that is certainly true with child porn detection as well. Indeed, it wouldn't take all that much coordination to pull off; a better-funded copyright office could handle this.
I don't agree with Google/internet being forced to do this due to the burden it creates, but it isn't that hard to pull off...
only because its not true?>
My opinion - pushing for such a legislation would make it a criminal act to err on the side of false negatives. Which means false positives will be very common.
This point even a monkey could understand. If it refuses, then it is wishing for bananas or something.
Technically untrue. Computers are able to correctly classify (and with good accuracy/precision) lots of things that even domain-expert humans have a tough time classifying.
One problem with those systems is that, depending on how they're trained up, humans aren't even able to pick apart how the software is making its decisions (and use that to advance the state of the relevant science, for example).
Creepy and awesome.
A 7 year old that can read could be trained to properly recognize spam, child porn, (allegedly) copyright infringement and also give you advice on purchases after seeing your spending habits. And a 7 year old may not have the bandwidth of a supercomputer to process hundreds of thousands of items at once, but he is able of greater accuracy and that's because the human brain is the most advanced pattern-matching processor in existence.
You can classify anything by means of statistics, sometimes with surprising results, however my point (and maybe I wasn't making myself clear) is that you'll get a lot of errors of judgment. Which is why algorithms will be trained to err on the side of false positives, because doing otherwise will put the business in jeopardy.
I was hoping the Wikipedia page on 'expert systems' would bail me out the examples front, but its 'disadvantages' section isn't all that clear or complete. It does touch on the issue I mentioned, though.
I think many high-frequency stock trading algorithms are examples. The trades may as well be magical, and as long as the program makes money the owner doesn't much care why.
How would you like it if the cops busted down your front door because you accidentally uploaded to Facebook a picture of your new baby having a bath? What if the pageant photo of your daughter was automatically classified as sufficiently pornographic? Handcuffs for everyone.
Once computers are involved, people blindly trust them, and from there there's nothing but trouble.
Google's smart. How could it ever get anything wrong?
a) it is possible to build an expert system to classify online content into 'infringing' vs. 'non-infringing'
b) were it possible to build such a system, attaching the deployment of such a system to a piece of federal legislation would be a legitimate use of government power
However, as kpanghmc pointed out, even if Google had the capability, that doesn't mean every company does. I'll take that a step further. Even if Google had the capability, the government doesn't have the right to step in and demand that a single private entity build something to stop it, even using government (taxpayer) money. This is how you know the government has gotten too big.
Furthermore, the article while trying to empathize, omits the fact that this is censorship. The banning of web pages or websites as a whole without undue process is not constitutional and un-American.
In other news, the police start withdrawing money from every bank account using an ATM and analyzing the $100 bills for signs of the account holders' criminal transactions.
If they wish to block child abuse, they basically admitted to wishing to kill the future of children just to hide evidence of abuse.
The ball's in their court.
</sarcasm>
Speaking of, I wonder how many children there are in the world.
Wow that is infuriating; ignorance masquerading as arrogance.
Sadly, even though it was carefully explained to him, data on whether or not something constitutes copyright infringement does not exist. That's not a technical problem. It's not something you can change with a magical computer program. Infringement rests upon whether or not the person has permission to do whatever it is they're doing. That data is simply not available, so there aren't any numbers for Google (or anyone else) to crunch. He can only say it's easy because he doesn't know anything.
If he worked in sales, I bet he'd promise customers a time machine, collect a huge bonus, then blame the geeks for failing to deliver.
Rep. Marino: "wakes up from daydream you nerds are smart, you can figure it out"
/rage...