11,865 karma · joined March 13, 2008
Research: https://www.cs.princeton.edu/~arvindn/
http://twitter.com/random_walker/status/11917456790
"within a few years there will be a real-time stream that aggregates essentially all activity anywhere on the web"
While Twitter, Wordpress, etc. have firehoses, this announcement is different -- aggregation across multiple services. I think the concept is very powerful and is something I've been looking into over the past few weeks.
I'm curious how many people know about Spinn3r (spinn3r.com) -- they claim to provide a real-time hose of all updates in the blogosphere. They quoted me $6k/mo. Unfortunately, Spinn3r's main product was deadpooled last year, so I'm not sure what their future is as a company. I think they should have made the real-time stream their main product. In a way they were too early.
Another startup called Superfeedr takes a different approach. They charge on a per-feed (or per-entry) basis, so their up-front costs aren't as huge. They currently cater to customers who might only want a few hundred or a few thousand specific feeds, but as they scale up they should be able to offer a bulk stream of essentially the entire blogosphere.
Then there's Google Buzz, which is also about aggregating different services. They don't currently offer a firehose but apparently you can build one yourself using PubSubHubbub, although developers report that it isn't working so great: http://groups.google.com/group/google-buzz-api/browse_thread...
But the long-term trend is clear. We're going to have an "uberhose" at some point! I have a few applications in mind of things that such a pipe would enable that are fundamentally not possible today, but I'm curious what ideas other people have.
Of course, my sample size is rather small; I'd love to see data supporting or rejecting my hunch.
That FAQ question is a big red herring. Some of the objections to our paper during peer-review were along the lines of "if Netflix had esentially duplicated every record in the database, how could you be sure you found the right record? What does right record even mean?" No really, it was that silly. So it was meant to be a way to justify the fact that de-anonymization can happen even if you didn't find "the right record."
* For the longest time we mostly stuck to doing the math. We certainly didn't call for the contest to be cancelled, and we had nothing to do with the lawsuit. But when some people implied that we were responsible for the mess that ensued, we were kind of pulled into it. We posted this as a way of explaining our point of view and reaching out to see if there's a possibility of collaboration.
* The sadness we expressed is genuine. The reason we brought up Netflix's response to our paper wasn't "snark" or "gloating." Rather, we were pointing out that the cancellation of this contest was rather needless, because if they had acknowledged the privacy risks back when we published the paper, they would have had more than enough time to deploy an opt-in system for this contest. I think it is really unfortunate that that didn't happen.
* Someone wanted to know exactly what I thought of the "greater good" argument. Well, I'll tell ya. I'm vehemently opposed to it and I think it's a dangerous slippery slope. I think this point of view is enshrined in the ethos of this country -- "better let ten guilty men walk free than to convict one innocent man," etc. I don't think anyone has the moral authority to decide that the privacy concerns of a few can be sacrificed.
* There is a specific reason we chose the open letter format rather than communicating with Netflix directly. Acutally, two reasons. First, there are many data privacy researchers who are at least as qualified as we are for this role. We wanted to make sure the community had the opportunity to participate in whatever ensures, rather than just us.
Second, I'm sure there are many companies other than Netflix who have a similar need for privacy preserving data mining. If Netflix doesn't take us up, perhaps one of the others will. Bottom line, since there are multiple parties on both sides, and we don't really know who they are, we felt it is better to have this dialog in public.
* Finally, it is understandable when something like this happens to want to find someone to blame. But think twice before shooting the messenger.
As for that part of the FAQ, it is intended as an explanation of some of the theorems proved in the paper and is a response to some of the theoretical objections we face from the data privacy community. It is not an issue that arises in practice.
Also, some thoughts on a closely related topic: http://news.ycombinator.com/item?id=1114578
tptacek commented that it was a "batshit crazy idea," and that is exactly right. This article is an example of how to (really) abuse history stealing. As promised, a stronger variant I've been working on is coming soon.
I actually spent 3 years doing research on the other side of the paywall. It is one of the things I'm "vocally critical" of. Here's me calling it evil just a couple of days ago: http://bit-player.org/2010/yet-another-spam-update#comment-2...
At any rate, there's not a single scientist who believes that publisher copyright for papers is a good thing. Unfortunately we can't change the system overnight. At least in CS, as I pointed out in that comment, it has already become a non-issue.
Science is actually admirably egalitarian. That's not one of the things I think is wrong with it. The heavily politicized branches might not be egalitarian; I have no experience with them. But in general, you don't need to be a member of any professional organizations or have any other credentials to submit your work to journals and conferences (which is how we transact our business.) Many journals employ double-blind reviewing, in which case even unintentional discrimination based on credentials is unlikely or impossible.
Nevertheless, scientists might often give the impression of not wanting to let amateurs into the 'club'. Why is this? It is simply an issue of bandwidth. Think about this: for each amateur, like the author, who had something to contribute, how many do you think thought they had something to contribute but were mistaken? If you guessed a hundred thousand you'd be in the right ballpark. I've seen it in many different areas.
For example, the number of people trying to submit "proofs" of P != NP, (or worse, P = NP) is just ridiculous. Some are well-intentioned although ignorant, and others are just cranks. Some fields of research attract more amateur claims than others, but whichever field you look at, the amateurs greatly outnumber scientists, and the vast majority of them are mistaken.
So what do we do? We use a simple filter. If you speak our language, we'll listen to you. This is what appears to the lay public as anti-amateur. But it's not, really. All you have to do is to learn some simple definitions and terminology set forth in a straightforward way in previous papers, to prove that you've done your homework. Then write up what you have to say and we'll be happy to give it a read. Sure, like every heuristic, it's not perfect; sometimes there are false negatives. But there really isn't an alternative. Without a filter, all we'd ever be doing is debunking crackpot theories.
I acknowledge that much of this probably doesn't apply to climate change "science," which is a special case. But there's been a lot of criticism of science in general, and I wanted to set the record straight. It's important to distinguish between what's really broken and what appears broken, otherwise we risk throwing the baby out with the bathwater.
"Elitism is the belief or attitude that those individuals who are considered members of the elite — a select group of people with outstanding personal abilities, intellect, wealth, specialized training or experience, or other distinctive attributes — are those whose views on a matter are to be taken the most seriously or carry the most weight.."
Having specialist and extraordinary designers instead of letting programmers or "the crowd" design the product has arguably been the key to Apple's success.
Well, somebody had to. I could get it to nest 10-fold, but after that the innermost browser window was unresponsive.
On the positive side, your design/stylesheet is much better than HN, IMO. (For example, on HN I keep voting up/down when I mean the opposite because the damn arrows are too close to each other.) Having 1-2 line summaries of articles is also very useful.
This issue has been popping up in different countries all over the world in the last few months. Most notably in China, of course, but also in India, France and Italy, where Google executives are in fact facing the possibility of jail time. See http://www.guardian.co.uk/commentisfree/libertycentral/2010/... and http://www.google.com/hostednews/afp/article/ALeqM5h-40ArK-e....
The role of Internet middlemen in enforcing copyright was a key legal issue in the last decade. It appears that censorship will be the analogous issue for the next decade. We seem to have reached a relatively happy middle ground with copyright—middlemen have some responsibility, but a strong form of copyright protection has proved unenforceable, forcing many industries to innovate or die. We can hope that a similar thing will happen with censorship, with oppressive governments either collapsing or being forced to allow free speech.
Wikipedia has the basics of setting up a new short code: "Common short codes in the U.S. are administered by NeuStar, under a deal with Common Short Code Administration - CTIA. Short codes can be leased at the rate of $1000 a month for a selected code or $500 for a random code." (from http://en.wikipedia.org/wiki/Short_code)
More info here: http://www.mmaglobal.com/shortcodeprimer.pdf
Ubiquitous surveillance is inevitable. The question is whether it is going to be whether or not the citizenry will have the same power that the authorities do. Technology is not the barrier; it is a matter of fighting for our legal rights.
The contrast is stark. One path is Orwellian, whereas the other path, while eroding some comforts that we're used to, promises us benefits that will more than compensate for it.
Primepages (http://primes.utm.edu/) lists the largest known primes of various forms. One of the categories is primes that have no special form, which are the hardest to prove prime.
Chris Caldwell, the primepages maintainer, called up Phil Carmody and said, "Phil, if you can turn DeCSS into a 'general prime' that's big enough for the top 20 ever discovered, I'll be happy to host it on primepages." (Just as he would host any prime, DeCSS or not, that was big enough for the top 20.) So Phil did, cuz he's a smart guy and had a lot of computrons, and the prime went in the database.
The point here is that if you're merely turning data into a number, you have no legitimate reason to distribute it other than whatever reason you had to distribute the data itself. However, primepages existed long before DeCSS became a problem, and was arguably simply doing what it always did, which is to list the top tens.
I don't know if that would hold up in court -- it was never tested -- but the argument is a lot more subtle than what people are criticizing here.
An aside: Phil and I worked together on some other primality records. I wrote most of that Wikipedia article back in 2002 or so, and they went and featured it.. this was all before the rules for citations/references on Wikipedia became tighter than a mouse's arsehole. So if anyone is in the mood to add some citations, that would be much appreciated.
Why am I mentioning this? Because of the "Myth of the Superuser," the tendency of the media to exaggerate and lawmakers to overreact to the power of pranksters online. I refer to this paper by Paul Ohm: http://papers.ssrn.com/sol3/papers.cfm?abstract_id=967372
4chan in particular tends to be sensationalized so often, so it's good to keep the right perspective.