HNHacker News
TopNewBestAskShowJobs

randomwalker

11,865 karma · joined March 13, 2008

Princeton prof: https://twitter.com/random_walker

Research: https://www.cs.princeton.edu/~arvindn/

submissionscomments
randomwalker··on No silver bullet: De-identification still doesn’t work
Differential privacy is a different way of releasing data that avoids the problems of re-identification. The theory is well developed, the tools are starting to get there, but the hardest part is that it requires a behavior change from data analysts -- you need to formulate the desired computation algorithmically instead of just poking around the data. This has proved to be a formidable barrier.
randomwalker··on No silver bullet: De-identification still doesn’t work
I'm the first author of this piece, happy to answer any questions. Much of my previous research on re-identification (http://33bits.org/about/) has been discussed on Hacker News.
randomwalker··on NIST Randomness Beacon
The other major one is reliability (but that is encompassed by sufficiently broad interpretations of "trusting NIST.") Also, for our specific application it's not clear how to assign precise timestamps to Bitcoin blocks that everyone can agree on; this is necessary to be able to use a beacon that is external to the block chain.
randomwalker··on NIST Randomness Beacon
I'm one of the authors. It's interesting that you cite our paper as an application of the NIST beacon. Perhaps a better reason to cite our paper is that we actually design a beacon using the Bitcoin blockchain itself (Section 4.4), because the NIST beacon has many problems.
randomwalker··on Reddit’s empire is founded on a flawed algorithm
tl;dr: Posts whose net score ever becomes negative essentially vanish permanently due to a quirk in the algorithm. So an attacker can disappear posts he doesn't like by constantly watching the "New" page and downvoting them as soon as they appear.
randomwalker··on Software Engineering Code of Ethics
I'm a CS researcher and educator (details in profile). Software engineering ethics has been a passion of mine lately. I teamed up with a philosophy prof to write an essay on why it's important to teach ethics in CS classes [1]. (It's an invited column for CACM.)

My coauthor has released a self-contained module with some theory and various hypotheticals that educators can use in classes [2].

I've been trying to crowdsource a set of real-world case studies with a broad coverage of various types of ethical issues [3]. I'm also gradually trying to incorporate this in my own teaching.

We'd appreciate feedback and suggestions.

[1] https://dl.dropboxusercontent.com/u/131764/web/sw_engg_ethic...

[2] http://www.scu.edu/ethics/practicing/focusareas/technology/s...

[3] https://freedom-to-tinker.com/blog/randomwalker/ethical-dile...

randomwalker··on RSA-210 factored
To clarify, this is not a new factoring record for products of two primes. RSA-768, a 232-digit number (22 digits longer) was factored in 2009, and that record still stands. http://en.wikipedia.org/wiki/RSA_numbers#RSA-768

The algorithm used here is GNFS (general number field sieve), which is the same algorithm that's been used for about two decades. In other words, this has no impact on the security of RSA.

More information: http://en.wikipedia.org/wiki/Integer_factorization_records

randomwalker··on Yale Computer Science Dept overworked, understaffed
CS enrollment numbers are way up in every school. Yale is complaining about 200... at Princeton our intro course has over 500. At Stanford almost every undergrad takes a CS course at some point.

We have a good number of lecturers in addition to tenure-track faculty members (at Princeton CS). They are extremely good and greatly decrease our load, so we haven't faced burnout so far. That said, enrollment was up so sharply this semester that we had to hire lecturers on a month's notice, which is kind of insane. Also, we have a new industrial Master's program which lets us increase our number of TAs.

In spite of having adapted in all these ways and money not being a problem, we know that this won't continue to scale because enrollment growth shows no signs of slowing down. We're not sure what the long-term solution is going to be. Online education is part of it (and we're on Coursera), but so far we're not using it in a way that decreases our teaching staff requirements.

Strange times.

randomwalker··on Unlocking Cellphones Becomes Illegal Saturday in the U.S.
I found out the answer to this just yesterday in a different context. Here's how it was explained to me:

When those who opposed the impending passage of the DMCA realized they couldn't defeat it, they (mainly EFF at that point) decided to salvage what they could, which is to stick in a clause to allow the Librarian to decide exemptions. The *AA didn't try to shut that clause down because they thought it was basically a joke and would never amount to anything. But in reality the Librarian has indeed exercised some power, so it's considered a minor win for consumer advocates.

randomwalker··on "Aaron's Law" Might Be Good
The eminently qualified Jennifer Granick calls it "a great first step" https://cyberlaw.stanford.edu/blog/2013/01/thoughts-zoe-lofg...

See also her writings over the last few days https://cyberlaw.stanford.edu/about/people/jennifer-granick

randomwalker··on Petition: require free access to publicly-funded research
That is correct, and many people are upset about the lack of response https://twitter.com/mattblaze/status/290503781093871616
randomwalker··on A Northwest Pipeline to Silicon Valley
The article is timely. A couple of additional data points:

UW CS had a huge increase in admissions (and yield) for the Ph.D. program this year. IIRC they mentioned the incoming class size is almost twice what they had last year.

Aided by a recent budget increase, the department also hired a massive number of new faculty this year.

randomwalker··on An Analysis of Anonymity in the Bitcoin System
Previous discussion: http://news.ycombinator.com/item?id=2800790
randomwalker··on About 33 bits
That's a great question with no simple answer. I've written two essays about this that look at it from two different sides:

http://33bits.org/2011/10/18/printer-dotspervasive-tracking-...

http://33bits.org/2011/06/08/the-many-ways-in-which-the-inte...

The synopses of the two posts are:

My opinion is that it impossible to put the genie back into the bottle — the cost of tracking every person, object and activity will continue to drop exponentially. ... If we accept that we cannot stop the invention and use of tracking technologies, what are our choices? Our best hope, I believe, is a world in which the ability to conduct tracking and surveillance is symmetrically distributed, a society in which ordinary citizens can and do turn the spotlight on those in power, keeping that power in check. On the other hand, a world in which only the government, large corporations and the rich are able to utilize these technologies, but themselves hide under a veil of secrecy, would be a true dystopia.

and from the other side:

There are many, many things that digital technology allows us to do more privately today than we ever could.

[examples snipped, but I recommend taking a look at the post]

Of course, I’ve only presented one half of the story. The other half, that technology is also allowing us to expose ourselves in ways never before, has been told so many times by so many people, and so loudly, that it is drowning out meaningful conversation about privacy.

Although these two opinions might at first sight seem contradictory, they are not. Some day I will get around to putting the two sides of the argument together into a coherent narrative that explains the nuanced scenario that I think we're heading towards, but for now I will offer you the above articles.

randomwalker··on About 33 bits
See <strike>comment #12</strike> the comment posted at February 12, 2010 at 5:15 am in the blog post.* The term entropy refers to uniqueness.

As for the development of algorithms to gather those bits, that's what my entire Ph.D. is about and what my blog is mostly about. This is what I've been proving for the last 6 years.

*Just realized comment numbers are unstable. Bad wordpress.

randomwalker··on About 33 bits
This is a well-known technology :-) See http://en.wikipedia.org/wiki/Keystroke_dynamics

In the research community it's a proven and accepted concept. There are products in the market that do two-factor authentication based on password + keystroke dynamics, but I don't know how well they work.

randomwalker··on About 33 bits
Hey, I'm the author of this blog. Much of my previous deanonymization research has been discussed on HN; see http://www.google.com/search?q=33bits.orgsite:news.ycombinat... Also, if you find the premise of the blog interesting check out the sitemap linked from the page.

But since this post is about the About page, let me share a couple of lessons I've learned from the blog, which has been more successful in communicating my research than I'd dared to hope for when I started it 3.5 years ago.

1. Those of us working on technical areas often struggle to explain our ideas to others not as technical, in a way that avoids oversimplification and losing essential meaning. Sometimes you'll discover an analogy or metaphor or phrase that does both. Seize those chances, they're powerful.

2. Coming up with a name is more important than you might think. If a good name will make your idea or product even 5% stickier, it follows that it may be worthwhile to spend 5% of your time just coming up with the name. One way to do it is to be constantly on the lookout for a good name while you're working on the product.

3. If you're writing about something that has policy implications, and want it to be read in Washington, it's hard but not impossible. Two important requirements are to network and build up an audience — they aren't going to read your blog just because it ranks high in Google searches — and to use language that non-technical people can understand.

Happy to answer any questions!

randomwalker··on Is Writing Style Sufficient to Deanonymize Material Posted Online?
The algorithm achieves significantly higher accuracy if it has more text per author. Also, if you're willing to do human analysis on a few dozen candidates after algorithmically shortlisting them, that gives you a further advantage. Finally, there is much room for straightforward algorithmic improvement (e.g., ensembles of classifiers) that we didn't have time to fully investigate. In short, IMO it's just a matter of more data and slightly better ML, not fundamental improvements.
randomwalker··on Is Writing Style Sufficient to Deanonymize Material Posted Online?
That's a good question. First, I believe that intelligence agencies are already well aware of the potential of technology like this, and at least some, like the NSA, could very well be ahead of public research. Second, research such as ours is intended to demonstrate a proof of concept, and it takes a lot of work to turn it into a reliable tool — for example, we restrict ourselves to English text. For those two reasons, I think our work does little to directly help governments and other oppressive entities. On the other hand, publicly available research is effective (we hope) in raising awareness of the threat, so on balance it does more good than harm to people writing political blogs.

As for practical tips to defeat stylometry and such, organizations like the EFF specialize in doing that, so I will leave that to them. Comparative advantage, etc. If you would like to help, you are more than welcome.

randomwalker··on Is Writing Style Sufficient to Deanonymize Material Posted Online?
Sure, Π is irrational, so it can't be exact :) What I meant is that 1147 is the closest integer to Π*365, which IMO is still an awesome coincidence!
randomwalker··on Is Writing Style Sufficient to Deanonymize Material Posted Online?
Lead author here. Since my serious thinking on this topic started when I responded to this Ask HN post[1] Π years ago[2], it's nice to see this posted here, to come full circle in a sense. Happy to answer any questions.

[1] http://news.ycombinator.com/item?id=413730

[2] No, really, it's been exactly Π years to the day :-)

randomwalker··on Fields medalist Tim Gowers: Elsevier — my part in its downfall
Summary: Gowers outlines the extraordinarily oppressive business practices of academic publisher Elsevier, explains why they are able to continue to do so in spite of widespread anger amongst the community (collective action problem), and goes on to explain how we might be able to solve this problem by publicizing the actions of people who've taken a stand.
randomwalker··on Computer Scientists and Google+: Something Interesting is Happening
This is bizarre. My post listed four factors. Your tl;dr lists one of them, and then claims that my post is simplistic.
randomwalker··on Computer Scientists and Google+: Something Interesting is Happening
Somebody asked me this question earlier. So I went and counted; it was under 10% (sample of around 100).

I find it interesting that whenever there's a post on HN that's supportive of Google+, the community collectively reacts with extreme skepticism. Just... interesting, that's all.

randomwalker··on Computer Scientists and Google+: Something Interesting is Happening
OP here. I'm wondering if parent commented in the wrong thread or something. This is completely irrelevant to what I was saying. You somehow seem to think I was complaining that it is hard to publish papers. No, that wasn't even remotely, _remotely_ related to what I was saying. The rest of my post is even less related.

I would like to hear some actual commentary on the issues I raised. If you have a problem with CS research in general, that's fine; I suggest you do a separate post about it. Thanks.

randomwalker··on [dead]
Debunked here: http://scienceblogs.com/pharyngula/2011/05/dichloroacetate_a...
randomwalker··on AdBlock Plus will soon allow "non-intrusive" ads by default
It is unfortunate that many people commenting here didn't bother to RTFA. In particular, from the FAQ:

Are you stupid? Nobody wants this!

The results of our user survey say something different. Only 25% of the Adblock Plus users seem to be strictly against any advertising. They will disable this feature and that's fine. The other users replied that they would accept some kinds of advertising to help websites. Some users are even asking for a way to enable Adblock Plus on some websites only.

https://adblockplus.org/en/acceptable-ads

That page is the real article, linked from the first paragraph of the submitted article, which is more like a changelog.

Please also note that Wladimir has been talking about this for years and didn't suddenly get this into his head. From a 2009 post:

As I stated many times before, my goal with Adblock Plus isn’t to destroy the advertising industry. ... So the idea is to give control back to the users by allowing them to block annoying ads. Since the non-intrusive ads would be blocked less often it would encourage webmasters to use such ads, balance restored.

http://adblockplus.org/blog/an-approach-to-fair-ad-blocking

randomwalker··on Complaints From a Single Doctor Caused Government to Take Down a Public Database
Thanks for the mention.

Unfortunately, all too often regulators and Government agencies take the wrong lessons from de-anonymization -- remove data altogether, try to ban de-anonymization, etc. [1] I'm actually visiting D.C. right now with my policy hat on.

In this case, I think we should be having a conversation about whether doctors' right to privacy is more important than public interest and patient safety. Ironically, a major reason why medical practitioners are often against public data release/reviews etc is apparently because they cannot publicly refute allegations or bad reviews, which is in turn because of patient privacy. Sometimes it feels like a morass of bad laws with unintended consequences.

[1] Recent proposed changes to HIPAA do exactly that, without even an exception for research.

randomwalker··on $200,000 from Kickstarter = 1 year of runway for Diaspora
It's a purely qualitative study.

You're right, it's a bit outside the normal ambit of computer science. I'm interested in multidisciplinary questions; we do also have non computer scientists among the authors, for example Prof. Nissenbaum http://www.nyu.edu/projects/nissenbaum/

Overall, as you suggest, our analysis is a mixture of CS, economics and social science.

randomwalker··on $200,000 from Kickstarter = 1 year of runway for Diaspora
Wikipedia lists close to 40(!) distributed social networks. [1] None have achieved a meaningful amount of adoption.

About a year ago I became interested in why this is the case. I'm an academic computer scientist with a strong interest in the startup/web tech scene (see profile for info). My colleagues and I have been studying not just decentralized social networks, but the more general concept of decentralized architectures for personal data. This includes "personal data store" efforts which are popping up all over the place, with similarly dismal adoption, "infomediaries" which were the rage in the late 90s (before the dot com bust ate them up), etc. Overall, we've looked at around 80 companies, projects, and proposals.

If the amount of reinvention in this space is surprising, the almost wanton refusal to learn from others' past mistakes is shocking. These projects seem to do the same things wrong and fail for very similar reasons, chief among them building more technology when it's not really technology that's holding things back.

There is the widespread — but often unstated and always unexamined — belief among the participants that moving to a decentralized setting is a magic cure-all for the problems that ail today's status quo, such as privacy and interoperability. Sadly, this belief is simply wrong.

I certainly sympathize with the urge to pat these guys on the back for trying, but is it really courage or foolhardiness? If 10,000 amateurs have tried to solve P =? NP and failed, do we encourage the 10,001th guy to give it a shot as well, or do we tell him to learn some math and CS first, and gain an appreciation for why the problem is hard and probably not worth taking on unless you really really know what you're doing?

[Our study is not out yet, but feel free to contact me if you're interested in discussing this.]

[1] http://en.wikipedia.org/wiki/Distributed_social_network

← PreviousPage 3 of 15Next →