Google acquires reCAPTCHA
googleblog.blogspot.com
googleblog.blogspot.com
The logical solution was to buy them and Google-size their infrastructure.
From a personal perspective, I don't think this really changes how I feel about reCaptcha either. They'll still have the same mission -- unless Google stops them reading out of copyright classics and puts them onto scanning in Dan Brown novels, which I think is unlikely.
I thought you meant that Google would have actually considered ditching their own captcha in favor of recaptcha for google sign-up.
Makes it a good fit for Google, doesn't it?
http://www.merriam-webster.com/dictionary/lucre
The most common phrase that uses the word lucre is "filthy lucre," a reference to money being something that distracts noble people from pursuing altruistic goals.
That said, he was a very good teacher. Great lectures with lots of student involvement. He also had some creative ways of catching cheaters.
When 251 started, everyone was required to sign an honesty policy about cheating, etc. The policy explicitly stated that students shouldn't use online (or other) resources for help while working on homework sets.
One of the early sets featured a fairly difficult problem. Many students succumbed to the power of the Google search box and typed in some keywords around the problem. In the first page of organic results there was a site full of algorithm question and answers, including this problem but none of the others in the set. The site's domain looked innocent; I think it even had the word math in it. Many clicked, found the answer and either used it to guide their thinking or copied it close to verbatim, cheating either way.
At the lecture following the assignment due date, Luis announced that he knew X number of people had cheated on the last homework set. He stated that anyone who came clean would receive a zero on the assignment but wouldn't be reported to the dean. Since many students worked on these problems as groups by splitting up the work, a few though that their friends had ratted them out. Others had cheated differently and thought they had been caught but didn't know how. Regardless of the method, everyone who cheated came clean.
At the next lecture, Luis told us how he did it. The innocent site was actually his, running off a box in his lab. A simple whois confirmed it. He had logged all IPs accessing the page and correlated them with student usernames. This wasn't that far-fetched: since everyone at CMU registers their computer with computing services in order to get network access, a reverse DNS to hostname, which usually contains the username, would be enough to identify. Besides, he could also have worked with computing services to get the original registration username as well.
Even if he had identified only 10% of X, he managed to get everyone who cheated to admit it and send a clear message at the same time.
The weirdest effect of this is that the most dangerous viruses tend to burn out. If one of those ever came along with a really long incubation time for the disease but a much longer time for contagion that might be a problem.
There are horror movies that scare me much less than something like that.
Basically they stuffed the ballots on the Time online voting page where users could vote for most influential person of the year. Most influential person ended up being Moot, founder of 4chan. Just to rub it in they managed to spell "marblecake also the game" out of the first letters of the top 21 entries.
Note: This hack shows more about Time's incompetence than recaptcha's.
In-depths story here: http://musicmachinery.com/2009/04/27/moot-wins-time-inc-lose...
The way I understand it, reCAPTCHA gives you 2 words to analyze, one which reCAPTCHA "knows" and one it is trying to learn about. As long as you get the word it knows about correct, it'll say you're a human. You're answer to it's unknown word is simply stored and (I'd assume) analyzed until it has enough responses to consider it a known word. Knowing this allowed 4chan users to nearly cut their response time to the captchas in half.
I'd be fascinated to hear more details (or be corrected) on how reCAPTCHA works if anyone has them.
Update – Just to be perfectly clear, anon didn’t hack reCAPTCHA. It did exactly what it was supposed to do. It shut down the auto voters instantly and effectively. The only option left after Time added reCAPTCHA to the poll was a brute force attack. Ben Maurer, (chief engineer on reCAPTCHA) comments on the hack: “reCAPTCHA put up a hard to break barrier that forced the attackers to spend hundreds of hours to obtain a relatively small number of votes. reCAPTCHA prevented numerous would-be attackers from engaging in an attack. In any high-profile system, it’s important to implement reCAPTCHA as part of a larger defense-in-depth strategy”. As Dr. von Ahn points out “had Time used reCAPTCHA from the beginning, this would have never happened — anon submitted tens of millions of votes before Time added reCAPTCHA, but they were only able to submit ~200k afterwards. And to do this, they had to resort to typing the CAPTCHAs by hand!” One thing that Time inc. did that made it much easier for the anonymous hack was to allow leave the door open for cross-site request forgeries which allowed anon to create a streamlined poll that never had to fetch data from Time.com.
Sorry if I didn't make it clear enough that the fault lay with Time and not recaptcha.
Edit - To clarify, I think that making information in the public domain more accessible is a very worthwhile project.
It's only very recently that they started talking publicly about their relationship with the NYT, and they've still said diddly-squat about public availability of the results.
For a while google had one of the hardest CAPTCHAs but bots just keep getting better, and they were cracking it more and more often. But you can't just make the CAPTCHA even more difficult, it was already fooling a lot of the humans.
Note that a bot does not need to successfully solve it 100% or even 90% of the time. I'm not sure of what the exact figure is but at some point the bot reaches parity with humans even if it only succeeds say 25% of the time. That's because on average now it only takes 4 guesses to guess correctly.
And I don't think reCAPTCHA is stronger then the multicolored CAPTCHA google had (still has?)
I think google is realizing that they just can not stop the best CAPTCHA cracking bots, maybe they can stop a lot of them, or a lot of the not so smart bots.
But they also can't just give on CAPTCHA, and let anyone and any bot in without even trying.
Thus reCAPTCHA because you might as well do a little public good while you're at it. If nothing else, you've at least made some old texts more readable.
As a trivial side-note: When encountered with a ReCAPTCHA, I'll fill out one of the words and put in gibberish (or other text) for the other. For some reason I find it satisfying to "pull a fast one" on any captcha service.
Why? What do you achieve except being told you're wrong every once in a while? I mean, you're trying to bother a machine. Isn't that a bit like a reverse Turing test?
I suspect the amount of noise that ReCAPTCHA filters out from automatic attempts is several orders of magnitude larger than anything any group of actual humans can generate.
As far as what satisfaction I gain from it... I suppose I find it to be a sort of rebellious act. Also, I did not get accepted into Carnegie Mellon.
Since computers have trouble reading squiggly words like these, CAPTCHAs are designed to allow humans in but prevent malicious programs from scalping tickets or obtain millions of email accounts for spamming. But there’s a twist — the words in many of the CAPTCHAs provided by reCAPTCHA come from scanned archival newspapers and old books. Computers find it hard to recognize these words because the ink and paper have degraded over time, but by typing them in as a CAPTCHA, crowds teach computers to read the scanned text.
The way I understand this, is that the user is presented letters from archival newspapers and must type in the text he sees, and recaptcha uses that text to improve OCR. But doesn't that imply that recaptcha was unable to interpret the scanned text before ? If so, how can it then verify the correctness of the text the user types in ? If not, how exactly is this helping OCR ?
> But if a computer can't read such a CAPTCHA, how does the
> system know the correct answer to the puzzle? Here's how:
> Each new word that cannot be read correctly by OCR is
> given to a user in conjunction with another word for which
> the answer is already known. The user is then asked to
> read both words. If they solve the one for which the
> answer is known, the system assumes their answer is
> correct for the new one. The system then gives the new
> image to a number of other people to determine, with
> higher confidence, whether the original answer was
> correct.
http://recaptcha.net/learnmore.htmlAlso reCaptcha's present two words. One of which is known, and used as a normal captcha, the other of which is an unreadable word. Thus if a someone passes the captcha word, its somewhat assumed that he also correctly typed in the unknown word because the words are presented in randomized order.
Finally, as a funny anecdote, Luis told me once of this time that they noticed there was a captcha farm in latin America paying pennies for poor latin american's to solve reCaptcha's for them in order to gain the ability to generate illegal spam. Luis started sending them whole sentences of reCaptcha's and eventually made it to full paragraph's. Ironically, the reCaptcha farm continued to solve these captcha's. He eventually stopped their activity, but he said the data they generated was some of the best he's ever gotten... :P
Ideas?
http://vonahn.blogspot.com/2009/07/hottest-people-in-cs.html
He's had a lot of neat ideas which align well with Google. I think he'd be a great addition to their team.