How to parse a sentence and decide if to answer "that's what she said"?
quora.com
quora.com
Now I want to build a "that's what she said" bot for Twitter, which will parse new tweets and reply accordingly. (Maybe I will, as a learning exercise.)
Actually, a fair part of the links about programming are just the same both here and on Reddit.
Edit: All the same, I'll try giving /r/programming a chance. As a whole, Reddit looks less civil than HN, but maybe I haven't been fair. People tend to criticize subjects that they don't know about or have never experienced 1st hand, and I'm guilty of that here.
HN makes my daily read list, even if it is a quick glance. Great community.
Well, the thing is that people on reddit know better than to take themselves too seriously, and as a result you get plenty of amusing things mixed in to the comments.
For HN, on the other hand, commenting is Serious Business(TM), not to be wasted on such frivolities.
Some people like the HN approach. Personally I prefer to have a little fun every now and again.
Not claiming it's accurate - I'm not a Redditer. Just clarifying.
I'd actually be intrigued at how simple the algorithm could be and still be mostly on the mark, rather than how much NLP you could squeeze in...
That being said, I don't think the discussion on Reddit is worthless at all. Yes, there's quite a bit of joking, which is good to read at 4pm when you are stuck at a bug. But this particular topic is an excellent example of the value of the Reddit discussion: reading through the jokes, etc. provides interesting thoughts about implementing the TWSS bot.
So, the two discussions are complementary. For particular stories, I head over to proggit and see what the Redditers have commented.
It seems, unfortunately, that the recordings aren't publicly available, although archive.org seems to have small sample of the transcripts: http://www.archive.org/details/SwitchboardCorpusSample
One particularly interesting section of the above readme was the section on technical issues (http://www.ldc.upenn.edu/Catalog/readme_files/switchboard.re...) including this example:
"ii.) The third problem was small changes in synchrony between A and B, due to a pseudorandom dropping of 2 ms chunks of data on either side. Over the course of a 10 minute conversation, these could accumulate to a differential of 30 or 40 msec between sides--enough to change a cross-channel echo from inaudible to audible, for example, or from barely audible to very noticeable, for a human listener.
When this bug was finally run down, it turned out to be a piece of code in the utility which extracts conversations ('messages') from the Robotoperator message master file. The code performed a check at each data block boundary to see if the first two bytes had the values 'FF FF'; if so, these were interpreted as header information, and the 16 bytes beginning with "FF FF" were discarded as not part of the speech data. This code was a relic from an earlier version of the Robotoperator which did not deal with mu-law values, and thus never encountered FF in data. In mu-law data, FF is one of two ways of representing zero signal level ('minus zero'). The offending lines of code were removed and the problem ceased."
http://www.cs.washington.edu/homes/brun/pubs/pubs/Kiddon11.p...
Enjoy. It's easily the most hilarious academic paper I've read.
As an example of other funny articles, this legal one on the word "fuck" is one of the best I've read recently (found it linked here on HN previously):
http://moritzlaw.osu.edu/faculty/articles/fairman_fuck
Of course, the annals of improbable research and Ig Nobel prizes yield several other examples :)
I think that answer at Quora, while awesome, should sufficiently kill TWSS.
Although now that I think about it... twitter probably often suffers from the exact same issue, although somewhat mitigated as people often include the original tweet in their reply.
Using average read/write times for a message - taking its length into account - it should be fairly easy to check if someone could have read and written a response to a message. Assuming it is a response if he could have written one in time should be fairly accurate.
You could probably even cheat and use a constant time while still getting good results.
"[4] I think I once heard of an MSN chat transcript dataset that was really awesome, but I can't seem to find mention of it anymore. Let me know if you know where I can find that or any other instant message datasets. I know that some IRC rooms get publicly logged — is there a single place where one could grab all of them at once?"
You could use pre or post filtering to weed out connect/disconnect and other noise.
I only thought of this because there was like two "TWSS" replies in the past couple of days on #rubyonrails
Reading this post I remembered a small project called sociograph a friend of mine created. It's an IRC bot that logs messages and draws a graph of the people communicating with each other in realtime, see for an example http://www.youtube.com/watch?v=A_ah-SE-cNY.
1) Ask people in mechanical turk to write those sentences. Then you can ask others to verify them - you get far with few dollars
2) Include higher level features - for example bi-grams, there is more information in them
Also: corpus at http://thatswhatshesaid.com/ (I have no relation with this site)
In any case, I think most of us are now primed to try out some implementation or another now. It would be a lot of fun regardless of the quality. Actually, the false positives are probably funnier.
An alternative solution is to attack the problem backward, training on terms (words or phrase) from sex-related conversations (such as adult chatroom transcripts). Then, from general corpus (Twitter or generic chats) identify terms that highly cooccur with those sex-terms. I would still use a Bayesian classifier, with strong prior against labelling something as a TWSS.
twss = lambda sentence: True
Now, will you stand over.... there?