Boxfish: Realtime Index of Every Word Spoken on TV
techcrunch.com
techcrunch.com
Doing frequency and sentiment analysis on this dataset would be pretty interesting.
Closed captioning does have a lot of noise in it, but we've done a lot of work to tidy that up. We also have the benefit of capturing so much data that the noise doesn't matter as much.
In real-time our systems extract and generate lots of information from the closed captions. Our NLP system identifies entities (e.g. whitehouse, chris brown, amanda berry), does frequency counts and we have some very large graphs of entity co-occurrences (leads to our statistical learning), e.g. Rihanna is commonly associated with Chris Brown.
We do a bunch more analysis on our graphs, including Latent Semantic Indexing (LSI), which helps drill down into quantifying the relationships between entities. Related to that we generate TFIDF scores for all identified entities, i.e. gives a sense of how "important" an entity is.
By combining our large scale entity graphs (both Frequency and LSI) with streams of closed captions we also do real-time topic extraction at multiple time scales, e.g. what is the topic of conversation on CNN for the last minute of conversation, for the last 5 minutes, for the whole program, etc.
Feel free to ask me more.
Part of the motivation for taking a statistical NLP approach is that it gives us more flexibility for processing foreign stations / languages (we don't yet do that).
I wonder could you time and geo-shift closed captions, i.e. show closed captions in two languages at once on the same TV program? That could make an interesting language learning tool and an interesting training set for machine translation.
Decision came down to speed is reasonable, library has some other math functions I need, and time taken to build.
What's your average accuracy?
This memory is kind of hazy, ISTR it's from 2005 or so.
In addition to capturing and indexing the subtitles, it also captured the video, and so allowed the captions to be used as an index to the video.
I don't doubt someone can come up with an even earlier incarnation!
A few years ago, someone documented how to use an Arduino + Video Experimenter Shield to easily log closed captioning data (http://blog.makezine.com/2011/08/16/enough-already-the-ardui...). Never got around to messing with it, but I can imagine 100 interesting things to do with that data.
Very cool company. I'm glad someone's doing this.
Boxfish, twitter, YouTube, Siri, and now with Ray Kurzweil @ Google... thinkers are converging on doing to every other form of content what Google did for structured documents.
The NLP trend is going to be amusing to watch at least (Siri, Summly), and whether its time has come in the next 5 years or not I'm not certain. But I know Ray Kurzweil knows this technology is inevitable.
--
As for BoxFish, I think this is a good example of a neatly executed, well funded startup with experienced founders and a solid space. No drama, no demo day, no immediate fires to put out, cool $3m in the bank, Deutsche Telekom AG subsidiary negotiating their deals for them, and "Yahoo just bought a kids startup for 17m" - the topic is hotter than others. This is the type of startup I for one daydream of having stock of or working at. Has high potential to be worth $mmms or $bn in the future - you know, that all depends and what not. But the makings are clearly there. Excellent work guys! Congratulations.
While the majority of TV is pre-recorded or repeated content (think of all the repeats of the Simpsons, Real Housewives of X etc). We know whether a show is recorded or live and the broad categories that a given show falls into. We also break up our trending calculations into different groups based, News, Sports etc and treat the data differently (as seen in the apps)
Also, bear in mind that while a show might be prerecorded it still may show useful data. For instance The Colbert Report and The O' Reilly Factor are usually recorded shows, however they can talk about drastically different things from show to show, and even between segments in shows.
I grant that useful trends are more difficult to extract from sitcoms and other things like that, but just because a show isn't live, doesn't that no useful trending information can be extracted.
We look at trending data over various periods of time, from minute length to longer so we can gather sentence level, show level, series level and even channel level topics.
We expose trending data via serach for the past 7 days in the app/website. Our API allows for longer term and more granular searches.
The page has since been taken down but here are two of our blog posts about the analysis.
* http://blog.boxfish.com/post/30997338037/obama-vs-romney-who... * http://blog.boxfish.com/post/32880728776/tvs-thoughts-on-our...
That's a really interesting question. Hopefully by opening up the API we'll enable more people to ask and answer those questions. How does social impact upon TV? How does TV impact upon social?
BTW we do trend identification across genres, channels and broader categories. I've often wondered what insights can be gained by looking at entities / things that are trending but not trending as strongly as leading news or sports events? Or looking at the rate of change of trends, i.e. identify slowly emerging trends?
Currently only for C-SPAN but that may change!
What is the reach? I know several people who would be interested in this for smaller countries.
I couldn't find this information on the homepage.