Seems like scraping all closed captioning would be very valuable data indeed. Is there anyone else doing something like this that provides an API or data feed?
We're providing API access to some select partners for data experiments. The API is centered around trending, topics, search and metrics but does include some restricted transcript access. Ping kevin at boxfish (me) if you've got some ideas.
We use TVEyes.com in our news monitoring product and it's very close to real-time. It does make mistakes sometimes, but I'm not aware of any perfect transcription software.
Copyright law would make such a feed of closed caption transcripts illegal to distribute.
That makes sense. Would be legal to capture the data and present it similar to a search engine? I'm guessing there is some sort of precedent for that sort of thing?