In Case You Wondered, a Real Human Wrote This Column
nytimes.com
nytimes.com
For those who don't know, if you see a story in your local paper, and it doesn't involve a car crash, crime, weather, or sports, it was probably placed there by a PR representative. Most of the things you read are not the result of random reporters deciding to cover X or Y, but a paid, concerted effort to place story X or Y in the paper by providing the paper with a fully pre-digested story to perhaps rewrite, or perhaps not.
The words "narrative science" appear 14 times in that story, including such clunkers as "To generate story “angles,” explains Mr. Hammond of Narrative Science...." when Mr. Hammond has already been introduced earlier in the story. It even includes pricing: hey readers, this is not only cool and will win the Pulitzer Prize, but it's cheap too! No mention of competitors... It reads like an ad because it is an ad.
This story was provided, probably almost word for word, by a PR person to the NYT reporter.
I'm not sure if computer-generated text will be better or worse than the media system we have now.
Also, it happens all the time. And with a publishing machine as big as the NYT, it's bound to happen there as well. Remember Jayson Blair? Different symptom, same cause.
Automating long-form narrative content is a relatively new concept. It's not perfect (yet) so I appreciate the skeptics.
Reporters are usually liberal. And they come from a time where newspapers were profitable and could be idealistic and still make money. They could answer to a higher standard because they didn't have to worry about losing their jobs for money reasons.
As far as right now the ship is going down but they still don't view the threat.
I have a question, though. Why don't they also create an expert system that writes PR pieces? I imagine it would be really nice tool for people in the business. You write down the variables (who, what, when)from an interview with a client in a form and you get to show him an instant first draft, you correct it together and you publish it after leaving the room.
Or even better yet, take that data and use advanced algorithms to embed mentions of clients, products and initiatives in bigger articles.
This is how a PR piece usually goes:
1. Case, problem, solution. Short story here.
2. Introduction of company
3. A Question and an Answer
4. Another look of the company, credentials.
5. This is really cool think about it. There may be a sound (text) bite here.
Do you see the patterns?
It's cheaper and easier to just use the endless supply of interns.
Although there are hidden costs even behind unpaid labor. Office space, HR costs (hiring isn't easy), training, etc, etc.
Then again, PR people have a strong commercial interest in making it appear so that only stories pitched by PR people get published.
My (few) experiences in getting stories through have been to the contrary, usually just journalists picking up something I blogged. An exotic example: http://bergie.iki.fi/blog/on_usb_fingers_and_world_news/
Last fall, the Big Ten Network began using Narrative Science for updates of football and basketball games. Those reports helped drive a surge in referrals to the Web site from Google’s search algorithm, which highly ranks new content on popular subjects, Mr. Calderon says. The network’s Web traffic for football games last season was 40 percent higher than in 2009.
This planted PR piece uses a bit of customer data to demonstrate that the technology makes great business sense. So far so good. This particular piece of data, though, will be read as "Markov chains creating content to spam up the search results" by folks inside the Googleplex, and it makes them look stupid. Don't make Google look stupid. If you do make Google look stupid, don't brag about it in the NYT. It will not end well.
There's an easy way to achieve that: have an actual human write it. This solution does not necessarily win one good will with Google: for one obvious example, most of the content farms were farming manually rather than farming with Markov chains.
I also think "good content" only craters the approach to the bridge of describing both a) what actually ranks on Google and b) what, in an ideal world, Google thinks would rank on Google.
This story was provided, probably almost word for word, by a PR person to the NYT reporter.
Definitely not. The idea for the story was provided to the reporter, probably by a P.R. person. The reporter conducted interviews with representatives of the company, some of whom are quoted in the story and some of whom aren't. The reporter went away and wrote up the story himself. It was edited by at least one line editor and at least one copy editor. The reporter, line editor, and copy editor have all beaten many competitors to obtain jobs at the most prestigious company in their field.
If anyone at the New York Times were found to have submitted a story that was "provided ... almost word for word by a PR person" that person would be fired and the paper would issue a public apology.
Again, I'm not defending the story, and I'm sure the PR person who pitched it was thrilled by it. But, you know, Steve Lohr's byline is on this story, and you've accused him of pretty bad professional misconduct, and I don't think that's warranted.
It happens more than you think.
Like every other daily newspaper, the Times produces some rushed, lazy journalism (as well as some very good journalism). But there are certain lines they don't typically cross, and literally taking dictation from PR is one of them. You might argue that the difference isn't meaningful, and that rules like "Don't just copy out someone else's text" are a fig-leaf to hide bigger problems. I'd have a lot of sympathy for that argument. But if we're going to criticize the Times we should criticize them accurately, for the things they're actually doing wrong, rather than accusing them of doing things they haven't done.
> primarily a low-cost tool ... for local youth sports .... and financial results of local public companies ... “Mostly, we’re doing things that are not being done otherwise,”
Then, once you have some customers - any customers! - you improve it, bit by bit. It doesn't need to be perfect in the first place; it doesn't need to be perfect in the end. It just needs to be good enough to be useful.
> [customer] worked with Narrative Science for months to fine-tune the software
As for the technology itself, we're not told anything of its details, just what it can do. This is a marketing article, not a tech report. It would be interesting to see the models they use for stories, and whether they use grammars for the overall structure. These are very narrow domains, which are the easiest to start with: you could enumerate all the standard cliches, understand when they apply, and tweak the model. That's where the journalist expert domain knowledge of the two founders would come in handy. BTW: "easiest" is only relative - it would still be very difficult (almost impossible), and kudos to these guys for actually doing it - and even better, making an actual business out of it.
It reads like a 50's Asimov story - the future is finally arriving.
But a Pulitzer in 5 years is absurd, either cynical puff or visionary bravado. Theoretically possible, I think, maybe in 50 years - the figure I've long given for strong AI. ;-)
It's clearest to see when a product exists to solve a problem, but it's too expensive for some people (or some situation). In the article, the problem of reporting on local sports/financials has a solution (reporters), but the value of that news isn't worth their time: their time is too expensive. So the newspaper doesn't "consume" a solution to the problem of reporting that particular news.
By targeting this non-consumption, the startup doesn't compete against reporters (yet...), so it provokes no desperate fight for survival.
It's an term from Clayton Christensen, who wrote The Innovator's Dilemma, though he doesn't use it til "The Innovator's Solution", and expands on it in "Seeing What's Next".
Reporting a day at the races or the markets is easy because we know which kinds of data are relevant and we have them available.
In some fields (eg, finance) one can conceive a computer based process that would do a better job than most investigative journalists. It can't deal with missing data, but it can discover inaccuracies, unusual events and suspicious patterns and in some limited fields this is enough.
For example, an AI based process might have been just as good at finding the problems at Enron as conventional journalists were (since it the problems there were mostly uncovered by forensic accounting on their public balance sheets):
But hard information was scarce. "It's almost as if you have to use forensic accountants when you're doing a company story because many companies are using very aggressive accounting techniques that are perfectly legal," Shepard says.
http://www.washingtonpost.com/wp-dyn/articles/A64769-2002Jan...
No matter how good the algorithms get, they are still limited by their input, the statistics. If for example a player scores a very unusual goal, say a bicycle kick in soccer, then a real writer who actually saw the match would surely mention it. An algorithm could not if there is no field for unusual goal in the match statistics.
Maybe though if this kind of thing continues to improve the humans will have to start doing more serious analysis instead of fluff coverage.
"I wrote this article with one mouse click"
http://coding.pressbin.com/60/I-wrote-this-article-with-one-...
I can't imagine the sort of code base that would be needed to make these stories not seem formulaic.
[1] http://globalmoxie.com/blog/page-one-safari-chrome-extension...
There are certain topical areas which lend themselves to automated content generation. Sports, financial news, weather, astronomy (astrology isn't worth mentioning), earthquakes and other severe events, machine monitoring.
Domains in which a quantified or measured outcome tied to a specific point in time or event (final score, market close, daily forecast, etc.) occurs. The important data has already been highlighted, all you've got to do is sprinkle some syntactic sugar around it.
Oddly enough, these are areas in which you're already most likely to find existing "AI"-type content generators.
In areas in which you've got to do significant determination of what is salient, the approach isn't nearly as successful.
----------------
We’re sorry, but we’re unable to process your request because another entity has made a previous request concerning this username. If you are still interested in claiming the username, you may contact us in 60 days for an update about its availability.
---
You have reached the right channel for these requests. As mentioned earlier, we have no further information to share with you concerning the username "xxxx" (marked out). We will be unable to assist you further from this alias.
----------------
What human being talks like that?
Naming collisions have to be a common occurrence for Facebook. It’s sufficient to write exactly one mail for such cases. There is no need to re-write or change things around, it’s always the same answer to the same question.
Using a robot to write stuff like that seems wasteful – I don’t even think it would currently be possible.
That's not to say their technology couldn't be improved to search the web and see what past events are relevant, but providing good insights about the implications of the facts will be a whole lot tougher. I don't think journalists need to be shaking in their boots unless they only deliver the quality and depth of results that this algorithm delivers.
Sure, there's no way that my profession and the great majority of jobs on the internet would be possible if we rely on human switchboard operators rather than relying on automation. That doesn't mean it will be true for the next advances in technology, does it?
I'm willing to believe the underlying machine learning technology is very clever, but I'm also willing to believe a specialised toy script could produce similar results, even if you had to hard code the minimum winning margin for a "rout".
As for the Freakonomics comparison, they seem to have missed the appeal of Levitt: that his ability to posit a plausible causal relationship between two apparently unrelated variables. Any idiot can summarise "remarkable findings" based on spurious correlations.
Bergie had a strong start on the office day, closing four bugs in row. Then luck turned and he broke the build...
http://techcrunch.com/2011/08/15/yc-funded-marketbrief-makes...