Programmer creates 800,000 books algorithmically, starts selling them on Amazon
extremetech.com
extremetech.com
When we first started seeing the automated content books, literally overnight our product set increased several million titles - I now estimate that about 20 million of the 50 million products we have are automated books (they come out with new "editions" all the time as well). This obviously has a massive impact on our search results - these books have keyword laden titles and descriptions, and without a solid identifier it was very difficult to get rid of them. Thankfully recently the suppliers that print these products have started flagging them as crapware.
As for the customer response to this type of product - it's definitely negative, with a massive return rate. As far as I am concerned, this is a massive scam - hedging on the fact that some people are to lazy to return the books.
In one question the interviewer asks: 'could you make a novel for instance?' to which he replies 'well, novels don't usually make money'. Wouldn't you think the correct answer would be 'well, it wouldn't be very good'?
Hauntology at its best[2].
Perhaps the same question can be posed to many startup founders in our community. Very often people in the startup scene seem much more interested in extracting profit (or really, just making the exit) rather than creating something good (in whatever way that you might personally define goodness).
First and foremost -- 1) Like Google it has a "see no evil" business model. ("If humans aren't in the approval loop, we're not responsible.")
2) Its behavior as a company is generally reprehensible (e.g. strong-arming state governments into not collecting sales taxes -- it wants to have its cake and eat it w.r.t. affiliate marketing)
3) Its business models reek of dumping (using profits in one domain to unfairly compete in others). The main difference is Amazon doesn't actually make profits, so it's burning investor capital to win new markets so it can burn more investor capital to win even more new markets, etc. (It will make it up on volume...) The fact that investors buy into this is odd, to say the least, but it's still pretty odious.
4) as a small publisher, I'm still ticked off by the way it allegedly "competes" in the ebook market. (Kindles are the hardest ebook readers on which to read ePubs, KDP does not in fact pay competitive royalties, but Amazon has bamboozled the tech press into believing it does, and even seems to have several governments helping it maintain its near monopoly on epublishing.)
5) Then there's the whole fulfilment / outsourcing thing (e.g. see here http://www.motherjones.com/politics/2012/02/mac-mcclelland-f... and the linked articles -- if you think Walmart screws its employees, Amazon et al are far worse).
He also has medical dictionaries which seem to confuse his cause, as he explains "For health titles, only the format editing and production side is automated. The text in the health books was written by medical professionals and edited by a professional editor; the computer expedited formatting using about 50 odd routines (the preface, chapter intros, glossaries, indexes, headings, margins, etc.)"
Is it it just me or does it seem odd that the agents with the most,
* technology
* resources (spanning all non-technology areas including business and legal)
* digital content access
* familiarity
* etc.
chose not to create such offerings, not even via subsidiary or ex-employees.
I'm talking, of course, of agents such as Google.
To me this speaks volumes of whether this is ultimately (legally) viable...
Note. Publicity around this particular offering has been around since at least 2007, from what I can tell.
Next we will have entire usenet archieves published as ebooks, least a few threads in there day of more interest.
At least you are not alone. Every sysadmin in the world knows how you feel.
Similarly, if I were to fund a human team of title-writers to create plausible original titles for a hundred original titles per hour, and fill the actual pages with random words from the OED... Advertising that as a book on a particular topic would likely be some type of fraud. The result is not far from what I've seen of this guy's work.
My favorite page: http://www.totopoetry.com/search.asp?word=truth I have used this approach to write definitions as well (www.websters-online-dictionary.org) The following contrasts definitions of zealously: 1. In a zealous manner. [Human] 2. In an enthusiastic, fervid, ardent or fervent manner. [graph theoretic] 3. In a fanatical manner. [graph theoretic]
A vid on fiction automation: http://vimeo.com/17168987 A debate/reaction amongst literature people: http://www.thepassivevoice.com/10/2012/can-robots-really-wri... Cheers Phil p.s. most of the “books” are used by businesses in narrow markets, and are econometrically estimated, not compiled from internet sources.
Also deviations in the data are recognized and highlited, but not (yet?) examined and elaborated upon. No doubt this will be possible in the near future.
So, kudos for the general approach. Can't wait for the automated reading programs for digesting these books. And that is meant only half jokingly. The data is there, the general knowledge is there and there is enough reasoning power to draw conclusions. The next step would be to make automated descisions based on the available data, so politicians could join the authors in beeing unemployed...
Well I for one welcome our new text blasting overlords.
[1] http://www.amazon.com/2009-2014-Outlook-Toilet-Seats-Greater...
Also, I didn't like the Scratch N Sniff parts.
Also, from the description: "…editorial decisions to include or exclude events is purely a linguistic process." Is it really correct to describe that as an editorial decision? (Not to mention "editorial decisions…is"?)
[0]: http://www.amazon.com/Basketry-Websters-Timeline-History-700...
I don't think that real authors would face any significant competition from this guy.
Here's an honest one that shines a light on the quality:
"The description for this book is TOTALLY misleading. It is NOT a book of quotations and phrases. It is a reference book of where to look to find possible quotations- like the old filecard cabinets in the library. On the few pages where you can actally find a quote, it reads like this one: Jack London, from Jerry of the Islands, "I am writing these lines in Honolulu, Hawaii." Huh?? That's it. That's all there is! I'm not sure who would use this book. Certianly not me! I was very disappointed as I was looking for a collection of quotes from notables like Mark Twain, Jack London, etc."
Edit: Sort by avg rating. Goldmine of comedy in the reviews for "Butts" and "Scrotum" books. Wow.
Here's the review for "The 2009-2014 Outlook for Plastics Lamp Shades in the United States":
"(4/5 stars)An instant classic in the Icon style
While this outlook hardly holds a candle to comparable classics such as The World Market for Silica Sands and Quartz Sands: A 2009 Global Trade Perspective, the information is invaluable for any red-blooded American. The five-year span is parsed in fascinating prose, and the 176 pages fly by, feeling almost like a 150-page work.
Don't let your lack of background knowledge deter you - there isn't too much reference to the 2004-2009 report, and most of the important information is explained in exposition.
Luther Blaze runs the show in this non-stop thrill ride of an economic adventure. His last outing ended in the government setting up a secret commission to investigate his possible wrongdoing in stopping the mysterious project Mantis, but now he's back, and ready to run roughshod over anyone in his way.
At a price of approximately 2.81 per page, you know you're getting your money's worth with this paperback. It's a perfect read for the park, or a lazy Sunday afternoon. On a side note, this book is a real pick-up gem! I personally attracted no less than three beautiful women, all of whom wanted my thoughts on the challenging themes and motifs. They all gave their numbers, and there's been no looking back!
The biggest problems with this book are largely physical aspects of the book. I didn't care for the font too much. And the beige background on the cover betrays the intrigue within.
The Icon International group has hit another classic out of the park. I look forward to the next book with baited breath, and I can't wait to see how Agent 71 and Dash get out of this jam."
The review is, indeed, ball-bustingly hilarious.
In this case algorithmically selecting stuff from around the web generates random books, but they have little value for the most part. Books that are researched and curated on the other hand have higher values. The only difference being the curation, not the basic information. So it gives you an insight into pricing curation.
Next up, a bot that sends out 800,000 DMCA takedown notices, oh wait we already have that for youtube.
Machine generated, or guided, curation may eventually become a good thing, but as you know all too well, there are far too many uses of such tech that are harmful (e.g. poisoning search results, &c).
It suggests that crooks in the 21st century don't rob banks, rather they rob a million people of 50 cents because none of them feels ripped off enough to prosecute and even if they did they are only out 50 cents.
Generate conference-ready CS papers in seconds!
I assume what pushed him into the market was the software's economic analysis of the latent market for spam books on Amazon 2010-2014
As a medical student, I particularly love books like this[0] with their description: "If your time is valuable, this book is for you. First, you will not waste time searching the Internet while missing a lot of relevant information. Second, the book also saves you time indexing and defining entries. Finally, you will not waste time and money printing hundreds of web pages."
How is this not fraud?
[0]: http://www.amazon.com/Stevens-Johnson-Syndrome-Dictionary-Bi...
For anyone not familiar with his work, I highly recommend his collection of short stories titled "The great automatic grammatizator". The major plot of the eponymous story is about a hacker who builds a novel-writing machine so that he can drive authors out of the market by out-producing them.
Very entertaining and, it seems, prophetic.
Because actually now I'm angry that tax dollars are spent buying, shipping, storing, garbage.
I just want to know, is this even legal? Is he selling books?
http://en.wikipedia.org/wiki/Book
"The body of all written works including books is literature. "
http://en.wikipedia.org/wiki/Literature
"Literature is commonly classified as having two major forms—fiction and non-fiction—and two major techniques—poetry and prose."
http://en.wikipedia.org/wiki/Prose
"Prose is a form of language which applies ordinary grammatical structure and natural flow of speech rather than rhythmic structure (as in traditional poetry)."
http://en.wikipedia.org/wiki/Natural_speech
"In the philosophy of language, a natural language (or ordinary language) is any language which arises in an unpremeditated fashion as the result of the innate facility for language possessed by the human intellect."
The answer is no, since the text did not come from human intellect but computer programming; they can't be classified as books.
Also, even assuming your definitions were reasonable, you fail to explore the possibility that his books are poetry.
I did not fail to consider the poetry route, I ignored it because poetry comes from human emotion and human feeling.
I've had a feeling for a long time that due to the predictability of humans and our processes that it this is inevitable. I think it's great that he can do this for things like instruction but if this were to get "smart" enough that would put a lot of people out of work.
I'm also confused about the super high price. Is this deliberate to avoid having to refund to very many unhappy customers?
And while these books are probably awful he's going to be known by the future people as one of the innovators of auto-generated content. At least he's not breaking spam filters with Markov chains.
It's gently odd that AI got stuck for a while; I very much hope that AI research and practice gets a bit more attention and funding.
Ahhhh.... a marketer "programmed" it. All cleared up now.
http://www.robotwisdom.com/ai/racterfaq.html
This is an interesting web page that convincingly makes the case that "The Policeman's Beard is Half Constructed" was written largely by the humans involved, with the program having a minimal role, and that the Racter program they sold to the public was incapable of reproducing the novel.
> "The Library of Babel" (Spanish: La biblioteca de Babel) is a short story by Argentine author and librarian Jorge Luis Borges (1899–1986), conceiving of a universe in the form of a vast library containing all possible 410-page books of a certain format.
definetly a hacker though
If I rely on his books which presumably have no editorial oversight, he'd better hope I don't suffer a loss as a result.
Putting this on amazon is definitely overkill for now. He should have proven the value by first generating books under his name and seeing the response.
"Beginning in March 1998, he launched a private initiative (dubbed the “K to 12 +2 project”) after directing a workshop for the World Bank which considered illiteracy. One aspect of this problem is the lack of educational materials in local languages (the smaller the language in population, the less likely the publishing industry will find it profitable to serve such communities, leaving some 1000 written languages without basic textbooks). Using automated authoring processes he pioneered, his international business publications have funded a variety of multilingual educational materials including a free online multilingual dictionary, PC games, videos and ebooks. He has applied this approach to support projects sponsored by the Bill and Melinda Gates Foundation creating thousands of factsheets (using meta analysis) on tropical plants, and is now working on automated rural radio scripts, call center materials serving smallholder farmers, and SMS content engines working with the GSM Association, the Grameen Foundation, and Farm Radio International in Kenya, Uganda, Malawi and India, among other developing counties. As a hobby, he has applied graph theory to automatically author hundreds of thousands of didactic poems (limericks, sonnets, haiku, acrostics, etc.), and is on working fiction and academic studies."
Interesting research area. Unfortunate that such laudable motivations ended up in so much ridiculous spam... Perhaps a separate category on Amazon or hosting the content on another site would be better for everyone?
Agreed, when I'm searching for books on the outlook for wooden toilet seats in China, I really want to be able to filter down to the legitimate guides only.
Want to bet? Of course they will.
http://www.jwz.org/blog/2012/07/apple-dicks/
(Oh. No. Apple had a very good reason: It was called xscreensaver and fuck you. I believe the second reason was the most important one.)
I'm going to charitably guess you're being sarcastic.
http://www.jwz.org/blog/2012/08/xscreensaver-for-ios-now-ava...
I fail to see the point of your tangential vitriol.
Am I the only one that sees the idiocy here? That jwz had to fight with them at all is the entire point.
And you can't see it as the fault of the red tape itself.
Copyright is supposed to protect an expression of an idea. I'm not sure how far you can bend the argument to say that computer reorganization of facts is expression.
For pixar and animations, the end result is an expression of artists' designs.
Can you have copyright without an expressive aspect? Is there any human expression in this output?
> Feist Publications, Inc., v. Rural Telephone Service Co., 499 U.S. 340 (1991),[1] commonly called Feist v. Rural, is an important United States Supreme Court case establishing that information alone without a minimum of original creativity cannot be protected by copyright.
http://en.wikipedia.org/wiki/Feist_v._Rural
(The specific holding was that you can't copyright the phone listings in a phone book because a simple alphabetical listing isn't creative enough.)
and Bridgeman v. Corel:
> Bridgeman Art Library v. Corel Corp., 36 F. Supp. 2d 191 (S.D.N.Y. 1999), was a decision by the United States District Court for the Southern District of New York, which ruled that exact photographic copies of public domain images could not be protected by copyright in the United States because the copies lack originality. Even if accurate reproductions require a great deal of skill, experience and effort, the key element for copyrightability under U.S. law is that copyrighted material must show sufficient originality.
http://en.wikipedia.org/wiki/Bridgeman_Art_Library_v._Corel_....
(Again, a court stated that in order to create a new copyright, the person had to exercise some creative spark; no machine alone can do that.)
> I'm not sure how far you can bend the argument to say that computer reorganization of facts is expression.
The Supreme Court seems to have said that it can't be a purely mechanical reordering. Of course, the whole point of having a judge is to create new judgements based on new fact patterns, so precedent isn't an absolute guide to future rulings.
If you are simply using public domain databases as the input, and passing them through an automated process with no creative input, then you probably won't be able to get (or more to the point, enforce) a copyright on it.
Now, whether the databases he's using are copyrighted or not is an interesting question; as well as whether database copyrights even apply. Not all jurisdictions have the notion of database copyrights. Individual facts cannot be copyrighted; but collections of them can, at least in some jurisdictions. However, I don't know if the databases he is using are copyrighted or not.
I guess it makes sense that these auto generated books may fall under the same umbrella.