To cash in on Kindle Unlimited, a cabal of authors gamed Amazon’s algorithm
theverge.com
theverge.com
Valderrama allegedly said that she organized newsletter swaps, in which authors would promote each other’s books to their respective mailing lists.
Funnily enough, I've seen this same behavior from developers, freelancers, and authors who are "internet famous" on HN and Reddit. I've signed up for quite a few newsletters from these people because they often provide useful information.
But since I signed up for so many, I started noticing that fairly often, they'd recommend each others' products. I realize it's entirely possible that these endorsements are authentic, but after a while, it starts to look like it's an interconnected web of people that met at conferences and agreed to cross-promote each each other on their respective newsletters. And so it ends up feeling like each recommendation isn't really authentic, but is instead more of a quid pro quo in exchange for a reciprocal recommendation.
And it's possible that I'm completely wrong and just imagining all of this. It's just one of my pet peeves, so perhaps I'm exaggerating how prevalent it is. And don't even get me started on developers who have resorted to hokey long-form sales copy to sell their books and services... :)
The typical person working in this niche is a solo creator who maybe splits there time creating free content and doing paid consulting projects. The business model is usually something like "write a lot of regular free content to build up a giant mailing list and then occasionally sell something to your mailing list to pay the bills".
The problem is that it takes a lot of time and effort to create good content. You are often so drained creating your "free" content that you don't have enough time or energy left to create enough "paid" content regularly enough to pay the bills. So you find friends who have created great stuff that you really like and you help each other out by cross-promoting.
In almost every case, it's a direct business relationship. The typical deal in this niche for reselling electronic content is a 50/50 split. So you help your friend sell extra copies (which is just free income to them) while you also make decent money to sustain your work.
Of course once you get "magic internet money" coming in this way, it's very tempting to abuse your audience and oversell. But then you lose subscribers quickly. Some people are very cynical about this and literally have spreadsheets to maximize how much they can sell while minimizing user attribution, but I'm not a fan of that. It just feels gross.
So yes, it is indeed a small, interconnected web of people who agreed to cross-promote each other. And sometimes the recommendation isn't authentic, but often it is. But in almost case, it's not an exchange of free recommendations, but it's literally an affiliate split between the authors.
As far as long-form sales copy to sell books, it is definitely very hokey. But the counterpoint is that most other approaches just don't convert. A cold lead who finds you via Google needs to be convinced that your stuff is worth money and long-form sales copy is one way to do that. I also hate it and completely avoid it, but I understand why people do it.
And I completely understand that long form sales copy is often what converts visitors into customers. It just doesn't convert me. But it would be totally irrational for content creators to waste time on landing pages that'll convert anti-long-form grumps like me. :)
But you are exactly right - that ad page isn't for you at all. The mailing list and free content is for you. You hopefully buy based on your relationship with the author and aren't going to be convinced by any ad copy anyway. They just hope you'll ignore or forgive them for the annoying ad copy and not unsubscribe.
That ad copy is aimed at the less-engaged customer - say a project manager at a large nameless consulting company somewhere in Germany or whatever who doesn't really care about topic X but was told by their boss to research topic X and has a training budget to buy a book on the topic (which they may or may not even read).
Notice also how every one of these pages will have an absurdly priced $800 bundle option. First, some people at companies who have training budgets will buy it because to them $800 vs. $100 is a rounding error when it's all just free money anyway. And second, having an $800 option makes the $100 option seem incredibly reasonable (when you are essentially asking someone to pay $100 to download a PDF).
They still do this with black hat SEO, since such sites aren't going to get any legitimate backlinks organically anyway.
It’s best to fix the gaming via technical means, but if you can’t, it’s still important to spell out unacceptable behavior in the terms of use. Even if it’s something you can’t enforce at scale (for instance: enforcement requires manual review by humans) you still need the policy so you can do one-off enforcements against egregious cases. Otherwise you sit around in these meetings going, “well we know they are gaming the system but we can’t do anything about it because it’s not against the rules!”
Generally speaking a business should find flaws and fix them before releasing. Iterative development is okay, but when issues like this become commonplace and nothing is done to combat it, that's when there's a problem. I have no solutions to offer.
Steam has exactly the same problem, although ar least they‘re trying to solve it in various ways.
Amazon is notorious for trying to sell you the same kind of product after you've already purchased it (e.g. I bought a ladder and now all of my recommendations are for more ladders). How hard can it be to uncover which items are more likely to be one-time buys and recommend the products people buy with those items instead (ex. show me painting supplies when I buy a ladder)?
It is hard. Dynamical systems.
My hunch is that they aren't trying to improve those algorithms very hard. But it would be fascinating if there is something about human nature that I am simply not seeing here.
First your information goes into a system. That system is used to reason about how you buy things, and how someone similar to you buys things. This happens for every individual in the system.
This eventually yields a cohesive model of behavior that can be used for probabilistic inferencing which effectively is establishing real time dynamic 'types' of both 'items being sold' and 'consumers making purchases'.
Because it's so heavily generalized to a probabilistic/stochastic model, things that shouldn't get connected can get connected. You know the age old adage, correlation is not causation? That's exactly the problem with a model that is based entirely on relational linking of items with little to no oversight in ways to explain how it's actually functioning. I want to use the word self-referential because that's very much what these systems seem to be doing, it's like it literally loops in making one too many inferences about it's predictive space. I'm not saying this is precisely why it returns your choices back to you.
But given that there's been some research (that I honestly haven't kept up with very much) that comes down to rigorously defining a data 'context' (for something like an RNN) and then rigorously searching the space for all data of what is a 'perfect match' for anything, all of this hyper perfect matching is bound to return you back to yourself in funny, ridiculous ways.
And yes, I realize the many layers of irony for those people out there who would like to point it out. I can try to explain it and I'm making all the same mistakes it does. Trying to reason about this stuff is NOT easy because it makes your brain feel broken. And I don't build RNNs or anything like that, I've done a very small amount of research on recommendation engines when it was early days, I've taken engineering classes that study probability and stochastic processes which are used to model systems where the 'types' of variables are not always easily distinguishable from one another, but they must be organized and ordered into a system that can be reasoned about rigorously. It's like building circuits for people's minds and behaviors when they in their weakest decision making spaces (for lots of people). It's sometimes screwed up but you know, you know what you like, right haha, lol. Just blame Bezos he's the one with all the money
For content, I also find that user curated lists can be far better than automated systems. They give me an upfront sense of the curator’s tastes and define niches that algorithms would never recognize
Not to mention, they probably have several hundred times the amount of data to use for recommendations since a user changes song around every three minutes. Multiply that by the number of users and average number of songs someone listens to in a session... that's going to be a big number.
Also, music choice is mostly preferential. Some preference is used when buying products but most of the time I feel like there's a need which influences the product purchase and this factors more heavily into the decision.
Netflix I would say is a good counter-example of a content provider which is based more on preference but again, the content takes longer to consume so there's likely less data than Spotify. Not to mention, it seems they've been giving more recommendation preference to their originals lately.
I bought a LaserJet printer. I went into my "Recommended for me" and the first four pages were probably 30% other printers, and 30% printer supplies (for my printers and others).
How is this still a problem?
This is the fallacy of machine learning: these systems aren't smart, they're incredibly dumb.
They also know which printer it was, when it was bought, and presumably they could know how long printers last, if not how long that specific printer lasts.
The first is certainly enough not to recommend supplies for a different printer.
> This is the fallacy of machine learning: these systems aren't smart, they're incredibly dumb.
Here, I agree. They're generally no smarter than the people programming them.. reminiscent of any computer system.
I'm not sure what you mean by "correctly," but in general, if you're talking about automation and recommendation engines, my answer is "no." Of course, I don't know the goals of the companies implementing them and whether my sense of "correct" recommendations matches theirs. I don't think any of us do.
Perhaps the answer is simple: Amazon is somehow making money from the example I cited - more money than showing the consumer relevant extras. Or they've done research to show there isn't much opportunity for more income by improving the algorithm for this.
I personally haven't had any issues finding the books I want, and I find Amazon's suggestion algorithms helpful precisely because it works as you describe. Almost all of my book purchases are made because somebody's recommended a specific book or author to me, or because I'm interested in a specific topic or author and searched for them. In both cases, I'm fine with getting generic suggestions based on the author or topic.
Edit: to be clear I think it's an improvement, just not a solution.
I'm not saying there's something intrinsically good about being contrarian, I'm just saying it feels like a whole lot of products, especially but not exclusively in tech, feel so "samey". It's all the same crap just differently branded in different colors, and I feel like algorithms are possibly to blame for that.
Going to stores in the old days used to be exciting, you'd always find something new and interesting, especially in Radioshack n such. Now it's just the same old bullshit every single time.
The more you know the less excited you get about marketing because you can actually see through the bullshit.
Not long ago, shopping for a laptop was interesting. Did you want the one with the swiveling screen, or the one that was liquid cooled? The one with the super HD screen or the one with the built-in fold-out legs to angle the keyboard better? Thick keys or thin keys? Optical or no? A novel visual design or just as-thin-as-can-be? What were the pros and cons to all this? On and on.
Now all laptops practically look identical, even something with a somewhat novel design becomes 2-3 times the price for no additional power, and all of them are designed to overheat and die within a couple of years, with the exception possibly of Macbooks, which cost a King's ransom. There's no novelty or interesting ideas, just the same old things.
No, they don't. Most of the options you point to from “not long ago" (and many more) are still available. Okay, I don't see liquid cooled or fold-out legs much, and swivel screens seem to have been displaced with the 360-degree hinges.
There was a sentence that Amazon's metric isn't influenced by big fonts or other formatting choices (they probably just use a plain text version of the text to define what a "page" is) and 'that the KENPC system (Kindle Edition Normalized Page Count) recorded pages read with “high precision”' - which doesn't actually mean that much.
Did I miss the interesting part?
My point was mostly based that I know it was that way for a long time, and that things like stuffing or those newsletter signups or competition in the back of the book wouldn't have any influence if it still wouldn't be this way.
A June blog post by the Kindle Direct Publishing Team assured authors that the KENPC system (Kindle Edition Normalized Page Count) recorded pages read with “high precision” and that the company was constantly working to improve its “fidelity.”
What makes me skeptical here is the use of "high precision" and what there could be to improve... a page was either viewed for the time necessary to read it or not. Over all readers of a book, it should not be too hard to find out the distribution of valid read times of a page.
Unless, of course, you don't have that data, but only some snapshots of which page is opened (e.g. when a sync to the cloud servers happens).
Shame the Kindle APIs are undocumented :/
The interesting question, that the article raises but does not answer, is how far Erom is demonstrating a pattern that'll be followed by other genres. There is a lot of similar self published genre fiction, but it is still a different order of magnitude to KU contemporary romance.
This is the real issue
http://twincitiesgeek.com/2017/07/24-romance-novels-for-geek...