Amazon appears to be tracking every tap on Kindle
twitter.com
twitter.com
Amazon actively touts this "Whispersync" feature in their marketing. (From the Kindle product page: "With Whispersync, switch from Kindle to the Kindle app without losing your place (requires Wi-Fi).") One would presume that Amazon achieves this by tracking whenever readers tap the screen to advance to the next page. (And having a timestamp for that tap matters for resolving merge conflicts.)
Also worth noting that in the case of Kindle Unlimited (Amazon's "Netflix for ebooks" program), authors get paid per page read. (If a person reads the first 5 pages of your book and drops it, the author gets paid less than if they read the whole thing.) One of the things that Amazon has to deal with is fraud prevention, to detect when authors are finding ways to game metrics: https://techcrunch.com/2018/06/11/notorious-kindle-unlimited...
You can't have your cake and eat it too, amazon needs to store something to make this feature work.
[0] I think Kindle on Android is the worst for this. Sometimes I don't get the position synced to my Kindle even 30min+ after leaving the Kindle app. Seems like the way to guarantee the server gets updated is to either exit the app or to return to the library.
The difference lies entirely in what WhisperSync is storing, which you can neither know nor control.
Of course this is impossible to do in fixed formats like PDF, but 4 years back I specifically worked in Kindle content to make PDF books reflowable :)
>KDP [Kindle Direct Publishing, Amazon's self-publishing platform] pays authors for both paid downloads as well as for pages read and it doesn’t sense reading speed, just the highest number of pages reached. ...
>The way that the book-stuffing con works is that scammers stuff lots of extra content into an ebook before uploading it to Kindle Unlimited, and then trick readers into jumping to the end of the book.
>Thanks to a flaw in the Kindle platform, namely that the platform knows your location in a book but not how many pages you have actually read, the scammers can get paid for a user having “read” a book in Kindle Unlimited by getting the user to jump to the last page. ...
>Interestingly, the flip-to-end scam doesn’t quite work on newer Kindles but still works on older, non-updated Kindles which makes it still a lucrative scam.
https://techcrunch.com/2018/06/11/notorious-kindle-unlimited...
I don't like this anyway. If I buy a movie from Amazon Prime and only watch part of it, do I get a partial refund? Seems like they are shafting authors.
For similar reasons, the quality and genre of made-for-TV movies currently airing on the Hallmark channel is probably different from what you could experience for the cost of a ticket at your local movie theater: one of these distribution channels caters to people who want to consume several hours of content every day at no marginal cost beyond the monthly subscription that is already part of their budget, while the other caters to people who are willing to pay $15 to spend 2 hours watching a film that a studio spent millions of dollars marketing to them.
Say a user reads the first page of 100 different books—they can do that, because opening each new book has zero marginal cost for the user. Does Amazon now need to pay authors for 100 books?
I don't know how Netflix payouts work but I have to imagine that viewer time is taken into account.
Because Netflix pays for their content up front, they have to take a bit of a gamble. (Maybe they spend a bunch of money for a new Coen Brothers film, but nobody watches it, so they take a loss on that project. Or, as was the case in 2008, maybe TV networks grossly under-estimate the value of their catalog of old shows, so Netflix gets to pay peanuts for the content that serves as the bread and butter.)
My understanding is that they tally that up and pay each of the authors at the end of the month with a single payment of the KU revenue, directly paid for books, etc. As each book can be bought multiple ways. Though note this only applies for self-published books as otherwise, it all goes back to the publisher.
The money that people spend on $10/month KU subscriptions is used to pay authors based on which authors people spent the most time reading (or, more accurately, which books you read the most pages of). If I read 400 pages of book A, and 5 pages of book B, then author A gets paid more than author B. I think the reasoning behind this should be pretty intuitive and obvious. Since every KU subscriber only spends $10/mo regardless of how much they read, there is a fixed "pie" to distribute to authors, and it makes sense to divide the pie based on which authors contributed the most to the readers' use of the KU platform.
Readers don't get a "refund" for dropping a KU book 5% of the way through, because even if you quit reading one book, the fact that you stopped reading a book does not change the fact that you still have access to tens of thousands of other ebooks in the KU library for the remainder of that month, which is the thing that you are ostensibly paying for. (I don't phone Netflix to request a partial refund if I start watching the first episode of Bojack Horseman and quit halfway through the first episode, I just start watching Stranger Things or Narcos instead.)
If authors don't like this arrangement, they are free to not participate in Kindle Unlimited, and sell their books under a more traditional model (where the author sets a price, and people can buy the book for that price irrespective of any participation in any sort of subscription program).
>Readers don't get a "refund" for dropping a KU book 5% of the way through,
But...
Authors get paid by the page?
So Amazon basically gets to stiff authors on the refund that customers aren't getting but Amazon is applying internally to products customers use through their services by way of just not paying authors for content? Seems pretty fucked up to me and the only one that benefits is Amazon. Customers are left with something they don't want and authors aren't paid for their work while Amazon keeps the change...
Lets say user pays $10/month fee. Amazon decides it wants to keep $2 and distribute $8 to authors. Now, there are different ways of accomplishing that:
(i) Distribute proportionately based on which books user downloaded. If user downloaded N books during the month, author of each of the books gets $8/N.
(ii) Distribute proportionately based on amount of time spend by user on each book. If user spent T hours in Kindle during the month, and T_1 time on one of the books, then the author of that book gets $8*T_1/ T.
There are pros and cons of both approaches in terms of fraud prevention, user engagement etc.
Presumably when they stop reading Book A after the 5% mark, they would move on to Book B and Amazon will then pay the author of book B. So Amazon is paying someone for the whole duration that the customer is reading from their collection.
Personally I think it's a fair arrangement. Some of the "books" are really low effort cash grab that you'd literally open, read 3 pages and drop - it'd be unfair if they got paid just as much as well written works of the same length that you finish reading through.
That seems illegal.
If someone pays $10/mo for access to a library of tens of thousands of books but doesn't actually read any of them, then Amazon gets income without having to pay royalties, in the same way that Netflix still gets your money if you subscribe but don't watch anything. This is true even if you decided to subscribe to KU/Netflix because "Oh, I should get around to reading Harry Potter/watching Stranger Things" and then don't get around to actually reading Harry Potter or watching Stranger Things. This is very much legal.
But does Netflix pay partial royalties on films, based on percentage viewed?
For example if user reads one page of one book this month this author should get 7usd.
Medium works the same way except with reading time and maybe claps.
This at least is the most logical, fraudfree and fair way to do this.
The "authors get paid per page read" model is only for Kindle Unlimited, not for regular Kindle ebook sales. When you buy a Kindle book for $6.99 or whatever, Amazon sends the money directly along to the author (or their publisher) after taking their cut, just like you'd expect.
But if you pay $10 a month for a Kindle Unlimited subscription, and read a dozen books by different authors, Amazon has to figure out how to split that fixed monthly subscription fee between all the authors that you read; paying authors based on page reads seems like the best way for your KU money to go to the authors/books that you actually read.
If I had a subscription with a book store that offered me N amount of books a month for a fee, the book store would still need to buy copies of said books from the publisher, who would pay the author whatever was worked out in their contract per sale, whether or not I read one page or the entire book. How is Amazon's model different than that?
It used to be that if Blockbuster wanted to rent out Raiders of the Lost Ark to six different customers simultaneously, they needed to own 6 VHS copies of Raiders of the Lost Ark. Now, with streaming, Netflix doesn't have a finite number of "copies" that they can lend out at a time; if every single Netflix subscriber in the country decides that they want to start watching Indiana Jones right now, the only thing preventing Netflix from providing that is their bandwidth, because Netflix has worked out an arrangement with Paramount Pictures that allows them to do this.
Likewise, Amazon has an arrangement with KDP authors that says, "We lend your ebook out to as many KU subscribers as want it. At the end of the month we pay you based on how much people read your books." If an author doesn't like the terms of the KU program for any other reason, they are free to decline the subscription model and sell their ebooks through the regular "customer pays fixed price for ebook, I get money from sale" model. In fact, most authors don't opt into KU; there are a millions of Kindle books, and only tens of thousands of books in the KU program.
Because you pay a fixed subscription fee and the money gets divided between all the authors you've read, it's effectively zero sum: if Amazon wants to give more money to authors who wrote books that people dropped after the 1st chapter, that means less money for the authors who wrote books that people actually liked enough to read past the first chapter. Amazon has structured their program to reward authors for writing books that people consume more of, which seems like a good way of rewarding creators based on the value that they contribute to the platform. If you don't like it, don't opt in and instead sell your books the "normal" way.
They still offer this service.
When was the last time you saw a sustainable private library (i.e., funded by membership fees, not by taxes)? Were the fees $10/person/month?
Does Amazon respect you and turn off data collection?
Or better yet - does it ask you before turning it on?
Jokes aside, i often find the excitement around "tracking" to be frustrating. Like fears of global conspiracies and mind control, the idea that somewhere a companies engineers are using this data to track your inner most thoughts is crazy in practice. Instead they're using it as a way to diagnose a bug when your ebook crashes or as a way to figure out how to make sure you're not getting stuck in a poorly design UI.
It can, for example, be that out of 1,000 companies tracking your every click, your every character typed, and your every web site visited, 999 are doing it for better bug tracking or feature development.
But the 1,000th company is Facebook.
So it feels reasonable for people to ask, "What are you tracking, how is it used, and how can I be confident it won't be abused?"
Because of the nature of computer security, you cannot be certain that any data that you give to a 3rd party will be secure. Once it leaves your machine, it's out of your control, forever. Maybe an initially privacy friendly policy vanishes once the company is bought out, or a data breach occurs and the data is suddenly indefinitely in the public domain.
This doesn't just apply to Facebook, it applies to every single company storing data. The only way to prevent is to have a data retention policy and invest heavily in security, which pretty much nobody does except for the really big players.
I do this mostly to save battery life, but also so it doesn't track (but of course it may save up these logs to transmit when I do turn the wifi/cell back on).
The Kindle at least keeps track of my last read position in all my books. Foxit only tracks for the last two pdfs. Ditto for every other reader I've tried. I've had to resort to keeping a separate note of my place in the book.
It absolutely does save up those logs. I worked on some of the code which saves up the advertisement view metrics. Admittedly that was more than 2 years ago, but knowing the team we handed that off to, I would be astonished if that's changed.
But why does it need to be calculated on Amazon's servers? AFAIK Kindles are running a linux kernel with a lot of busybox, and calculating a running average doesn't appear (to me) to be a particularly difficult calculation.
Perhaps it can be argued that this calculation uses battery, but so does sending all of this telemetry to el Amazon.
What I'm saying here is that I think that we shouldn't concede privacy in return for convenient little UI widgets, especially when the computing power is available, cheaply, locally.
Also, slightly OT: has Amazon ever said whether that reading time is calculated locally or on their servers?
The kindle tracks progress via the actual amount of text read instead of pages, so the screen and text size should be irrelevant. It can still be switched to display the page number, but that is also independent of the amount of text on the screen(meaning it doesn't necessarily increase with every swipe to the next "page").
This doesn't seem to me like what the tweet is describing, wherein the kindle is registering every tap instead of distinct "x words/chars progress made".
If they were capturing and storing "X words in fiction read in n seconds", I could understand it, but they're not: they're registering every tap. I'd be interested to see how this matches up with "userChangedTextSizeToBlah" data if this is how they're calculating reading speed.
"This is useful for legitimate reasons and is an industry standard practice --- even though it has rife potential for abuse, too."
It doesn't affect the utility of the trade-off, but all else being equal we usually accept what exists now. Fighting what already defacto exists is hard-- either as lone consumers, where we're tilting at windmills to little effects... or as regulators, where we risk unintended consequences.
> "Kindle tracks every page turn in order to determine how long you have left in the book" would be a statement of the trade-off
Nah, Kindle tracks all these actions so that developers can improve ux and understand how the device is being used. Wonder how often people turn pages by mistake? Look for quick pairs of page forward with page backward. Then maybe you can think about touch sensitivity. How much will be people be annoyed if you remove the buttons? See what proportion of users exclusively, mostly, or don't use buttons.
> If it's wrong for one person to do it, it's wrong for everyone, even if all people are doing it.
It's wrong for anyone to abuse the information. If we have an industry of people using the information ethically, and then one bad actor misuses the information, we need to consider that background. Do we take measures solely against misuse, or do we attempt to stop the collection to have more certainty in stopping the misuse?
You could have addressed that instead of parroting the same point. :P
If everyone pushed their customers off a bridge should Amazon, too?
This data doesn't only have a very limited utility. It's more limited than I thought it was from the title, but still pretty powerful stuff.
So good call holding your data close here.
Edit: Exited the interview half way through, would never work there
The only defense against that and attacks that become possible in the future is to not share any data with the attackers. Amazon, Google and the government are obvious attackers in this scenario.
That's not to say it hasn't started to be abused but at the time it was completely from a "customer/UX first" stance.
Generally every product I am aware of tracks interaction based data such as where someone clicks or taps and what context they are in. Consider things like `utm` parameters which suffix most links people click to determine the context they clicked on something and what they clicked on.
I do not see this as sinister. I imagine somewhere in settings one can turn this feature off but I don't know for sure.
* Disclaimer: I am currently employed at a subsidiary of Amazon. These views are my own.
Would you need to log every page turn with every book and time and date for something like that though? Wouldn't that be a more specific event like "turns page forward, turns page back within x seconds"? This sounds more like "we don't know what we might use this data for, but it's better to have it and not need it than to need it and not have it ... who knows, maybe we can deduct some profile from knowing how quick the user read through that chapter in that book" than legitimate use cases.
From what I've seen tracking simple events and then piecing them together en masse tends to show up significantly more frequently.
Given the very private nature of the data ("he read Marx and Mao, and read some sections carefully!"), vacuuming up as much as possible doesn't sound like a good idea. Add to that the almost chronic inability of large corporations to protect data, they really should start treating data collection as a liability rather than an opportunity.
In this age of Big Data, when it's relatively inexpensive to weave together a large number of small data points to come up with an overall profile that is truly invasive, I have to consider every byte that is sent to be a risk, and to be avoided when at all possible.
There's also the question of autonomy. I actively resent data collection without my informed consent. Companies that do this are, in my opinion, being abusive and infringing on my right to autonomy and, to a degree, to have control over aspects of my existence that matter to me.
Librarians will never give out your checkout history to anybody without a national security letter or a warrant. Meanwhile, Amazon is probably passing your reading records around to 50 different analysts and storing them in databases where dozens or hundreds of engineers have access. When I go to amazon.com, I see recommendations for other books similar to those that I've read, including those that I didn't buy through their store.
You could always just get all your reading material from Library Genesis instead. I have owned several Kindle models over the years, but I have never had to interact with Amazon. I put the device in airplane mode the moment I took it out of the box and kept it that way, and I have downloaded all my reading material from LibGen (or a publisher who provides DRM-free ebook files) and moved it to the device via USB.
Which wouldn’t be as convenient, but you could buy books from Amazon, and then use Calibre to strip DRM and send to kindle over USB.
Brought it home, did a factory reset, and only push DRM free books or books after stripping DRM.
You shouldn't have to break the law to protect your privacy.
I'm wondering if this person on twitter had any kind of usage-data opt-out that was on their kindle and they just hadn't opted out. I haven't looked at my kindle's settings in a long while, but I'll take a look when I get home tonight.
It gives insight into how things are used and what's used most often etc. Granted this isn't necessarily directly applicable to this case with the kindle, but similar in concept.
Users are not your experimental group. This attitude should have died years ago, latest with the GDPR.
One is optimizing/improving experience based on anonymous data, the other is building user profiles for targeted ads.
Mapbox is a fantastic example of mass data aggregation of users that has been anonymized.
EDIT: Why should we do clinical studies on medicines when that could be invasive to a persons privacy collecting such personal information? Is it necessarily wrong? Tools can be used for good and evil, that's the problem here, not the tool itself.
Clinical trials require informed consent and institutional review. Regulating software telemetry like other human subject research would make privacy advocates very happy.
I explicitly said what I'm talking about isn't necessarily directly applicable to amazon kindle. I also agreed with you regarding software telemetry being used responsibly and irresponsibly.
So I'll repeat I don't see where we disagree.
I'd not worry about tracking if the experience is good enough. Please someone make sane controls for these devices again
Gigawatts of electricity are used to run sophisticated neural nets, advanced data pipelines built to funnel all that data to the mother ship and thousands of dev hours went into this elaborate tracking mechanism. You need that of course, "to improve UX". I do not have a problem with that for such a device, per se. Fair enough, go real overkill on your "user research" then.
And yet, so many UX aspect of a Kindle are just plain bad. Problems which have varying degrees of complexity, to be fair. But some of them are so trivial yet impactful that it is hard to imagine they would escape the attention of a single UX engineer worth his salt looking at a Kindle for two days. Some examples:
* I have a cheap Kindle, which means the lockscreen shows ads. They are so hilariously anti-personalized I can't even. Like, you have all this Big Data and you think I will ever buy a run-of-the-mill cliche romance?
* The recommendation system itself. When you "start out" on your Kindle, your recommendations are literally just every single other book the few authors you read have ever written.
* Having airplane mode off seems to drain the battery massively even if I am not connected to WiFi/Bluetooth nor actively attempting to
* One can view interesting usage statistics, but only if you declare yourself as your own child and activate a password based content lock
* I have to manually flip through sometimes dozens of pages of imprint, one-line-per-page copyright notices and so on until the Kindle realizes I have indeed "read" the book so it stops displaying it at the top at "99% progress"
What I want to say: Do all the Edge AI and IOT and Orwellian Surveillance for all I care. But maybe fix the boring old low hanging fruits first?
I think you can technically unregister a Kindle and use it only with sideloaded books.
I'm not sure, but Onyx (http://onyxboox.com/) probably don't track.
I do have the Alexa app installed on my phone so.
I can't think how many times on the product side of services I've run that we've been grateful to have this kind of data. Sometimes it's for obvious reasons that we always knew, but often it's for suddenly crucial reasons that we never could have anticipated when we first started collecting the data.
Again, that doesn't mean it's OK and I strongly support GDPR-style data privacy legislation in the US. But in the meantime I guarantee you that just about every service you use is gathering data like this and a whole bunch more.