Kindle, ePub, and Amazon’s love of reinventing wheels
hackaday.com
hackaday.com
Amazon has a culture of promotion focused “engineering”. This rings true across product management and software development.
Rather than solving real customer facing problems, many organizations including Kindle, emphasize “what do I need to do to get promoted?”
This is not only SDEs, but L7/Senior engineering managers and L8/directors, who latch onto a sales pitch and put all their eggs in one basket.
The template is like this:
> I’m going to build a new framework, platform, format, or some other “thing” that everyone should now use to solve {{problem}}. Then, I can claim a larger scope and significant impact, no matter how pointless, costly, or annoying this thing is to customers.
They write documents. These are narratives where they spend hours and hours reviewing and nit picking how “crisp” a sentence is, or how there are too many commas, or some other superfluous bullshit.
There’s an OP1 document, where an L7 writes a fancy narrative, justify more head count and growing their organization. It’s their ticket to a director promotion.
If you’re a developer, you flock to the same project. For L4s and L5s, it’s your ticket to the next level. For an L6, you spend your time selling the absolute hell out of this thing, going on a road show and pitching your idea to get buy in from other teams. You’re careful not to write much of the code and mainly come off as being the architect and consultant. This is your ticket to getting to the Principal level. If this is someone else’s idea, you’ll do whatever it takes to take credit when this succeeds but blame everyone else when it fails.
Best case, the thing you deliver gets you promoted. Then, you simply just move onto to another organization. You were in Kindle and this new piece of shit has made on call hell? Oh well, you already got your promotion. Now you can just move on to AWS (or find another team in any of Amazons legacy businesses where you can coast).
Kindle was an army of H1Bs with an extremely toxic culture. They refuse to hire or even interview white, black, or Hispanic candidates. You can pay H1Bs less, and they’re happy with it because they get to live in Redmond, WA instead of India.
I could go on and on, but I feel anxiety right now even thinking about my time working there. Absolute shithole.
Just to rant about both companies, Amazon Prime's Chromecast support seems to be currently bugged out. It plays for a few seconds then gives up with a content not available error. Maybe it will work again in a year?
This is complete BS, the org has as many white developers as asian and indian and even opened a big branch in Madrid
> You can pay H1Bs less, and they’re happy with it because they get to live in Redmond, WA instead of India.
Salary is a _direct_ computation from performance ratings which are the same % for every org. For H1Bs to be paid lower, they need to be systemically rated lower by the same managers who "refuse to hire white, black, hispanic" candidates.
So if that was the _only_ anecdote that made you rethink, hopefully this is a second anecdote that helps you make a choice!
(side note: If you dislike working with Indians it can be a terrible place, a good number of her colleagues have been Indians (she is one as well).)
EPUB has a lot of features. A lot. I don't think anything (except maybe Adobe Digital Edition) support everything. I actually just realised yesterday that Moon Reader+, one of the best epub reader for Android, doesn't support vertical-rl layout (mostly used in Japanese book).
Even Kobo doesn't use EPUB as-is. Its native format is KEPUB, which has tons of id-tagged <span> inside the content. Moreover, KEPUB and EPUB on Kobo use different renderer (EPUB use RMSDK, KEPUB use Kobo's own) with different supported features. Isn't this a mess?
The new KFX format has one major advantage over EPUB. The images can be stored out-of-band. This allows you to start reading image-heavy book quickly, while images continue to load in the background. For fixed layout magazine/comic/manga where the entire file can be 100MB or more, this helps a lot.
Not saying there might not be any issues with EPUB but the fact that in 10 years they can't be bothered to add at least a basic reader for me just mean they want to keep this incompatibility and keep people in this separated ecosystem with amazon/Kindle. I believe this choice is anything but a technical one.
Since I never had much trouble using Calibre it means I'll probably never buy a Kindle just because of it but many people can't be hassled to do that.
If you're going with just the very basic stuff: text, maybe a few images - it really doesn't matter what kind of bare-bones reader you use.
BUT when you're opening a PDF that actually uses the features it's a different story. You've got a file with umpteen layers of information that needs to be displayed in a very specific way (I've seen blueprints do this for example). Most bare-bones PDF viewers either choke completely or display everything wrong.
Which I completely believe, PDF support on readers is always an issue (there is even a discussion on it in this thread).
Then it support the position that if they added support for PDF but not a basic EPUB then this choice was not made for technical reason.
1) Why Amazon keep re-inventing the wheel?
As I explained in my original comment, KFX does have technical advantages over EPUB.
2) Why doesn't Kindle support EPUB natively?
Probably because there isn't a need for one, as they are using KFX anyway. Why invest time in something that is only for people who aren't using your platform?
Kindle is out since 2007, KFX was released in 2015, hopefully when you create your own format there is some advantages.
In the end it appears we agree that the position of Amazon is : "Why invest time in something that is only for people who aren't using your platform?" But we probably disagree on the answer to it :D.
Other online media, notably streaming services and Xbox, DRM everything. I don't think Amazon is an outlier here.
EPUB is just a ZIP file. It's perfectly feasible to read the HTML/CSS and image portions separately.
Not sure if that is strictly easier than just new format (which allow you to add more features you wanted too)
https://kavitareader.com for those interested.
Amazon backend accept EPUB (either from KDP or from Send to Kindle). Amazon then process EPUB into Amazon own format, and deliver that to Kindle device. Which sounds a lot like what you described.
Btw your project is interesting. Does it support vertical-rl EPUB? After finding out Moon Reader+ doesn't it make me doubting everything now.
The biggest downside of Kobo for me has been Rakuten's update cadence. The system updates opt-out doesn't actually work from what I can tell and my device only comes out of airplane mode once every few months. Every damn time I do so it has to download an update, lagging out for a long time, and changing the UI in exciting new ways.
Calibre is king, set things right and you can just email books to your Kindle.
Works perfectly.
I don't want to send a copy of "And the Band Played On" (for example) and start getting Chick Tracts.
The MobileRead forums are where most of the unofficial discussion/development happens: https://www.mobileread.com/forums/forumdisplay.php?f=223
That's good to hear. I have no problem with a device vendor focusing on just making sure their device works well, as long as they're not trying to close it off and actively working against external developers. I'd heard stuff like Kindle blows physical fuses so that one can't install older firmware and stuff like that so they actively prevent users from having control of their device. I have no desire to be a customer of such types of vendors.
[0]: https://www.amazon.com/Pocketbook-Touch-HD-Metallic-Grey/dp/...
Having used it, I don't think I could go back to Kobo's reader (and even less to Amazon's, which is a joke in term of configuration options).
Oh, and not to mention, how easier... is to just poke things? At least with Clara HD, don't have to fear about bricking it, because worst case scenario, I can just reflash the system image, and have it working again.
And few days ago I just went and replaced internal microSD card with a bigger one (First making image of original one, then flashing it to bigger, and extending partition) and it just worked.
Sadly or not, I like to hoard things locally, on the reader itself. And that includes manga. So yeah, being able to replace microSD Card was very convenient.
And the device itself is a tiny bit more responsive overall too, alongside better transfer speeds :D
I bought a 2" thick sci-fi book a few years ago, "Pandora's Star". It seemed pretty interesting, but way too much book to carry around. I got a kobo and put what I vaguely recall was named "koreader" on it. It was pretty nice. It stopped working after 6 months, though. I took it apart, and at least one power management IC is shorted out. Its e-ink display permanently says, "Sleeping..." It's been a paperweight on my nightstand for about 2 years, and a minor source of anxiety for me. I don't plan on buying another kobo device any time soon because I hate e-waste.
My wife got me a kindle paperwhite to replace it, and it didn't take much work to get the same book onto it. I think I'm maybe 25% of the way through the book, but I can't really say because I don't understand what the numbers in the lower corner mean. I also printed the Rust programming language manual to PDF and managed to copy it onto the kindle without much fuss.
If you tap on the numbers in the lower corner, you can cycle through various options including % read. My personal preference is for time left in chapter.
The interesting question is why if the systems support ePub, Amazon doesn't update the "Send to Kindle" tools to also support the extension? Those tools support .doc, .txt, .pdf, and other random things, just explicitly not .epub. It could just be that those tools are low maintenance and don't update often in general. Maybe it's simply as sinister as they don't want people to know kindles support ePubs, even though they do (because they have to) and the "Send to Kindle" tools are the primary interface people use for "what does my kindle support?" as very few users bother to directly USB their kindles and rely on wifi updates.
It's strange.
My only complaint is no equivalent of Send to Kindle for Kobo, so I always have to plug the device in if I want to copy an EPUB to it.
If you compare the size of an e-reader to actual published books, you'll find that the smallest devices are roughly the size of a 4x6 index card. Most trade books measure about 9" diagonally, and in my experience, you'd want a 9" -- 10" device for most reading, possibly a 6" -- 8" minimum. Given that an actual device loses some area to bezels and controls, the translation isn't direct, and you're suffering either larger physical device dimensions or smaller screen size.
I purchased a 13.3" device specifically for use with textbooks and scans of small-font, multi-column articles. It's a bit on the large size for standard text (though can be read landscape-mode in 2-up format), but the ability to read almost anything without requiring in-page zoom and navigation (which is well-supported) is a major benefit.
The device has also convinced me that virtually all the supposed disadvantages of PDF or similar fixed-dimension document formats (e.g., DJVU) is not the document format but the display properties.
I'd recommend:
- Typical fiction or nonfiction text: 6" minimmum, 8" preferred.
- Technical books, some tables / diagrams: 8" minimum, 10" preferred.
- Textbooks, technical articles, old scans: 10" minimum, 13" preferred.
Comic and manga fans will also probably prefer 10" devices.
Then there's colour, though AFAIU size formats are more limited (8" and/or 10"). The fidelity is low --- more a low-saturation 1950's-era palette, though with much better resolution. Colour cuts resolution by a factor of three, so 220--300 dpi (typical of current devices) falls to about 75--100 dpi. I've not experienced this directly, so can't say what the net appearance is.
I wouldn't mind colour for some articles, mostly for graphics and visuals where colour differentiates values being presented. Though I'm rather enjoying my desaturated B&W world.
Kindles need to be jailbroken to use koreader, and Amazon apparantly makes that harder with each device :-/
I have this as a right-click action for files matching *.pdf in my file browser (thunar):
k2pdfopt -ui- -dev kp3 %F
So if I right-click on a pdf I can immediately reflow it into a kindle-readable version.(%F is the quoted file name, in your terminal you would `k2pdfopt -ui- -dev kp3 'something.pdf'`)
Alternative phrasing: Amazon is continuing to fix security bugs in Kindle device.
For stuff you read from start to beginning the Kindle (and other eInk readers) are amazing.
For books you need to jump around in it's just painful. Use an iPad or something else with a smoother UI and display.
A general rule of thumb when deciding whether to purchase an e-book: If the physical book is larger than a small paperback size, then it is unlikely to be suitable for an e-ink screen (unless the book is text-only).
Small e-ink readers are not suitable for any books with complex layouts, layouts designed for larger books, or books with colour graphics: charts, diagrams, photos, etc. Amazon, in particular, encourage publishers to convert as many books as possible to Kindle format regardless of whether those books are suitable for e-ink screens.
Then just get a bigger e-reader.
8" Kobo Sage is the largest I still find comfortable holding in hands for longer periods of time:
https://us.kobobooks.com/products/kobo-sage
10" Kobo Elipsa should be even better for PDFs, but I would only use it with a stand:
honestly, I am a bit perplexed by your wording. there are no issues, you are simply using a very very very small screen. just buy a big e-ink reader like the Kobo Elipsa
I'm trying to understand your meaning of "format that relies on scrolling". How can this be the case since pdf was designed for printers?
"A PDF file is actually a PostScript file which has already been interpreted by a RIP and made into clearly defined objects.
Postscript is a file format supported by almost all high-end printers and many business-class laser printers. With these printers, you can simply send a postscript file to them over USB—no drivers required—and they’ll print it perfectly.
"
Files meant for screen display used to be standardized around a 72dpi resolution. Any detail not visible at the normal viewing distance is lost unless you are able to re-render the file at a higher resolution a.k.a. zooming in. And once you start doing it, scrolling becomes necessary in order to read through the document since the formatting is fixed or as you say "made into clearly defined objects". And this often works very poorly due to the inherent limitations of eink display when it comes to frequent screen refreshes.
Nowadays most displays, big or small, can do better than 72dpi but they still have a long way to go before they could match printed material.
Small screen e-readers get around the problem by aggressively reflowing the text of books with larger, sharper font and more line spacing to improve readability . Yet non of these techniques really work for PDF. Kindle DX failed rather miserably in the textbook and technical document market because of its limited resolution. I do not own the Kobo Elipsa, however I have spent some time with a couple of Dasung eink displays of a similar pixel density. And my current personal verdict is these panels are still not good enough for reading pdf files.
Shop around. You should be able to load up a test PDF from the web on demo units at stores like Best Buy, etc.
Is the screen on the Fujitsu lit?
Does the note taking application index your text so you can search? Is it easy to get the notes off? Can you highlight and annotate books and then easily export your annotations and highlights?
Although the font could be better, prettier maybe. So much better than Kindle for daily use.
I like the TeX Gyre Pagella version of Palatino, myself.
Kindles are actively walled to force customers into the amazon ecosystem.
I could just plug this sony e-reader to my PC and copy whatever I wanted to read. The e-ink panels have gotten substantially better but the experience is now objectively worse with Amazon and more expensive, most customers not don't realize what it could be.
There's also Gutenberg, Archive.org, and numerous other public-domain e-book collections such as StandardEbooks (https://standardebooks.org/)
https://news.ycombinator.com/item?id=31220553
I gave it a shot and it works well although instructions wern't obvious at first.
Here's a clear how-to guide on Jailbreaking your Kindle https://www.mobileread.com/forums/showthread.php?t=346037&hi...
The article ends with a discussion of Calibre. It would also be very interesting to know how much Calibre is used with pirated books.
Amazon must be following an interesting dance with copyright, piracy and the Kindle. They know if it gets too out of hand that's the end. But if they were to entirely lock down the Kindle to anything without DRM they would also know that many people would shift devices.
The objective of antipiracy restrictions shouldn't be to restrict the top 5% of most technical power users. That takes far too much time and effort, and frankly power users provide word of mouth recommendations.
The objective of antipiracy restrictions is to stop 95% of users going to the trouble of piracy.
This obviously feels very sketchy, not only because you might be including Amazon in some very transparent piracy, but because it works TOO well. By which I mean, these files aren't simply added to your device, but to your Amazon library; these are then readily available OTA to all devices you own and any you add later.
I wonder if the difference is that with the exception of comic books (sigh), most ebooks are small enough that they download quickly, yet have good metadata and often excellent cover art, so they replicate quickly. And they rarely get buried inside of unnecessary folders the way other files do.
It'd be pretty awesome if say, you could visit Humble Bundle and after buying a DRM-free collection of ebooks, just hit a button to transfer the entire bundle to a Google Drive folder or iCloud Drive folder without having to download and upload things. That's still maybe a missing thing. But it's not terrible to have to send files manually. It's just extra work. (How lazy can I get, apparently?)
I don't mind paying $15 for a book if the author(s) are paid well.
I don't see how Amazon could know for sure. I have never uploaded a pirated book to my Kindles, but have uploaded dozens of books via Calibre, either public library books that aren't available in native Kindle format for some reason, or obtained elsewhere (notably Tor.com's regular offers of free ebooks).
Perhaps because it is relatively simple, this has been one of the longest-lasting and best-aging devices I've ever owned.
Otherwise, I much prefer real books. I find them more comfortable to read, more pleasant to read (I get a better feeling turning pages and looking at paper than pressing buttons and looking at a screen) and I can pass them to my sons.
Also, for children interacting with a real book is a better experience.
And for the rest of us Calibre converts anything in 30 seconds.
The reason you fight monopolies, is because monopolies cannot last forever. There are all sorts of bad consequences with having digital publishing owned by a trillion dollar monopoly.
In order to have a free press at all, which is important in democracy, which is important for society to stay resilient, we need a diverse ecosystem of digital publishing options.
Calibre, with the DeDRM tools installed, is unreliable for Amazon's newest books. There are workarounds that involve getting the file in an older format, but then you lose the typography improvements that are only offered in KFX files.
It feels like Amazon is finally starting to close the DeDRM hole.
I read all my Kindle books on it, it's entirely replaced my paper notebooks for note taking, and all my technical books, manuals and so on have been excellent on it.
Your intuition that this form factor is better for technical PDFs is correct, I think.
Past that, most books aren't available for DRM-free purchase. Breaking the DRM on an otherwise legally-purchased book is the next easiest thing. DeDRM tools are just a search away!
For free epub ebooks, there's overdrive and the library ecosystem; legal libraries also DRM-wrap their epubs.
This is the one case where I'm OK with DRM -- you're receiving the library's book on loan (it's never your book), and they need a way to enforce the length of the loan.
ebooks.com
edit: actually, every something I .com is pretty good. Haven't done that since the early days.
I really liked the wireless mechanism of pushing books to my kindle, is there a similar feature for koreader / remarkable / Kobo / anything else?
> Beginning in late 2022, you will be able to send EPUB (.EPUB) documents to your Kindle library from the Kindle app.
Before this point, it was never possible to send e-pub documents via Send to Kindle.
With that being said, I have an Oasis and keep it in airplane mode 99% of the time.
It's probably the one "killer" feature for me.
That image brings to mind: How much space would you need to store every out-of-copyright book written in English?
As of last year, someone on reddit determined that l--g-- had 2.7 TB of fiction and 40 TB of nonfiction, and double that size of scientific articles and magazines.
It's missing a fair amount of books, which would make the true total size larger, but it also has a lot of duplicates and scans that are unnecessarily large compared to true ebooks, which would make the true total size smaller.
This is much later historically, but in medieval conception 2 identical texts could be worth vastly different amounts based on the quality of the binding and the illumination in the book.
This is an example of why applying modern thinking in a historical concept can lead us astray. The Library of Alexandria wasn't pro or anti copyright; that makes as much sense as saying they were pro x86 and anti ARM. It just wasn't even a concept that had been developed.
Edit: upon rereading, my tone seemed a bit combative. Not trying to call anyone out, just wanted to point out that applying modern sensibilities to history can cause distortions. It is a pet peeve of mine that people consider ancient people's stupid, and this is one of the major reasons why. We forget that we have the benefit of thousands of generations of previous people improving their understanding of the world, and what seems obvious to us is only obvious in hindsight with the benefit of that prior knowledge.
By modern standards, there's hardly anything more anti-copyright than forcing someone to let you copy their personal library. It disrespects not only the modern rent-extraction copyright by the author, but also the idea that written material in any form is anyone's property at all.
zstandard via the ZIM [0] format.
Honestly I wouldn’t be surprised if the modern version is just two SD cards. At least if you only want the text.
Bowker, the US issuer of ISBNs, has registered roughly 300k new titles annually, since at least the 1950s, and beginning in the aughts, about 1--2 million "nontraditional" (self-published) titles.
In straight text, assuming 250 pages of 250 words at 8 bytes/word, 40 million books works out to about 20 TB. This is a likely close to the lower bound as there are texts with image and other content which bump this value up. It works out to about 0.5 MB/book, which in my experience is too low.
At 150 million books, we'd get 75 TB, and at 200 million, 100 TB. Nice round numbers....
A more realistic per-book size is 5 MB. This is typical of a mostly-text, formatted PDF. At 40, 150, and 200 million texts: 250, 750, and 1,000 TB.
A data-heavy / graphics-heavy book might weigh in at 20--100 MB, though those tend to be less common.
Scanned-in books with both page images and OCR text typically run about 5--10 MB. The largest I've seen (a copy of Lyell's Geography was >300 MB.
Apple has offered iPads with > 1 TB onboard storage for at least several years. This seems to be the higher end of available tablet storage. (For some reason, the Cloud-and-Surveillance-oriented Google skimps with many devices still at 16--32 MB. Remarkable, which also went Cloud, tops out at 16 MB.) If fully devoted to texts at 5 MB each, that would offer storage for up to 200,000 documents, roughly the size of a well-stocked municipal public library.
As to how many public-domain works might exist, the 1927 Report of the Librarian of Congress notes that the Library had assigned 4,503,164 copyrights over the preceding 57 years (presumably 1869--1926), the time in which such registrations were the work of the Library. Those works should all presently be in the Public Domain, and would total 22.5 TB at 5 MB/work.
Annual report of the Librarian of Congress yr.1927, Washington, Library of Congress; for sale by the Supt. of Docs., U.S. Govt. Print. Off., p. 22.
https://babel.hathitrust.org/cgi/pt?id=uc1.a0013520523&view=...
The Library of Congress's total holdings was 3,420,345 million as of end of FY 1926 (p. 23 ibid). Presently, works published prior to 1 January 1927 in the United States are in the public domain.
An additional roughly 1 million maps, 1 million volumes and/or pieces of music, and ~500k prints were also in the library's collection as of 1926.
A full set of the Librarian's reports, dating to the 1860s, is available at Hathi Trust. From at least the 1950s, annual holdings and aquisitions are broken out by Library of Congress classification, which makes for some interesting data, for those who find that sort of thing interesting.
I will add this study published in 2010, which estimated 129 million 'editions' published in all languages. Any text can have multiple editions (I don't recall how 'edition' was defined, but IIRC the study was of published objects not of texts), and "includes serials and sets but excludes kits, mixed media, and periodicals such as newspapers". It also contains a useful description of methods to bulk process bibliographic data. (Though maybe this is the Google study you talked about, because IIRC their data is involved.)
J-B Michel, et al, "Quantitative Analysis of Culture Using Millions of Digitized Books", Science (16 Dec 2010) https://www.science.org/doi/10.1126/science.1199644
By the way, I greatly appreciate your informed comments. It's some of the most substantive and interestin reading on HN.
Does it really? I think it just shows that you didn't know about something they offer, and decided to attribute that to them being malicious.
I prefer reading on my Kindle to paper books but I also love browsing and supporting book stores and the people who work there. I've always wished that I could browse a bookstore, take a stack of books to the front and have them scan a QR code (or something like that) on my ereader and sell me digital copies of books rather than the paper copies.