Tell HN: Today I learned Epub is just HTML/CSS
en.wikipedia.org
en.wikipedia.org
You can make copy and paste work normally in Apple Books with the solutions on https://apple.stackexchange.com/q/137047 ‘Don't want iBooks to always paste the “Excerpt From” of what I have copied’.
Presumably if Google decided to add EPUB support, Edge would get it back too, but Microsoft hasn't decided that feature is valuable enough to add onto their modifications.
I also think it's particularly sad that Windows 10 had a perfectly good EPUB app, Reader, that Microsoft deprecated aggressively to force everyone to read eBooks in Edge... only to remove eBook support from Edge too.
An ebook, unzip, is rarely bigger than 1Mo, which is lower than what most page are today.
Oh no, we can't have that. Here are some "beautiful, performant and lightweight" Electron ePUB readers/organizers: https://www.electronjs.org/apps?q=epub
Rather than complaining about the Electron apps, write better cross-platform apps that don't use Electron.
I know it's fashionable to dog on Electron, but if it didn't fill a legitimate need, people wouldn't use it.
It's easy to compare an Electron app, that actually exists, with some imaginary native app that doesn't.
It's not so easy to find the budget and personnel to actually build dedicated apps for minority platforms. The choice generally isn't between "bloated Electron app" and "sleek native app". It's between "bloated Electron app" and "no app at all".
I suspect Electron frequently fills a need for the developers, instead of their users. It's easy to deploy, cross-platform and stable, I give you that, but the users pay for it with RAM, disk, CPU and energy.
A single user might not mean much, but multiply that by the millions of Electron programs installed, that's the scale of lost resources that pay for the advantages of Electron.
If there's "no need" for them, why are people using them? How come you get to decide what other people "need"?
> I suspect Electron frequently fills a need for the developers, instead of their users.
It fills the need of the users to have actual apps they can install and run, rather than imaginary ones.
> that's the scale of lost resources that pay for the advantages of Electron.
People don't hand-write programs in assembly language any more, either, even though that means that you can no longer write a word processor that runs in 12K of RAM.
I'm not saying all Electron programs are bad. VSCode, for example, is surprisingly good for many use cases, and being an IDE with many features, its resource usage is pretty justified.
I'm also not buying that people need software that only exists as Electron programs. Check the categories in https://www.electronjs.org/apps, there's even taskbar notifications and app launchers, do those merit running a dedicated browser?
It's not that users need "non-imaginary" software and Electron fills that need, it's that most users don't know about native and web frameworks, and they will install software as long as their computer can run it, even if better alternatives exist right now.
I see this sort of assertion pop up often, but it never is accompanied by specific verifiable examples of what those options are.
Can you point out a single example of said native software that fills that role in every platform? A single one.
(For instance : VLC, Spyder...)
What is python qt support like for android and so forth by the way?
Compared with simple html+javascript running on a webview, Qt is relatively hard to maintain and develop, relies on source code generations and processors to work, prototyping tools are subpar and undermaintained, has no support for centralized theming, it's basic support for component-specific theming is already CSS shoved in a convoluted way, its model/view take is absurd and very poorly engineered, and yeah there's the fact that it forces you to write frontend code in C++. But wait, it's not even C++ because it requires code to be preprocessed to generate boilerplate code.
And let's not pretend that Qt's widgets successor is already a markup+javascript combo that takes the bad parts of javascript and bundles it with the bad parts of a custom markup language.
Calibre seems to use all the memory after some time and then it starts to use swap memory...
Not the fault of the library developer - they’re just trying to help, but it’s like instead of fixing holes in the ship, we build pumps to dump the water out. Too many pumps on the ship and it gets bloated and can’t take any cargo. This is the current web in a nutshell. We need a ship captain that can guide us authoritatively.
I have tried a lot of browser extension based and standalone readers in the past and they all render things differently with different bugs. Something doesn't add up.
It's XHTML to be precise, and that doesn't change no matter what processing is done by readers, it's still (just) XHTML. The format is literally described in the submission :)
> they all render things differently with different bugs
Just like HTML did in the beginning.
> Just like HTML did in the beginning.
ePub wasn't born yesterday. It's revision 3, first released ~2006/7. XHTML was proposed to correct and prevent the kind of problems html4 had because of how it organically grew, specifically relying on its XML root (no pun intended). Epub should definitely not suffer from bugs like HTML had in the pre-XHTML and pre-HTML5 era.
I sometimes have bugs like:
- whole book is black
- some pages can't be loaded/read so I have to skip them
- some toc and back link don't work like they should (probable bad markup)
There's also some readers oddities:
- completely inconsistent line-height
- aligned setting not working at all
etc.
Anyway, there's a reason XHTML2 didn't happen and we got HTML5 instead. Either ePub has some extensions that are not trivial to implement or most readers are buggy. Or both.
The three main components of the ePub (aside from the actual pages) are the TOC, the spine and the manifest. The manifest basically tells you where everything is, the TOC is the table of contents which can link to various pages and the spine gives you the traversal order.
Some mistakes I've seen are using the TOC to traverse the book. Using the spine to traverse the book but not handling hidden pages properly. Not handling two page spread properly.
So yeah the spec is nuanced and it would be easy to make a reader that worked with a lot of books but then had weird issues on another set of books that aren't particularly different. We ended up writing our own parser because we kept finding issues with the main open source ones.
I recall using this repo (https://github.com/IDPF/epub3-samples) to test specific functionality to make sure it was in line with the spec.
So it's not quite just HTML/CSS when packaged up, but it is just HTML/CSS when it comes to the actual text content.
Other than getting the various constituent files zipped up, with their interrelated contents synced, the only other oddity is that the renamed ZIP must always start with an uncompressed 'mimetype' file.
The IDPF maintain the standard. The easy one is v2 (http://idpf.org/epub/201) but that is now deprecated. Unfortunately v3 allows more interactivity and scripting - and we all know how bad the tech industry is at keeping that kind of stuff secure.
Firefox addon: Download files and read and modify the browser’s download history, Access browser activity during navigation, Access your data for all web sites
Chrome Web Store obscures the full permission list, but the comments for that extension admits: right "read and change all your data on the websites you visit" needed
It's 2021, you should view all browser addons as a threat.
Wait, so you could actually do all of that and then let it interact with APIs as well? When JS gets involved like this, I can see some crazy applications in my mind packaged as an ".epub" book.
I guess it depends on what reader you're targetting then. A quick cursory search shows that not all of them support JS. Makes sense to me.
With epub? I hope not!
> I guess it depends on what reader you're targetting then. A quick cursory search shows that not all of them support JS. Makes sense to me.
Any epub reader supporting javascript would very much be an antifeature.
https://www.w3.org/publishing/epub3/epub-contentdocs.html#se...
https://github.com/phuff/epub_builder
I built it because I wanted to have something that made a daily brief news paper that was personalized and sent to my kindle. It makes an epub and uses kindlegen to convert it to a .mobi. There's a lot of fun epub formatting stuff you can do.
Here's the system that makes the daily newspaper, but it's been so long I'm not sure it's actually functional code outside of my production version:
For example, rendering math on the web has been a solved problem for many years thanks to MathJax and KaTeX, but these require JS, so cannot be used in ePubs (unless you know the reader supports scripting).
If anyone is interested, I wrote a mega blog post about my journey to produce decent-looking math equations inside a ePub (and mobi files): https://minireference.com/blog/generating-epub-from-latex/ some discussion from when I posted on HN https://news.ycombinator.com/item?id=26356903
I'm doing testing though, and hopefully going to see more MathML in the future (in browsers and ePub readers).
Out of all my epub math textbooks I have exactly one done with MathML and it's fucking glorious compared to the dogshit image based ones most publishers put out with blurry images intended for the 800x600 readers of 2007 and not modern 300 dpi ones. I had one book where literally 2/3rds of the equations were just missing from the file and unviewable on any device. This started on PAGE 7. I then had to ask the publisher to fix it, which they sort of did by replacing with a PDF version.
The struggle to get better at math and find time + motivation for it is real.
I'm the high school textbook guy, if that rings any bells.
Despite being HTML/CSS, the layouts aren't particularly interesting though. Most content reads from top to bottom, and the formatting is identical whether you read it on a phone or a tablet.
EPUB3 is a dog's breakfast -- it's hard to think of a better example of "second system effect". As far as I know, there's still not even one reference implementation that supports the full standard, even though it's been out for nearly 10 years. It gains you very little over EPUB2 for standard novels written in western scripts. EPUB3 is only needed if you require embedded scripting, support for non-alphabetical or bidirectional scripts, etc. I believe that most commercial "EPUB3" files still have an EPUB2 toc.ncx file and are designed to fall back to EPUB2 if the reader doesn't support EPUB3 (there are a lot of readers like this).
Something that's easy to overlook: "The mimetype file must be a text document in ASCII that contains the string application/epub+zip. It must also be uncompressed, unencrypted, and the first file in the ZIP archive".
All the other files in the ZIP can be compressed normally.
What this means in practice is that uncompressing an EPUB is easy (just rename it to .zip, if necessary, and run unzip), but recompressing it requires some care.
Assuming you've got your book's content in an OEBPS folder, and the container XML file in the META-INF folder, you can do it like this:
zip -X0 test.epub mimetype
zip -X9Dr test.epub META-INF OEBPS
(edit to fix code formatting)The first is the "Standard Ebooks"[1] toolset, which is a suite of Python scripts to create, process, and build ebooks in all common formats. The results on the Standard Ebooks site speak for themselves. They're impeccable in every way, and far better than many big name, commercially produced efforts.
GitHub: https://github.com/standardebooks/tools
How to use: https://standardebooks.org/contribute/producing-an-ebook-ste...
The second is Sigil, which is a great editor if you prefer to work with a GUI:
GitHub: https://github.com/Sigil-Ebook/Sigil
Homepage: https://sigil-ebook.com/about/
Note: You don't have to be a full-time emacs user to use nov.el.
What's even worse - almost all Epub readers don't do proper sandboxing.
I recently published a book and going through the w3c epub specifications was a pain. Instead, I bought a book I wanted to read then reversed engineered it.
For small files you can use the w3c online validator, which will give you an overwhelming list of errors.
Note: The kindle does not support epub, instead it uses kpf. For that you have to download a 333 MB program to convert your epubs.
Care to explain yourselves this time?
EDIT: It would be ok in a live discussion. But on a forum, not so much.
Am I 'not allowed' to express my own opinion of this post via typing in text, even if it says 'cool'.
I think I am getting tone-policed.
Java software (jar, war): zip files
Android packages (apk): zip files
OpenDocument Format (odt, ods, odp): zip files
Quake 3 / OpenArena / Urban Terror / etc. (pk3): zip files
Firefox/Thunderbird/Chromium extensions (xpi, crx): zip files
EPUB: :D
If you need to cram a bunch of files into one package, zip is the obvious candidate. There are well-tested libraries and apps for dealing with zips for essentially every language and operating system.
As the saying goes, "don't mess with success".
I can recommend Kobo as an e-ink e-reader that supports ePub with one caveat: Kobo requires you to sign-up for a Kobo account before you can even use the device - horrible. It's easy to search online to find a way to bypass this.
Although Kobo is an alternative to Kindle, you won't find the range of titles that Amazon sells. However, I think e-readers are best for text-only, small paperback-sized books. Anything else simply doesn't fit the small screen and is inferior to the physical version of a title. (Amazon sells a lot of Kindle titles that are simply unsuitable for small e-reader screens.)
At least it's possible to strip the DRM on Amazon books with the right set of tools, and Calibre is able to convert them to EPUB.
I do feel that the e-ink reader screen size does matter because reflowed text only works well for small, paperback-sized books. Any book larger than this small size that also features tables, charts, images, diagrams, code listings and more, will not display well on a small e-ink screen.
Is the kobo better in this regard?
- Shows as an external drive when connected via a micro-USB cable
- Copy and paste books, then when you unplug from your laptop your Kobo sees them
- With Calibre you can tag your books; tags become collections on-device
- Understands more formats, but Calibre can also convert to EPUB during copying
- You can also read the Adobe-DRM protected books (or ask Apprentice Alf)
I am thinking of a scenario where, if there is a collapse (societal, economic, political, or technological), how can knowledge be disseminated and preserved in a resilient way?
Even better if the paper and ink can be made onsite, and the printer can is repairable by someone within the nearby geographical region.
Sure, we all have pleasant visions of truly interactive ebooks driven by creatively built JS content. But in the real world, ads would be the first thing to be added if JS was supported.
Isn't epub a zip file of a bunch of html docs, metadata and images?
If that's the case, you didn't have a genuine EPUB to start with. To meet the spec it needs to be in a container (a renamed ZIP file) and have a handful of related metadata and navigation files alongside it.
That said, the actual text of the book is done by HTML/CSS, but within the EPUB container file.
No, why would I want Electron ??
IMHO EPUB should specifically be script-free.
And if I wanted scripts in my document¤, then instead of packaging the whole browser (that would be overkill!!) with the document like Electron does (?), I would rather use MHTML.
¤ The following was supposed to be my example, but since this website uses Flash, these scripts are no longer easily ran. But since they are described, you should have a good idea what they are going for :
http://resonanceswavesandfields.blogspot.com/2007/08/phasors...
Yeah maybe I am just knee jerking, with AV1. Broadly speaking however there would be two schools of image and video compression, one for computer generated images and one for images of natural concepts. So AV1 isn't really suited for animated graphs and the like at the end of the day.
On second thought maybe allowing in AV1 would open the door for sound, which really would take us to far away from the book format.
I do definitely see the need for having animations such as the ones you link in your post however. So many things that just takes sentence and sentence to describe can still be described way more efficiently with an animation.
CSS Animations or animated SVG maybe?
Yeah, animated SVGs would be even better for this use case - but you need raster graphics support too for other use cases (some animations might require the inclusion of photographs).
I don't see how sound support is an issue - no more than color and high frame rate support are issues for a format that might also end up displayed on grayscale only displays incapable of high framerates. It's up to the creator of the media to take these into account (or not).
I think that we should define a book as something that you can use only with your eyes.
If I bought a book and I would need to use headphones to "read" all the content in it I would feel a bit defrauded.
(Though a (sub?)format dedicated to "e-ink" readers might actually be a good idea ?)
http://idpf.org/epub/30/spec/epub30-changes.html#sec-new-cha...
First, there's opportunity for that web endpoint to stop functioning. Second, there's opportunity for that web endpoint to become taken over by malice. And third, there's opportunity to turn that caption space into an advertisement.
So, to put it succinctly: fuck no.
> This would allow you to display an "obsolete sample code" warning below the examples.
So now the book isn't timeless. It changes. It's no longer a book.
A better idea: include the "obsolete sample code" warning in the book and ask the user check for the latest practices at a URL also included in the book.
> When the user is not connected to the internet, you could display "Get online to know code snippet status". And so on.
When the user is not connected to the internet should be the only case ever considered for a book. Otherwise you're not writing a book. You're writing an app.
Cheers.
I learned this a while back when I was deep in @font-face and web fonts and style sheets.
I was almost immediately discouraged because, modern features like @font-face were inconsistently supported.
Haven’t checked I a while now, maybe it’s better.
- points gun Always has been.