How does Firefox's Reader View work? (2020)
videoinu.com
videoinu.com
The heuristics are indeed brittle, which is why I had to add:
<span class="work-around-firefox-reader-mode-bug" rel="author">‌</span>
... to my website. Without that, Reader Mode would look for the first HTML element that contained the string "author", assume its content was the name of the author of the article, and make it a byline right below the title. Of course the element it picked was a section header, with the content "Completing an authorization". So not only did I end up with a blog post written by "Completing an authorization", but also the section that was supposed to have that header no longer had a header.So the hack is to add another element before the website content that it picks instead. ‌ is a ZWNJ, which I used because Reader Mode is too smart for its own good and will ignore an element that has empty text content or just an NBSP. So a ZWNJ is the next-best thing.
But it's great when it works, and I'm glad it exists. I have JS disabled by default and have lost count of the number of websites that use CSS to hide their content and JS to make it visible. Reader Mode defeats them easily.
It's bad enough that browsers assume my website, that's made up entirely of free-flowing text inside p tags, won't fit on phone screens unless I add a meta viewport tag. I'm not going to also make changes to the content to satisfy them.
This anoys me to no end as well. It's especially ridiculous now that mobile browsers provide a "Desktop" mode that can be used as a fallback for those sites that do need a larger page width.
ZWNJ and its other zero-width friends is my ninja secret. Oh, are you demanding this text box be filled with something? No thank you. Have a ZWNJ.
Few, if any, sites that are "checking" for content being in some widget "correctly" filter out the entire Unicode space class. (I mean, I assume somewhere out there there is one, but so far this has never failed me.) Most of them just filter out ASCII space 32. A few are smart enough to filter out tabs as well; if Reader is also filtering out NBSP that already puts it at the head of the class. I'm not sure I've yet encountered one that is filtering out the entire Unicode class.
(On the other hand, I would submit that by the time someone has taken the time to find one and copy & paste it into your text widget, you should consider your options carefully. The user has just demonstrated significant technical sophistication. It may not be your wisest move to also filter that out. You want some Zalgo emojis? Because I can give you Zalgo emojis. You may not want Zalgo emojis; there are still some systems out there that emojis can break. You want to find out?)
Why would a section header in your article contain "author" if you didn't mean for it to designate who the author of the article was?
It's bad substring matching. It's "authorization", but apparently Firefox just interprets that as "author".
There's no reliable way to do it. The right way would be to only use explicitly defined metadata and if it's not there to default to showing nothing.
class="authorization"
or something like that in the enclosing element, so it's not that they're inferring metadata from unstructured content - it's that they're drawing bad inferences from the metadata provided.That said, inferring metadata from unstructured content is literally the whole goal of Reader mode - to make pages more readable even if the original source didn't design it to be - so while this particular bug is avoidable, others may not be.
Is that the whole goal, or is it to get around the design choices of the original?
These are two different things.
Ha, you wish. Skyscraper is the name of a standard Internet advertising unit. It’s fairly common for adblockers to block it. https://www.marketingterms.com/dictionary/skyscraper_ad/
Really, let's switch to a user perspective once and consider - what if I always want reader mode? This is technologically complete impossible and all the solutions are a band-aid.
Firefox and others' attempts rely on the page authors' goodwill. But some pages will always attempt to frustrate reader modes.
Alternative approaches for content extraction use machine learning such as [1], but they of course need to be updated for culture- language- and technology-specific changes.
It's a mess and will remain so for the foreseeable future.
In fact we wouldn't even need any new web standards to implement something like this. Any website following the latest Web Content Accessibility Guidelines and using the correct page structure and aria- tags should be fully semantically parsable and displayable in any format or medium.
Just got a link to a documentation for an API that I can't read because of too low contrast. Dark Reader plugin nicely fixed it so didn't even have to try reader mode. It sucks to get old anyway, you young people don't have to make it worse. You will be in the same seat in 20-odd years so plan for it :-)
The purpose of ads is often to get you to just recognize the brand (and/or associate positive feelings with it) so that if you are ever in the market for their products you'll be more likely to choose their brand over a "no-name brand" (ie. one that you haven't seen ads for).
So you might think you're immune to advertising, but you probably aren't.
I’ve been out of web for a long time but it seems to of mostly died with the advent of React and the single page web application. Now it looks like everyone apart from public government sites with strict accessibility requirements just don’t bother putting the effort in.
For example, somebody had build an app that scrapes free recipes from the web, so that it's shown without ads and the ridiculous fluffy lectures that are typically part of it.
Great for users of that app, not so great for the recipe authors depending on this model.
This is also why RSS is dead. Both Twitter and Facebook had it before, but killed it. You don't want to give away your goods for free.
It's easy to put this problem at the recipe bloggers or Facebook, but we're just as responsible for this. We don't want to pay for anything so it's ads. Then we avoid and block ads.
We lack a payment culture.
Then there'd be no reason to optimize the web to serve up ads, because legally you couldn't.
The world needs to switch away from advertising for the good of humanity.
* A significant number of people do not really use file names. https://jayfax.neocities.org/mediocrity/gnome-has-no-thumbna...
* Twitter offers no official way to make bold or italic text, but some people make it work using math symbols. They either don't know, or don't care, about all the accessibility and search tooling this breaks.
* Try downloading a C program off the internet and compiling it on anything other than the original developer's computer.
In none of these cases does anybody benefit from the lack of usable metadata, yet nonetheless the metadata doesn't exist.
This would happen if page authors just participated fully in the Semantic Web.
:(
The future we could have had, but did not get
Reader mode is a reasonable low-overhead script-stripper that mostly does its job. My expectations aren't high, but it always delivers.
Except not on Firefox Mobile. This has been the most frustrating regression I've experienced, suddenly losing my favorite things about Firefox.
So STOP READING THOSE PAGES!!!
the reason the pages are full of ads is because people keep reading them.
[1] Meaning I've spent in the ballpark of five figures in power bill extracting content from html. Not my personal money, of course.
I'm not gonna stop reading the newspapers or blogs I consider most trustworthy (or less bad) just because I don't like their design choices!
I turn the stuff on on a per site basis when I want to use an actual web app, something like github, and every now and then I get an article that wants to fight me on it and my usual response is "well I guess you don't want to influence me with your words bad enough" and I back out. But it works 99% of the time exactly how I expect it to.
I assume you only want always-on reader mode for articles -- and detecting what's an article is another NLP problem. Yet both the completeness and article detection can probably be solved through heuristics in 90% of cases (the evidence is that we DO use reader modes). Maybe it depends on how much the last 10% frustrate you.
[0] I'm working on a browser extension that does this: https://github.com/lindylearn/unclutter
Sigh… welcome to Opera versions 9 to 12. Not only did this complete technological impossibility exist, it did so over a decade ago.
Tools → Preferences… (ctrl+f12) → Advanced → Content → Style Options… → Presentation Modes → Default mode [User mode]
Why do you think Opera got rid of it?
http://enwp.org/Opera_(web_browser)#History
Opera A.S. threw the whole browser with its countless innovations and usability affordances away. The next major version after 12 was a Chromium derivative.
Yet another example of "worse is better".
javascript:( function(){ let i, elements = document.querySelectorAll('body *'); for (i = 0; i < elements.length; i++) { if(getComputedStyle(elements[i]).position === 'fixed' || getComputedStyle(elements[i]).position === 'sticky') { elements[i].parentNode.removeChild(elements[i]); } } document.body.style.overflow = "auto"; document.body.style.position = "static"; })()
that's the bookmarklette i can gist it if someone wants a nicely formatted version but either way it works.
https://til.simonwillison.net/shot-scraper/readability
pip install shot-scraper
shot-scraper install
shot-scraper javascript https://simonwillison.net/2022/Mar/24/datasette-061/ "
async () => {
const readability = await import('https://cdn.skypack.dev/@mozilla/readability');
return (new readability.Readability(document)).parse();
}"Right now I'm using a bash script instead of YAML file for screenshotting multiple sites I maintain, as there is no option to add a timestamp to the filenames, something like this (simplified):
declare -A capture
capture[www.foo.com]=https://www.foo.com/
capture[www.foo.com-bar]=https://www.foo.com/bar
for key in "${!capture[@]}" ; do
shot-scraper ${capture[$key]} -o - > $key_$(date +"%Y-%m-%d_%H_%M")
done
Otherwise it is a terrific tool, able to screenshot websites that are otherwise very difficult due to CDNs, javascript, etc. Thank you very much!The way I'd solve this problem would be to develop a new file format that was similar to Markdown but contained semantic elements (eg "this image is relevant to this bit of text", or "this text explains that text".
This would be rendered into HTML by a browser extension so the browser could display it, which would help keep authors honest (because if the server renders it, we're just using HTML again).
The format would contain no styling information at all, apart from the semantic tags, so styling would be entirely on the client. You could pick the way you wanted your articles to look, which would probably be different per device.
This sounds a bit like epub and would probably be great for book reader devices as well, while being extremely light to load because there isn't much you can do with it.
Unfortunately the easy part is designing the protocol, the hard part is universal rendering and adoption.
Plain HTML is great for documents. It’s just the browsers‘ failure to render it nicely that makes authors feel like it’s necessary to style to make it presentable.
I think the overall best solution would be to switch from a layout language to a programming language where each program manages special input and output variables that are marked semantically and given display and layout hints. No HTML, no CSS, just programs on a virtual "semantic display" with input and output of data to the user.
It's not like W3C and others hadn't tried to renovate HTML - that's what W3C's XML initiative was about in 1998, with a generic namespacing mechanism in anticipation of a wealth of new vocabularies (of which SVG and MathML, but not XHTML made it).
I'd argue W3C people were so wound up in XML and "enterprise" tech for like ten years that the web stack was unprepared when the iPhone with requirements for mobile sites came around, and W3C became an easy target to take HTML away from.
Mentioning this to warn against inventing new meta "formats" all the time - everything needed for HTML was already in existence before 1986 when SGML (on which HTML is based) was published.
Including markdown and custom shortform syntax. Large parts of markdown can be specified using SGML shortref, yielding a shortform syntax that expands to HTML proper, capturing exactly the way markdown was originally specified.
It sounds like you're reinventing the Semantic Web:
I find most people who consider their projects to be "no-frills" have quite a few frills indeed. This includes personal blogs, but Sourcehut and the constant praise for its "clean" UI is one other example of this sort of thing. Load it up in w3m or Lynx and it very much feels like a site intended to be read in a conventional desktop browser but filtered and constrained to a text-only medium.
Actually clean design would start with, "How would I expect this to feel if I made a custom, terminal-based (but not command-line driven) Sourcehut client?" and then figuring out the appropriate markup you'd need to spit out to achieve that in a line mode browser that understands the basics—forms, content ordering, etc.
(PS: Now transport that same artifact to a conventional browser. Is the experience better or worse that what's currently available?)
That’s if you are fine with using JavaScript on the backend.
So I had to read a ton about this. I ended up using a heavily modified Kotlin version of Readability:
https://github.com/dankito/Readability4J
https://play.google.com/store/apps/details?id=com.pranapps.h...
Currently I'm using it to parse a web article, package into a .mobi file and send it directly to my Kindle Paperwhite.
If you want to read your favorite content on Kindle, give it a try[0]
[0]: https://ktool.io
I just tried sending this article[2] and 10 minutes later it still hasn't delivered yet. This article contains a lot of images, so I understand it's not going to be quick, but with KTool, it does deliver within 2 minutes.
But that is not the only reason why I build KTool. I don't want the articles to include links when reading on my Kindle because it's bad UX (easy to mis-tap), distracting and reduce comprehension (see here[3] and here[4]). Instead, I push all these links to the bottom of the article, which make it still accessible yet improve the reading experience.
[0]: https://ktool.io
[1]: https://chrome.google.com/webstore/detail/send-to-kindle-for...
[2]: https://waitbutwhy.com/2017/04/neuralink.html
But I'm working on a feature where you can send content from client side. Meaning you can use any other extensions to modify that page's content[0]. KTool Extension will take that content as input, parse and then send to your Kindle.
[0]: https://chrome.google.com/webstore/detail/bypass-paywall/kko...
Hey you can sign up for an account to receive updates, or just reach me directly via email. My email is on my profile. Thx
Edit: so much for the illusion of "semantic HTML" ie where you need heuristics and are entering an arm's race vs publishers to make your HTML even readable
1. https://addons.mozilla.org/en-US/firefox/addon/automatic-rea...
There's Gemini and gopher, protocols designed around transferring and rendering plain documents.
There are terminal based browsers like Lynx that render pages in a terminal, of course it's all text on your end.
If you're talking about a GUI program to do what reader view does and only that, the reader view code is open source and available from Mozilla, I'm sure it wouldn't be much to build a webview app for mobile or a simple GUI that sits on top of curl or wget or something like that, fetches the page, processes a URL and renders the text. You're probably going to be manually entering URLs though, I'm not sure in the modern web how you even click from one site to another, how you even wind up at a URL for a written article using only reader view.
There's also a huge long tail of parsing issues because web pages are not static documents. You'd want a fallback to the original HTML -- so why not use your main browser? There are extensions that activate reader mode automatically on pages where it's supported [0].
[0] I'm working on one of them: https://github.com/lindylearn/unclutter
It was far less code and worked perfectly 95% of the time (though, the average web-page was also a little simpler 12 years ago). But that code would have quickly ballooned out if our use-case had called for addressing the other 5% of webpages, as the Firefox Reader View must do.
[1] https://www.alfajango.com/blog/create-a-printable-format-for...
Since it turns a website into very plain (X)HTML it‘s fairly easy to use it to make a browsing proxy or automatically produce epub files for e-readers, which is what I do.
Edit: Here’s the proof of concept type code I use: https://gist.github.com/solarkraft/d6306f17a761fcb5ce47f2be7...
It’s a bit crappy, but it works for me :-)
https://github.com/masukomi/arc90-readability/#readability
Please submit a PR if there's something i don't have listed there.
I'm surprised at how incredibly mono-cultured web developers have become in testing browsers other than Chrome.
OR .... they know, and they don't care ....
Seems like it's a really difficult task.
But generally it's awesome.
There are several open source projects for extracting web contents. However, there are three extractors that I've worked with and give us good result:
- readability.js[1], web extractor by Mozilla that used in Firefox.
- dom-distiller[2], web extractor by Chromium team, written in Java.
- trafilatura[3], Python package by Adrien Barbaresi from BBAW[4].
First, readability.js, as expected is the most famous extractor. It's a single file Javascript library with modest 2,000+ lines of code, released under Apache license. Since it's in JS, you can use it wherever you want, either in web page using `script` tag or by using it in Node project.
Next, DomDistiller is extractor that used in Chromium. It uses Java language with whopping 14,000+ lines of code and can only be used as part of Chromium browser, so you can't exactly use it as standalone library or CLI.
Finally, Trafilatura is a Python package released under GPLv3 license. Created in order to build a text databases[5] for NLP research, it mainly intended for German web pages. However, as development continues, it works really great with other languages. It's a bit slow though compared to Readability.js.
All of those three work in similar way: extract metadata, remove unneeded contents, and finally returns the cleaned up content. Their differences (that I remembered) are:
- In Readability, they insist to make no special rules for any website, while DomDistiller and Trafilatura give a small exception for popular sites like Wikipedia. Thanks to this, if you use Readability.js in Wikipedia pages, it will shows `[edit]` button thorough the extracted content.
- Readability has a small function to detect whether a web page can be converted to reader mode. While it's not really accurate, it's quite convenient to have.
- In DomDistiller, the metadata extraction is more thorough than the others. It supports OpenGraph, Schema.org, and even the old IE Reading View mark up tags.
- Since DomDistiller is only usable within Chromium, it has the advantage to be able to use CSS styling to determine if an element is important or not. If an element is styled to be invisible (e.g. `display: none`) then it will be deemed unimportant. However, according to a research[6] this step is actually doesn't really affect the extraction result.
- DomDistiller also has an experimental feature to find and extract next page in sites that separated its article to several partial pages.
- For Trafilatura, since it was created for collecting web corpus, it main ability is extracting text and the publication date of a web page. For the latter, they've created a Python package named htmldate[7] whose only purpose is to extract the publication or modification date for a web page.
- Trafilatura also has an experimental feature to remove elements that repeated too often. The idea is if the element occured too often, then it's not important to the reader.
I've found benchmark[8] that compare the performance between the extractors, and it said that Trafilatura has the best accuracy compared to the others. However, before you start rushing to use Trafilatura, you should remember that Trafilatura is intended for gathering web corpus, so it's really great for extracting text content, but IIRC is not as good as Readability.js and DomDistiller for extracting a proper article with images and embedded iframes (depending on how you look, it could be a feature though).
By the way, if you are using Go and need to use a web extractor, I already ported all three of them to Go[9][10][11] including their dependencies[12][13], so have fun with it.
[1]: https://github.com/mozilla/readability
[2]: https://github.com/chromium/dom-distiller
[3]: https://github.com/adbar/trafilatura
[5]: https://www.dwds.de/d/k-web
[6]: https://arxiv.org/abs/1811.03661
[7]: https://github.com/adbar/htmldate
[8]: https://github.com/scrapinghub/article-extraction-benchmark
[9]: https://github.com/go-shiori/go-readability
[10]: https://github.com/markusmobius/go-domdistiller
[11]: https://github.com/markusmobius/go-trafilatura