Pocket gets worse the more you use it (2019)
web.archive.org
web.archive.org
Sharing why I decided to develop my own "read it later" software.
1. I have a habit of saving web pages I like, most of which are in MHTML format, some are saved as single-file HTML, and others are web archives saved on iOS. Altogether, I have accumulated thousands of them.
2. On my computer, I can preview them one by one, but I cannot search through them. So, I developed a Node.js service that parses web pages locally and stores them in an SQLite FTS for full-text search. I deployed the service using Docker on my NAS.
3. To enhance my learning experience, I also developed an annotation feature that allows me to make notes and annotations directly on the offline HTML. For good articles, I save them to read slowly over time.
4. Gradually, the app gained some users. Since they were not familiar with Docker, I wrapped it with Electron and developed a standalone desktop version. The desktop version and Docker version use CRDT for peer-to-peer synchronization.
5. Some users provided feedback that it was inconvenient to annotate after saving, so I developed an open-source browser extension. Users can now annotate web pages directly in the browser, and the annotations and snapshots are saved automatically. When visiting the page again in the future, previous annotations can be restored.
I’ve my own app which can save pages as I saw them and create reader views out of them, but no offline access yet. This is a pretty niche use case but I feel these days more and more decent writing is behind paywalls (Substack and newsletters) so the current read it later apps are losing usefulness and they’re all frozen on features.
Nice to hear about hamsterbase, will check it out.
I'm definitely going to give it a try, although as a Firefox user the workflow is going to be a bit less nice.
EDIT: just tried it, it's very nice. Right now, the UI makes it a bit better than just using in SingleFiles, but not enough to justify $50.
However, if this would become a way to fully manage my Firefox bookmarks, I would pay for this, in the spirit of https://www.bitecode.dev/p/i-checked-if-browsers-could-cache.
Right now, the UI to manage the bookmarks sucks. Plus there is no way to:
- save their content for offline reading - search their content
This means it's are very hard to get value out of them.
Right-click on the webpage > select Singlefile > choose "Batch save URLs".
Using this method, I have converted hundreds of my bookmarks into html files.
I also think that Hamsterbase is not worth $60, so I offered a free trial. Only after it is released on iOS and Android will I consider charging.
FYI - your Obsidian and Logseq links on the honmepage are swapped.
I have updated the wording to: “Free during the beta period, Provide Believer plan, one-time purchase (60$) with lifetime updates.”
Kudos. I'll consider becoming a believer, since the 'scrapbook' extension got murdered by FF leaving me with years of inaccessible notes I've been looking for a replacement, even if not perfect it will definitely have to be based on open standards and open source.
Price seems to be reasonable, good luck with your project!
But your $45 price seems really reasonable, so maybe worth showing it clearly on the mobile landing page?
Not my area of expertise; just my personal 2 cents.
Currently hamsterbase provides rss subscription export, you can take any cell phone rss client subscription.
Could you elaborate on how you are doing this?
You can take a look at this article.
I will sync web archives to my Mac through iCloud, and then import it in bulk into Hamsterbase.
I like being able to read my bookmarks a week, or even a month later.
Browsers these days have settings which automatically unload tabs to make sure you don't run out of RAM. I personally don't run into issues with 16GB of RAM, so I don't have those settings enabled.
> and never turn off your computer nor close your browser. I like being able to read my bookmarks a week, or even a month later.
Browsers allow persisting the set of windows and tabs you have open across sessions, so this isn't an issue. I probably have a few hundred tabs open at any given moment across a few subject-specific windows, some of those tabs several months old, and I close my browser and shut down my computer all the time. Firefox simply restores the windows and tabs when I open it, so it all works out just fine.
I can't remember a significant improvement to Pocket in 2 years. Meanwhile, here's a non-exhaustive list of major, basic features that don't work: -Searching my own archive. This works maybe 75% of the time on articles I am 100% sure are in there. -Having a "permanent copy" of articles. This is Pocket's version of "full self driving," in that it doesn't mean what you think it means. -The ability to do any sort of advanced search. Pocket's web app has had a regression for exact match searching or multiple word searching for the longest time. I've raised numerous tickets with support and never get any meaningful responses. -Reliability of text to speech. There have been weeks-long spells where the better TTS voices just don't work. -Highlight exports or really any decent API interactions.
I'm guessing that this software is just on life support now. I'd love to hear from someone actually at Mozilla on what's happening with this software. What's been the headcount on this project?
Anyway, anyone on Hacker News is almost certainly served better by Readwise. Matter also looked promising but they didn't have an Android app so I never bought in.
I’m too baffled to be mad about this.
That said, yes, curiously little community around Pocket.
The Idea of Pocket is great - seems everyone needs a personal-knowledge-base. And the archival feature was/is cool (I added the same to my own PKB).
And Pocket for organizing, messy like the article says. I do tagging and SQLite FTS - works a treat.
<https://github.com/Pocket/pocket-ios>
And nothing server-side either, which I suspect is what you're referring to.
What is that? Do I need it?
Otherwise you're limited by what you can hold in your head - and as you get older and take on more ambitious work you'll find that's a pretty big restriction.
My version of that is spread out across way too many places now - it's mainly https://til.simonwillison.net/ and https://simonwillison.net/ for published notes, 700+ public GitHub repositories for code and issue notes, 100+ private repositories for private code and private issue notes, a Pocket account with 8700 items in it, plus the contents of my 15+ year old Dropbox folder structure, an old Evernote dump and a continually growing S3 bucket of larger things.
I tried to tie it all together with a unified search engine, but that needs some more work: https://simonwillison.net/2020/Nov/14/personal-data-warehous...
This feels like a comment without substance and in a way, it is. But I think it’s important to push back on this idea that if you aren’t managing a Second Brain while you’re GTD’ing a side project and Tools for Thought-ing your way through the western canon you’re wasting your life.
I don't usually need to scan back more than a few days, but one book lasts several months. You can use it standing up, e.g. in a SCRUM. It's resistant to having a mug of coffee spilled over it. It never crashes, and it doesn't need RAID or backups. You can attach post-its to the pages. You can use a stack of 4-5 of them as a monitor stand.
Downsides: you can't paste screencaps into it. And there's no search function.
Yes, it's a glorified grocery list. But it works better for me than any electronic solution.
otherwise all the shine and chrome in the world won't make it superior to a filing cabinet and some free time.
that's the difficulty point for all of these personal PIM type tools, they all have their own biased view on things like file format and 'touchability' of stored data in the (unannounced) hopes of trapping a captive audience; no good.
Open standards, open formats, and a revenue stream that isn't user-hostile -- then i'm in.
The main script is:
~$ cat share/bin/n
#! /usr/bin/env bash
cd $HOME/share # or wherever you want to put your notes
echo >> notes.txt
echo -n '# ' >> notes.txt
date '+%a %d %b %Y %T %Z' | tr -d "\n" >> notes.txt
echo ' #' >> notes.txt
if [[ $# -ne 0 ]]; then
echo $@ >> notes.txt
tail notes.txt
else
vim "+normal Go" +startinsert notes.txt # or nano: nano -R +-1 notes.txt
fi
This script is accessible from all my personal machines as well as appropriate work machines. It works well enough. I've devised a small system of tags and keywords, but that part is not particularly interesting. I just put the keywords at the beginning of a line and search for them with /^thekeyword/.More structured notes go into dokuwiki or just a LibreOffice file in the project folder.
All this is almost guaranteed to be readable in 20 years. The LibreOffice files may be at more risk, but at least there is more than one tool available to read them.
I use them for cooking. I refer to this one every time I want to make a Pisco Sour cocktail, for example: https://til.simonwillison.net/cocktails/pisco-sour
I see notes as a way of accumulating micro-skills - things I know how to do, but not well enough to remember off the top of my head.
This frequency informed the design, in fact. Easy to write, easy to read the last couple entries, moderate difficulty in finding older entries (in exchange for lower cost of making individual notes).
I think for me this is why it's so important for the system to be almost zero effort, else I'll just kinda stop using it. Just having an ubiquitous default place to dump something out of my brain helps.
Of note. A very slightly modified version of this runs on an ancient Thinkpad in my bedroom for diary/mental health reasons. I find the separation helpful. Those I almost never ever read, but I think that's the point of those.
Edit: at work I didn't really use this system, it's mostly for my own life. At work I used the ticket tracker, wikis, and calendar. I found that taking personal notes at work was best substituted by making actual documents that my team could find and use. There were much fewer uncategorized tasks at work.
I have a channel where my various IOT devices tell me stuff, a channel where my algorithmic trading bot notifies me, and a reminders channel that basically serves as an append-only, always-synched memopad.
When I produce images, it's nearly always in the context of a project, which has a wiki or other system. But that's outside of what I'd personally call "notes".
I believe that charts and diagrams convey a lot of information that cannot be easily reproduced with raw text. I guess it ultimately depends on the topic you are reviewing.
I have ambitions to do this occasionally, and I tracked things and took notes a lot for a while, but I never found myself using it.
https://til.simonwillison.net/linux/allow-sudo-without-passw...
Yup, I read that just a few weeks ago. Thank you!
For me it's a web-app tied to a PG database. But I started on SQLite. And it keeps notes, links (tagged), archive-page, put my own notes on it, etc.
Many folk use things like Evernote, Pocket, Notion, DokuWiki, etc for their "stuff".
My six-year-old rant still mostly applies (there've been some improvements to tag editing, and there's now in-article search for the Android app), but Pocket still massively underdelivers, increasingly as one relies on it more.
Most recently, I found I had to revert to a six-month-old version of the app in order to have both paginated navigation and to be able to open the "web view" in app ... which is where occasionally useful things such as tag editors live. Thread here: <https://toot.cat/@dredmorbius/110680672833582286>
That said ... I've largely given up on Pocket delivering what I would like to see, and if I can steer discussion here, I'd really like to know what Android-specific and/or e-ink tablet tools people have for managing large article repositories on device.
I'm familiar with Zotero and Calibre, though both are desktop only, as well as Wallabag.
Mendelay is a no-op given Elsevere's ownership.
Anything I don't yet know of?
My long-standing dream is to create KFC, Krell Functional Context, a document management tool that would include webfs and docfs components, inspired somewhat by Plan9OS and its P9.
More on that in another "dreddit" rant: <https://web.archive.org/web/20200918215041/https://old.reddi...>
There's a third-party Zotero client for Android named "Zoo for Zotero" that has worked well for me: https://play.google.com/store/apps/details?id=com.mickstarif.... However, I usually use it to read academic papers on my phone that I have already added to my library using my computer. I tried adding something new to my Zotero library on my phone as I was typing this comment, and sadly the interface seems a little clumsy: I can add papers to a collection but not a subcollection
I used to do a trick with Remarkable 2/kobo - exporting webpage on iPhone into Dropbox as pdf - resulting pdf is formatted for a small screen
For pdfs and web reading I use Boox Onyx Air - with 10.3” screen and full android browsing experience… But I need to sit and hold larger screens with two arms, while with Kobo Forma I can hold in one hand while layout down. Boox recently released Page - 7.8” lightweight android ereader, but I’m not sure that android device would hold battery as good as Kobo Forma.
Smaller devices can work, though 8" is probably the smallest I'd consider for serious reading.
My experience with Onyx is that when used just for e-books the battery life is excellent. Once you start web surfing or playing podcasts, battery starts getting chewed pretty quickly, though a charge still lasts most of a day. My preferred web browser (Einkbro) easily consumes 10x the power of Neoreader (Onyx's ebook reader software).
I'm increasingly finding the Web just ... generally less interesting (though of course it is compelling to a painful degree), and have amassed a large library of books and articles mostly in PDF, ePub, DVJU, and a smattering of other formats. Organising those usefully is the 2nd 90% of my content-management frustration. Writing is the 3rd 90%. Oddly enough, Termux provides a number of tools there that I find useful, including a healthy chunk of the libpoppler PDF tools, though that really needs an external keyboard (which I have, it's still somewhat awkward to use).
What I really wish Pocket would do is:
- Render Web-based articles in a fixed-layout, paginated, comfortably-margined (sides AND top/bottom) layout, with e-book-reader style touch-based page turns. (Whatever your favourite ebook reader does, that's what I'm talking about, Neoreader is quite good, I've also used FBReader and Pocketbook, and looked at the Koboreader software a bit.)
- Let me tag, highlight, and annotate those to my heart's content.
- Let me SEARCH the damned archive, by metadata and full text.
- Let me edit metadata where automated tools (or poor original creation) have fouled things up.
- Let me export lists of references with relevant metadata: title, author, URL, dates, etc.
I've adopted a number of practices to make up for holes in Pocket's offerings.
- I started tagging items with a "filed: <datespec>" tag a couple of years ago. I'd really like for that to be a built-in feature, and for search by time-period to be A Thing: past day / week / month / year, as well as spans. With my tags and an exported dataset I can at least build my own tools to do that.
- I've got a handful of projects that I'll tag a piece with if it pertains to that, so "project: <projectname>" Given the 25-character tag limitation, "proj:" would probably be a better prefix.
- Similar concepts for "task: <stuff>" and "error: <description>", where tasks are specific to-do type things (e.g., "hn" for "post to Hacker News"), or "print" for make a hardcopy printout. "share: <email-or-name>" could also be reminders to ping somebody with an article.
- "BOTI" is "best of the interval", the notion being that this is an exceptionally good item. More recently I've been relying on Einkbro's "save as ePub" feature which enables saving multiple articles, over time, to a single document. So you can build up a book of articles. (One thing I've realised from doing that is just how much reading I put on my plate. A small sampling of articles from the current year already runs to ~400 pages as an ePub.)
- I've got a notation for indicating my assessed article quality, on a scale of 0 to (for now) 5, with higher being better. 0 indicates content that makes you dumber (usually archived as examples of bogosity or negative propaganda). 1 merely establishes a basic fact, 2 gives some detail, 3 is a high-quality article, 4 exemplary article or a good book, 5 effectively establishes a new field or is otherwise a definitive reference (say: Shannon's articles on information theory). I try very hard to not over-rate content, and one of my projects is downrating a bunch of stuff I'd initially ranked too high.
Onyx made annoying product decision to not include microSD to expensive 10 and 13 inch models, while still keeping it on 8” Leaf2 and Page.
Are you using written notes while reading on your Lumi Max or typing only?
With such reading system I would look into Obsidian/Notion/Readwise/Logseq/Tana running on Lumi Max - at least Obsidian and Logseq has android apps. Note that Logseq is local first system - it stores files on your disk then the option to sync on your server for free or their paid sync service.
I too would have preferred an integrated MicroSD slot. Current cards are available in 1TB+ sizes, which should satisfy even my prodigious appetite.
I do make use of handwritten notes as well, both in documents (Neoreader) and in the notetaking app. I've found that far more useful than I'd anticipated, though searching notes is of course challenging.
With two layers of adblock (local LAN, IP-based, adblock on EinkBro itself), I see no ads in saved content, and very rarely in full Web view either. AFAIU it filters the source article through Readability.js, so the main article is what's saved, not navigation, asides, etc.
[1] https://www.wired.com/story/tiktok-platforms-cory-doctorow/
Specifically as regards Pocket, there seem to be both a constant war against those using the tool, and/or who are inflicted with it (it's packaged/bundled with Firefox these days), neither of which do much to engender good relations.
There's also what I see as a terribly misguided goal of trying to be a content recommendation engine, when I suspect that the problem most of us face isn't not having sufficient gobs of content shoveled at our faces, but in trying to make sense of that which we have acquired.
Tools such as workflows, date-based tagging, multi-tag search (being able to search, say "german" and "comedy", rather than "german" and "comedy" separately --- basic AND logic), date-stamping articles (when they were input, when they were read, when they were highlighted / forwarded), the ability to compile article lists to send via email, metadata exports, the ability to directly edit metadata (automated tools often get the basics wrong: title, author, date, etc.). Export of articles to a durable document format --- saving a single article to a PDF or ePub document, or multiple documents, etc., etc.
I've discussed most of this at length with Pocket's support for years, or rather, had some years ago and largely gave up the effort when there seemed to be no discernible movement in directions I'd encouraged.
iOS of course lacks an e-ink based device.
SingleFile even has an iOS extension (and easySearch is very nice for iOS search - although not quite as nice as a desktop experience, owing to iOS limitations).
Presumably you can roll your own pocket now.
I don't want to sign up anywhere. I don't want algorithms and suggestions. My app is on LAN network. I can add links, download data, search, add tags, highlight.
Some things do not work. It is work in progress. I am not a web dev so it is bare bones. I do not use JavaScript, as I hate it. It uses Django and celery.
Once a day I make my bookmarks public [2].
I am still learning. I just wanted to say you do not have to rely on anything to host a link aggregation software.
I copy/paste/append the url to a text file, links.txt, along with a cut&paste of what its about.
I can have multiples of those files. Search them with grep, whatever.
I abandoned bookmarking tools long ago, because they reinvented (poorly) things like search, backups, etc., besides probably selling my bookmarks to some data aggregator.
Perhaps the respective article authors didn’t use html markup semantically correct, so a naive reader runs into trouble. But if such a Reader is part of your core business proposal and you are, by now, Mozilla, one would hope articles didn’t get butchered in reader view.
I have no idea of how a product can be so unreliable at its main job for so long. Instapaper has been rock solid so far.
One a month they algorithmically select articles from your Pocket feed, have them printed and bound into a paperback book, and then ship them to you.
It's $10-14/mo, depending on if you want 1, 2, or 4 hrs of content. You can tag items as 'wpmustprint' and they will be in your next edition.
They use a url shortener to preserve links, and images, diagrams, and formatting are excellent.
I don't appreciate their move towards "social" feed sure, but the reader has genuinely gotten worse. I remember that it used to have the options to adjust:
- Line Height
- Page Width
- Text Alignment
- Fonts beyond the defaults (for paid subscribers, which I am one).
In the new app, it was all gone. Yet, I'm still paying the same premium price for it. The web app still has those perks, but it's probably because there hasn't been any update to it for the past 2 years or so.
I haven't had any use for the app for some years now. Like what many say in this thread, it's not that hard to build a system that replaces it well enough. You just need to make it scrape links, turn them to markdown (personal preference), and index the entire content for FTS in SQLite.
Thanks to the OP for reminding me to cancel my subscription. Now I just need to find a way to export my huge pocket archive before the subscription renews.
This sounds like the sort of thing the Pocket engineer[s] probably understand and desire, but it doesn't fit the vision of some designer who vetoes it and tells the engineer[s] they don't know anything about design and that users don't want to be cluttered with extraneous information and features.
Or maybe I'm just projecting my experiences.
I had really bought into pocket. I paid for the extra reading features and fonts, and I bought a kobo Clara specifically to use pocket.
I built a proxy to use omnivore (https://omnivore.app), an open source alternative on my Kobo ereader. https://github.com/Podginator/KoboOmnivoreConverter
I have used pocket for over a decade. They finally broke the camels back.
Why a company would completely gut the reading features, the primary reason for the apps existence, is a mystery to me.
It could be trained to become your personal pre-reading assistant, similar to information accumulators and pre-evaluators in governments agencies or company hqs.
Pocket is my default save - anything that looks remotely interesting gets sent there for me to browser at my leisure. If something I read looks like I'll want to revisit it, and I save the page to zotero, which is where most of the note-taking, tagging, and organisation takes place. Lastly I export my notes to LogSeq, which has zotero integration, so they can form part of my knowledge graph.
The browser plugins are the nightmare (Chrome comes to mind) You need to manually generate an API key and configure it in the browser. The process is very confusing, and too much work to configure on multiple machines with multiple browsers. The plug-ins are written by different developers, all open source as a hobby, so I while I think the effort is admirable and I hope it continues, it is still a beta experience.
I use it every day to get web content onto a Kobo ereader. It strips ads and often gets stuff behind paywalls.
My family shares the same pocket account and we often end up discussing things that get synced to all the kobos in the house.
This functionality is so useful and seems like a well kept secret. If it went away I’d really be upset.
https://www.theverge.com/23778208/kobo-pocket-ebook-reader-i...
Kobo can potentially fix it, but the short notice and breaking important functionality in one of their revenue generating products/services adds to my pessimism about Mozilla’s leadership, which is already very pessimistic.
Another example, look how much better Thunderbird is doing outside of Mozilla. Donations have more than doubled, but the product has a user centered roadmap, so imagine that. https://blog.thunderbird.net/2023/05/thunderbird-is-thriving...
(Original author.)
I just realized I have like 12,000+ links I've saved (many of them from HN over the years), and Airtable is great for tagging, searching, etc!
It extracts the page's title, the url, and even prompts you to give it some tags if you need some.
You still need to host your own webhook to save it, but that too should be straightforward. Just use firefox's readability.js to strip the bloat elements and save it in your database. You could save it into Airtable like what OP's been doing, or even save it as a file in your cloud storage.
> It takes me 45 seconds just to scroll from the top of the list to the end on the Android app.
There's the problem with over classification. Having had the same problem, I've reduced the tags to around 10 which is also a bit too much. Sometimes I think 3 should be enough: low, medium and high priority.
Far easier to search and more instantly readable than pocket is
The main challenge there is that it's difficult to reorganise information once it's captured, but search tools, the ability to reply to threads, and the automatic capture of basic metadata (sender, subject, date) are useful.
Mutt specifically has a highly robust threading algorithm. I'm not sure what happens if this becomes rentrant, that is, a given child node might have multiple parents from different threads. Mutt as given presumes that any given parent reference (indicated through headers) is either parent to, or a child of, other referenced emails. I'm picturing a web rather than a tree.
Email is also extensible with regard to headers, and with mutt it's trivial to write hooks and tools to work with those.
And once email is stored in a standard format (say, maildir), there are other tools such as mh which can operate on the corpus at the shell level.
For me the solution was to migrate to DEVONthink. It supports WebArchives for saving whole pages but also a few other formats that can preserve only content. Its search is great. It also has an iOS app. It’s web clip browser extension can save the page as you see it (most of the time) so can save pages behind a paywall or login.
I don’t think there are apps for other platforms so probably not a solution for the OP.
This Joplin? <https://joplinapp.org/>
Please note that I'm specifically trying to capture online content --- full articles with references, URIs, etc. Not simply take notes.
on mobile, do something clever, or just email yourself .