Managing my personal knowledge base
tkainrad.dev
tkainrad.dev
This video covers some of the things you can do with it in an academic setting but relatable. https://m.youtube.com/watch?v=JQD5c8A_D2g
0:33 - Math equations
1:45 - Replay text
2:45 - Ink to text
3:20 - Research tools
4:39 - Immersive reader
6:01 - Web clipper browser extension
7:12 - Save emails to OneNote
NB: https://www.microsoft.com/en-us/microsoft-365/blog/2015/11/1...
OneNote has a really scary number of obvious bugs, with pen input especially. Writing disappears immediately or jumps left and right, the document randomly stops updating the screen, the program gets stuck in a high-CPU loop making huge input lag and needs restarting, selecting writing doesn't select what you draw around but a smaller area inside that, etc. It's incredibly frustrating and difficult compared to the Android pen-based notes app I use, Squid (which has a much narrower focus, but is much better at it).
One drive is blocked at work so onenote is effectively useless to me now.
For now, I've started moving my notes to a static site using hugo. It's a bit less convenient, but until I settle on a better permanent solution at least I have text files I can move around and manipulated in bulk if needed.
it blows my mind that there is no good migration strategy for pc base OneNote databases to cloud based ones.
Export is really the weakest area of OneNote. As far as I can tell, there is no way of batch exporting that also exports embedded attachments (for example, if you embed a PDF or Word document into a page rather than "printing" to the page).
I'll give you two guesses as to why that is.
1. Uncertainty about data export. I don't want to be stuck depending on a third-party proprietary solution for something so important to me. (That said, Microsoft is typically quite good about supporting products for a long time)
2. I actually find Windows unpleasant to use, and much prefer linux. (Arch linux + i3/KDE happens to be my preferred setup).
The documentation isn't very good (it's unclear which version) and I ended up having to install the local version on my Mac and manually copy/past the pages I wanted because the copy notebook dialog didn't work with notebooks that have many pages. The worst part is not being able to access the raw data and simply copy that from one place to another.
I've switched to markdown files on Sublime with the Markdown Editing package for now, because they are easier to read and edit, portable and I can convert my notes to wiki pages, blog posts or documentation very easily. I'm still not 100% happy because of the different flavours and how pandoc converts files differently sometimes (especially lists).
My goal for now is to have a simple way to record (the files) and manage (the folders) notes, collaborate (I currently use gitlab repositories) and potentially use them later, e.g. by copying relevant parts to make posts, documentation etc, or by writing a script that finds relationships based on the content (something I've been thinking about for links I save in Safari or Firefox too - a script that scrapes the pages and builds a tree to organize links automatically).
Maybe a tool that could navigate the files, extract headlines and content to figure out relationships out of the files would be ideal to help me compose new documents and posts.
However, some of the most important concepts covered in my post are databases, relations between those databases, and editing workflows (markdown, slash commands). As far as I can tell, OneNote does not even attempt to do these things.
Some things that ease my mind:
Notion's export and backup features are quite good. You can export everything in various formats, markdown being the most relevant for me. This only becomes messy with complex databases that use formulas and are related to each other. You would not have these with local solutions in the first place.
Notion is already profitable and growing rapidly, I do not see them shutting down in the foreseeable future. However, the other concerns are relevant.
The API is supposed to be released soon. I intend to either build a backup workflow myself or use other tools that will get developed then.
I looked at their options before using Notion and was happy they had a full export available - but then I actually went ahead and used it because they are putting zero focus into improving the performance of the apps - and the export format is a horrific mess!
They split out code blocks from your notes in separate notes, the filenames are a mess and images aren't saved inline (as they could be with something like base64).
For workspaces of any reasonable size, you're going to be doing a lot of work to make the export actually useful outside the context of Notion - which should be obvious given how they treat things as blocks (and all the abstractions that come with that make it into their underlying data structures)
They've also been saying the "API is coming soon" for about a year now.
But yeah, lack of API is a major problem, and main reason I haven't even considered trying it yet.
Edit: I guess you meant human-readable as opposed to things like base64, and not as a critique against Markdown.
By now, there are several unofficial APIs at least, for read-access they are quite good apparently, did not yet use one myself though. E.g. https://github.com/kjk/notionapi
My post describes how I use it to draft my blog posts, where I also use the Markdown export feature. The code blocks are fine and within the same file, so I am not sure what you mean there.
Images are not ideal, but I don't know how it could be done much better. Everything is exported into a .zip and the images are also there.
Admittedly, my long term hope for a reliable and automated way of backing up everything relies on the API. Their promise that it will be released soon is a little awkward by now, but I am sure that it will come eventually. I don't think its true that it is already a year since they first said it would be here soon.
I've tried to optimize things so that while working, I can get things jotted down with close to zero resistance. If the barrier of entry is too high then I don't bother doing it.
That means often being able to spawn a new terminal, write a command and be done with it. Or to be able to pipe something from my clipboard into an auto-dated file.
I ended up putting together this ~15 line Bash script which seems to do the trick: https://github.com/nickjj/notes
I've been using plain text notes since 2001, although I only recently created this script. So far so good.
At one point I was trying to do a Luhmann's Zettelkasten with one note per txt file, internal linking, etc. I concatenated every piece of txt content I had. It's much simpler.
I have no system anymore. I try to create internal "hashtags" that are ctrl-F searchable. The goal is having something that's always open and at hand. I got plugins for autosaving when the window loses focus, autoloading when changed in disk and it syncs to Dropbox.
I cobbled together in Notepad++ (which has a GUI for this) a small syntax coloring definition that sort of matches my spontaneous habits from old salvaged text. I intend to let this evolve too. E.g. I'm using {braces} to have a little collapse button that hides pieces of text.
Among all the organization/notetaking methods I've tried, this is really the best. No fancy app in the browser, but really effective and fast.
I've written a small article about it here:
How do you access your notes on your phone? How do you edit?
That's basically what this script does except it uses separate files. You run `notes something really cool and important` and it creates a YYYY-MM-DD.txt file for you in a configured directory and appends the arguments to the file on a new line. So every day you get a new file.
Or, if you just run `notes` it opens your configured $EDITOR for that specific day so you can go free form.
I dont use ST3 for proper dev as much cause I use JetBrains IDEs more.
I still would prefer something better for my more permanent notes.
The less temptation to play with CSS, formatting and markup the better. One day it'll all go in a self-contained minimal wiki. Maybe.
On my android phone I use an ancient app called "unote" (installed from an .apk which I keep, it's long since disappeared from Play). Simple, hierarchical plaintext notes, kept in a local sqlite file. Sadly no search, but that just means I have to be disciplined about what, where and pruning.
Academic papers, standards documents, white papers etc... get downloaded as PDF and renamed according to a (human-readable) standard scheme that I use (year, title, [author(s)], [paper-type]), and stored in a 'meaningful' filesystem hierarchy. Podcasts and some videos are also stored in the same hierarchy.
Documentation for the software libraries and APIs that I use is downloaded from readthedocs (where possible) and stored in a parallel system that takes account of versioning. (So I can concurrently store differently versioned copies of the documentation for a single library).
I have a simple python script that iterates through my directory hierarchy and produces a sqlite database and a couple of xlsx files with a row-per document (one spreadsheet for reporting document-management metadata and another that allows me to assign labels and write precis notes). The script also extracts the content of the PDFs as plain-text and feeds some NLP tools that I'm playing with.
I use the spreadsheets to keep notes on the files, and the act of manually renaming and sorting the PDFs into the 'right' place in the hierarchy helps me to understand what's in them and remember what I've got. (I'm constantly reorganising the hierarchy as my understanding develops and evolves. My python script keeps everything -- notes and documents and other metadata - in sync).
So far this has scaled OK to around 26,000 PDFs.
Ps. Another question: can you tell which nlp libs are best for this purpose in your experience (and how do you eventually search the generated index?)
I have quite a big system myself, using bibdesk as my interface to filesystem, and searchability would be very nice indeed. Atm i only use default system tools like spotlight (macos) or mdfind. More custom nlp solution, your post inspires me to think more harder abt that.
Has anyone tried out using a personal database like this?
I played with it briefly a few years ago, and just tried it again now. It's definitely not geared towards note taking the way other apps in this thread are, but it could have value as a self-contained and portable wiki/kb. I'd be interested to hear of people's experience with this, either for its main purpose or even just as a note taking system.
I started with excel a lot of year ago, migrated to access and now i'm trying nixoxdb for android (the only other acceptable alternative was mobidb).
The main problem is the ease to use, you need to create a table for everything you want to store, and nothing support an hybrid free-form AND structured data(while also been able to store files), so the solutions with good text editing are terrible at structured data and viceversa.
Other problem are the ability to sync or even open the database in more than a couple of platform and the ability to access the file embedded in the database from other application (there are others but this are the one that bother me the most).
I'm also evaluating notion, but i really don't want a cloud service since i store basically everything inside my main database and the ability to access years from now is a MUST.
Also i still not have a good solution for my email that are stored mostly in thunderbird and only some message are "exported" to my main database.
I'm thinking of building something myself based on sqlite, .net, dokan/fuse and something similar to syncthing. But it's a big project and i don't have a lot of time
edit: couple of fixes
There were two things I didn't like about it. I realized that browsability is really important - an SQL database works well if you know what to query, but it's just easier to have an old Yahoo-style index. I went with a wiki approach instead. The other thing is that it's easy to use version control with a pile of text files, but not so much with a database.
These days I do have some things in a "database" but that is a big text file that I query using Linux tools like grep and sed. Does everything I need and plays well with version control.
Having Notion, and also other alternatives, this does not seem viable anymore to me. For example, when I want to edit the "Post" column of my bookmarks database, where I relate blog posts to bookmarks, I just start to type and Notion will suggest matching posts. I can not think of any SQLite frontend that offers such usability features, which is not surprising, as they are clearly not built for using them as personal knowledge base GUIs.
The reason why one might choose not to do this is that a) it takes a bit of work to get syncing across devices, and b) it takes a bit of work to get markdown, images, search, videos etc working.
It's a self-contained Wiki/Notebook/Journal with tags, it works instantly, and all of it is in a single HTML file with magic JS in it. Or you can use it with a server.
However, I do not see how this compares to Notion. How do I edit this HTML file on my phone? How do I add bookmarks to it via a hotkey? What is the equivalent of Notion's databases? Does it have Slash commands? This list of questions could go on for much longer. I really don't want to sound like a Notion shill, especially since I also use other tools (e.g. GitLab) for my knowledge base, but many of the suggested alternatives lack basically all the features that I describe in my post.
>How do I edit this HTML file on my phone?
Open https://tiddlywiki.com/, click the pencil icon on any of the entries. Edit away. The edits can persist with a self-hosted solution. Or you can just download the edited HTML file.
>Does it have Slash commands?
You probably mean "keyboard shortcuts". Yes, a plenty. Get started here:
https://tiddlywiki.com/static/KeyboardShortcuts.html
This doesn't cover all of it; there are shortcuts in the editor (Ctrl+B, for example to make text bold), and they are fully customizeable. Everything can be customized.
>How do I add bookmarks to it via a hotkey?
TiddlyWiki is tag-based. You can add tags with a keyboard. Tags define all the structure (table of contents and search).
>What is the equivalent of Notion's databases?
I am not familiar enough with Notion to answer this. If you ask "can I do ____ with TW", I can tell you.
>but many of the suggested alternatives lack basically all the features that I describe in my post
Sure, because they are different products. Feature-by-feature comparison doesn't make sense; if you want Notion, you need Notion.
If you want to organize your knowledge, Tiddly Wiki is one very fine tool to do that.
TW is not a CRM, it is not a Calendar or Google Sheets alternative (as Notion claims to be). It is not an all-in-one tool.
What it is, it's a personal knowledge base tool - and that's what the title of your post says. At that, it can do better than Notion. Or worse. User's call.
>This list of questions could go on for much longer.
The same can be true going from the other side. Here is one:
How do you access your knowledge in Notion 10 years after Notion-the-company goes the way of the dodo?
It basically means that you can access all kinds of features by just typing ahead after a '/' without knowing a precise shortcut. E.g. if I want my text to appear blue I would type /blue. I would also get there with /color or even by just starting to type either of those words and Notion will make suggestions. Slack, Confluence, and others do it in a similar way.
I think I do understand Tiddly Wiki works, as I have quite a bit of experience with Wiki systems, didn't use TW in particular though.
Some of the features you put aside as not needed for personal knowledge management are actually very nice for personal knowledge management. You might want to look a bit into my post to see what those databases can do ;)
> How do you access your knowledge in Notion 10 years after Notion-the-company goes the way of the dodo?
This is a very valid concern, that was also raised by some other commenters. In my opinion, Notion's exporting capabilities are quite good and I use them regularly. However, I do hope that we will get more automated and sophisticated third-party backup solutions when they have an official API.
Other options do things like public hosting.
Personally, I just use the HTML file with the Timimi Firefox plugin installed for autosaving the current file.
> Don't attempt to use the browser File/Save menu option to save changes (it doesn't work)
It looks like TiddlyWiki doesn't recommend that exact method anymore.
You "Save As.." that.
To save changes, click the red checkmark, and "save as..." the updated HTML over the old copy.
Switch to other possibilities if you use it enough to justify that.
I, a lifetime worth of knowledge in reasonable HTML of Tiddly Wiki would take less toll on your PC then the current CNN front page.
https://github.com/djmaze/tiddlywiki-docker
if you want to host on arm device, https://github.com/djmaze/tiddlywiki-docker/issues/15
Big thumbs up from me so far.
The hardest part was setting up a CouchDB instance for the backend, but I found some instructions and fiddled with it until it worked.
I have a single file called "engineering-notebook" and I store pretty much everything in it:
- bookmarks to resources
- personal notes
- personal tutorials/step-by-steps
- cheatsheets
It works fantastically for me
For Links/Bookmarks I use Google bookmarks - https://www.google.com/bookmarks/ which surprisingly has not been killed and looks like might be left alone by Google. Other bookmarking services delicious.com & posterous etc all died in due time.
By using browser extensions and a 3rd party Mobile App. I am able sync/store all the links. Though I would love to be able to sync bookmarks to my org-mode notes. Looking at org-capture now to see if that could possibly work (https://github.com/alphapapa/org-protocol-capture-html)
Chrome Extension - https://play.google.com/store/apps/details?id=com.amlegate.g...
Firefox extension - https://addons.mozilla.org/ru/firefox/addon/fess-google-book...
Android App for Google Bookmarks (3rd Party) https://play.google.com/store/apps/details?id=com.amlegate.g...
I use a separate file for non-technical stuff.
TODO alone is a game changer imho.
Notable's categories (tags) for me include things like manpage snippets, code snippets, workspaces/scratchpads, thoughts, etc.
Of note, a lot of the things in my head that I need to jot down on my mobile, even at length, are usually stored in Simplenote[1] then transferred to Notable if they're important enough, via Simplenote's desktop app.
I would love a system that treats all media the same as just a source of knowledge instead of being hung up on the source type i.e. is it a video, or a book mark or a pdf....
It doesn't handle video/bookmark/pdf. If I have to save those, I put them in a separate directory and make a note of it in the text file.
It's pretty well organized, thanks to org mode. It has a section that is by date, like journal entries, and a section that's by topic. Since it's just a single file, it's very easily searchable. It will never "go down". I don't need to run a database (I tried using wiki software for the same purpose). I can even "link" different sections, by having plain text labels. For example, I can refer to "ZFS Setup 2009-01-01" which is another place in the document I can search for.
I can "jump" to sections by using the orgmode annotation. Searching for "* Computers" will go to that top level section.
I liked this system so much that I also use it at work. People are very impressed that I can find things in my "notebook" so quickly. It's just text search.
Websites disappear. Web apps disappear. Apps disappear. My text file, does not.
I've used a similar system in the past (though not based on org mode), but was too brittle if/when "links" broke or were altered. For an over simplistic example, if the references that you made to a video within your master text file(s) breaks because the video's containing folder name has changed even slightly, well that becomes annoying. For a small number of references, types of destination media types, etc., maybe not an issue, but after some point, it can become too much maintenance/annoyance. Is this an issue that you encounter? If so, do you have mitigations for avoiding/resolving the broken links? Just curious.
P.S. It's also relevant to bookmark management as well: you get to see not only notes on the bookmark, but also on the 'child' URLs, e.g. if you open someone's twitter profile you'd be able to see their tweets you've favorited. Or if you open some blog, you'd see posts from that blog you saved on Reddit.
There is not much support for PDFs or videos though.
Together with some plugins for managing tasks, git and some script for synchronizing repositories, it has been working great.
There are some limitations on this approach (ie, two persons editing the same page concurrently is a no-no), but for my use-case it works perfectly.
I find it much more intuitive than org-mode, and the fact that it auto-generates a "global" task list based on items spread across the whole notebook makes it much easier to prioritize things.
I'm hoping (and optimistic) that the distributed web brings with it some epiphanies about how to do better local knowledge management!
More specifically in the section Changing results on the fly
Killer feature being the encrypted sync via more or less any kind of storage back end (where presently I use Box via WebDav, but that's easily switched). Seemless work across all my devices, all data in nice and portable markdown.
It fits my needs so well that I have taken the extraordinary step of allowing the desktop version on my pc's, even though it is Electron based.
A tool is great for collaboration, but how much do you really need for a personal KB? And how much effort do you want to spend migrating?
Then there's vendor lock-in, learning curves, etc.
A well-organized directory structure with plain text files gives you a data store that won't become obsolete and can be used on any OS. I use CLI tools to search (e.g., grep) and that's all the features I need right there.
Jupyter has helped me take notes when I need to "think" quantitatively, and be able to look at answers quickly, plus maintain a record that I can read later on.
Jupyter notebooks are great for capturing thought process and showing intermediate state of whatever it is you're doing, which means coming back to them even years later it's a lot easier to recall your own thoughts and not have to reverse engineer a monolithic block of code.
Most of my stuff is on dropbox/google drive for accessing them across different devices. I prefer tools that let me control where the data is and not the other way around.
For all academic papers, documents, I use Zotero. you can use your favorite pdf reader to annotate or take notes on pdf's and zotero will sync these. I also love the feature where Zotero can automatically extract all annotations from the pdf. (I some times save these as an org file)
If I am reading a longform web article or a blog post that I really think is useful and helpful, I also save these to Zotero. The push to kindle extension from fivefilters is an awesome tool that converts webpages to pdf.
I'm currently testing the memex extension from worldbrain to annotate and organize my browsing history.
For all notes, journals, random thoughts, ideas, (both work and personal), I use orgmode. (I recently switched from ZimWiki). Its been amazing so far. So many things are easier to do orgmode although the learning curve for emacs is pretty steep in the beginning.
On mobile, I use the orgzly app for accessing and taking notes! Its by far the best android app I use so far.
practices >> tools!
I also found https://dabble.me/ the other day which seemed like a nice slot in for quick thoughts with the ability to email to your journal. I've heard that journaling manually works better but I really dislike writing unless necessary.
One of my biggest gripes is with how many places I have my data in currently, and as a dev, I should have done something already but that's another story.
I'll definitely have a look at Notion and org-mode having read the feedback here along with a few other mentions elsewhere.
Don't use bookmarks as a 'I might come back and read this later'. Use them as 'I will need this later'.
I've used Onenote, Org Mode in Emacs but I am very happy to have discovered Bear Notes (https://bear.app/) last. year. It has really changed how much I document, work with text, store info etc.
Bear handles images, lists etc inline. It allows for very easy tagging (creating hierarchies) etc. And it looks really nice. Even with good fonts, Org Mode never really looked pleasing to me [0]. And I found that easthetics is really important for me to actually use the tool.
Bear uses SQLite for storage, which can be accessed with any SQLite tool (I've tested it). And it can also do batch export to a number or formats. The exported files looks good and can easily be imported into other tools, for example Org Mode (I've tested it). Bear can also export to PDF with styles. I actually use this to export my notes as reports to the company board, customers etc. And I've gotten compliments on how good the reports looks. ;-)
Finally Bear syncs easily. I can now work with my notes on the phone, iPad and laptop.
If I only could get the tool to use a solid, non-blinking caret life would be a bliss. Go Bear!
Not looking good.
https://twitter.com/conaw/status/1214855473876201472?s=21
“Still working out specifics
Right now looking like it'll be $30/month - cheaper for annual - cheaper for students/non-profits/unemployed
I personally think unemployed autodidacts should always get student benefits - scam that they don't”
[1] https://emvi.com
[0] https://zettelkasten.de/the-archive/
My use case: before Google+ folded I saved all my history there, because 99% of this was links to pages on stuff I found interesting/fun/relevant.
So now I have a large amount of urls (some will probably be defunct by now but it is not relevant). What I would like to do is to feed these urls to something (DevonThink?) that could access the article, and index it so that if - as an example - I want to prepare a RPG campaign on modern day pirates I can just write in the search box ["pirates" "shipping" "modern"] and hopefully get a list of web articles that are relevant.
Now, I know that Devon has some sort of automatic indexing/clustering facility for documents you have on your HD, but it is not really clear to me if it works also with stuff you only have an URL to.
(If anyone has some alternatives to suggest I will be very interested - I toyed with the idea of putting together an ElasticSearch VM for this but it remained on the backburner for years).
Here is my original request here, btw: https://news.ycombinator.com/item?id=18882167
For my personal needs I use https://github.com/vimwiki/vimwiki for notes, ideas, todos, articles etc. I used Buku for few years, but then I realised that I'm not using any bookmarks at all, so moved to native Safari Bookmarks. :)
At the end I'm using GIT to have complete archive of each change. :)
For public notes, I use GitBook https://www.aizatto.com/why-gitbook
I've got quite similar system in spirit and aims, just instead of lower Notion layer, using org-mode. It ends up searchable, synchronised with all devices, available offline and with great tooling for organizing and processing information.
> While documenting my configuration, especially my command-line workflows, I identified some shortcomings
Can't agree more! I've considerably simplified so many things while writing up on my setups and publishing code -- I guess it's easier to sneak in unnecessary complexity when it's all in your head.
> They are, however, not well suited for keeping extensive bookmarks libraries.
Agree state of browser bookmarking is a bit of a shame, interfaces are restricted and bloated with non-functional features. At first I switched to Pinboard [0], but after a while I realized even that was not enough for me and wrote Grasp [1], browser extension to capture stuff directly in org-mode.
> Examples of things that should not be Chrome bookmarks
Yep, eventually reached exactly the same conclusion! I often want to add private notes, more context etc and it's just not compatible with standard bookmark solutions.
> Taking Notes
Also big fan of notetaking, I basically write down any remotely meaningful thought if I don't have time to exercise it at the moment (via org-capture). Also using same file for everything and just processing it now and then. Same for ideas for new projects, blog updates, etc -- these are just entries in org-mode.
> ..it requires some discipline. I tag and annotate new entries, that I added via the web clipper, about twice a month
Yep, doing same, going through clipped links in org-mode, tagging and putting priorities. Then I can sort by priority and start reading/refiling etc. Eventually non-important stuff just sinks down as I don't have time to process everything, but at least it's searchable.
> For example, it is quite easy to search through questions you have answered on Stack Overflow. It is also not a problem to go back to your Hacker News posts or search through projects you starred on GitHub.
I actually find that even though in theory it's easy, in practice sometimes you don't remember where exactly you need to search for something, so you end up going through 'check reddit saves', 'check HN saves', 'check twitter faves' cycles, etc. So I'm automatically converting these into org-mode and they are searchable as any other org-mode file. I describe it in more detail here [2].
P.S. Great and clean design, especially the sticky navigation on the right! I might borrow the idea :) By the way, I tried in responsive mode and it seems to disappear alltogether, perhaps you could display it on top if the screen is too small?
> Ad using native bookmark sources, such as HN/Twitter/...
You are right, this is not ideal and it also happens to me that I am no longer sure where to look first. On the plus side, I do find it eventually, even if it's only the second native source I check. Atm, it seems to me that this is still better than duplicating everything in my other bookmarking layers. I will read your posts to find out more about your system, honestly don't know much about org-mode yet.
PS: Very glad to hear that! Especially, since I don't do much frontend work usually. I will think of a way to include a ToC also on small screens. Please, feel free to steal any idea you like. Then I don't have to feel as bad for looking into some of yours. I really like the pilcrows next to the headings and the dotted lines for highlighting sections on your site.
(as an alternative to grasp, because I'm on MacOS/Safari).
Yes depend up server and paid subscription. But use tag and sync across multiple devices
It's flexible enough for re-organisation and I can convert to lots of formats with pandoc.
Their support team is great too, had an issue with an update using Brave recently and they were very helpful with fixing the error.
I have every interesting article or web reference I've ever read since 1997, sitting in a PDF file alongside tens of thousands of others.
It is so convenient to be able to "ls -l | grep subject" and find all the pages related to that subject, and then to mine the data out of the PDF's for further reference.
Or, I just open the PDF and gain the knowledge again.
This works very well and doesn't require an active Internet connection. Every year I spend a few weeks up in a mountain retreat, going through the collection and getting an idea for the variety of topics I've read about for the year.
Its pretty neat to see, also, the changes in my interests over the years - and as well, its pretty wild to go back to the sites after time and realise I've got the only copy of the site - because its gone now.
I'd prefer to see more tools for manipulating PDF's become mainstream in the future. I can't recommend highly enough the productivity I've gained by being able to organise things this way. Its like having my own private Internet of 80,000+ pages, tailored to all the things I'm interested in ..
I "print" to markdown instead, even though pdfgrep sort of works I much prefer text files.
A little friction, but if that's too much, then the article is not worth keeping. But yes, someone with the time on their hands could contribute to markdown-clipper, saving images locally.
pdftotext -layout -eol unix -nopgbrk $PDF | egrep ...
Many PDFs have compressed content streams, plain text utilities only see metadata in that case. Cached, compressed text-only output is usually tiny, and can be zgrep-ed.pdfinfo shows document metadata (title, subject, keywords and more), but it's quite uncommon for these to be useful (Adobe and LᴬTᴇX-sourced PDFs tend to have this data).
Both come with xpdf.
It kind of amazes me that you've been doing this for two decades already.
Another thing it allows me to do is generalise my reading in real time - I know I can always come back to the subject of the PDF in a few days time, or whenever really, without needing to know where to find the details: I just grep my file tree, and mine the data that way.
Key thing is, though, that when saving the PDF, I always check that it is named a decent subject, derived from the <title> tags .. if there isn't one, I add it myself. This is the only place I pay attention to 'tagging an article' - in naming the file of a poorly-named page, but if the page title fits the subject, I don't change a thing. And it means I have a huge set of data also in the filenames, not just the contents themselves .. Periodically, I use 'detox' on my PDFArchive/ directory, also .. this helps with consistent-naming for regexes, and so on.
I keep getting the itch to put all this into a DB with proper indexing or whatever, but .. actually using plain ol' shell tools is turning out to be just as productive. Its a big set of data, but plain ol' files and pipes is proving to be all I really need to get the data mined ..
https://github.com/luizdepra/hugo-coder/
My blog itself is not open source. I thought about this for a while, but I think the pressure to keep everything cleaned up, documented, and organized is not worth the benefits in this case. If you have some specific requests, feel free to get in touch and I might be able to help you.
Its your life, and your code. You can post it, and keep it messy if you want to.