Using sqlite3 as a notekeeping document graph
epilys.github.io
epilys.github.io
Sick of people reading your journal? Do your worst, MOM.
Personally I use Emacs, org-mode, and org-roam for Zettel-style links between documents. Org-roam uses sqlite to store its links and index, but documents remain in plain text .org files.
+1 on org-roam, and may I add org-roam-server, which visualises that graph in the browser very nicely, in an interactive way: <https://github.com/org-roam/org-roam-server>
I had no idea indexing could be used this way. Very cool find.
My TILs site runs uses a GitHub repository where the notes live in markdown: https://github.com/simonw/til
Plus a build script running in a GitHub actions workflow that compiles the notes into a SQLite file using my markdown-to-sqlite tool and publishes the resulting SQLite file using Datasette to https://til.simonwillison.net - which gives me search and an Atom feed and suchlike.
The site has custom templates so it Durant look like regular Datasette, but you can run custom queries against it at https://til.simonwillison.net/tils
I can see the markdown format being powerful. But how do you parse it and create a document graph ? (or rather...what do u use to persist the document graph)
I have been trying to use Firebase on a side project of mine for this markdown -> graph problem
Presumably for each build, you’d parse the names of each doc to produce the list of nodes (and their paths), and then as you encounter the markdown reference, update it accordingly.
If you start afresh each time, handling updates would be trivial, but you’re at risk of destroying existing URLs when changing doc titles/reorganizing — so you probably want a canonical name (to generate the final URL for), and the symbolic names (to refer by in markdown).
And then eventually you realize you want multiple ways to organize your docs, so you start tagging, and then eventually you realize you want ways to reference groups of notes, so you’ll go ahead and implement hierarchal tagging, and you’ll finally have be able to treat your graph as a graph, instead of a forest (list of trees)
And finally you can simplify the whole thing to only reference tags (or parent tags), which may happen to link to a single doc, or multiple. Probably in the special case of 1 doc, the link directly takes you to the doc instead of some collection page.
A document store seems like a terrible idea, because you’re really not modeling a tree, and that’s probably the majority of your modeling problem — I imagine you’re trying to currently stuff a graph into a tree.
Extracting links from markdown and using them to populate some additional columns or tables at build time would be pretty straight forward.
The advantage is an easy-to-learn search syntax that is flexible and similar to Google: https://lucene.apache.org/core/2_9_4/queryparsersyntax.html In addition to a document field for your notes, you can manage and query any number of meta-data fields e.g.
Tom M* birthday:July
retrieves all notes that mention Tom with a last name starting with "M" as long as the notes also have a meta-data field called birthday where the
value must be July. (Try compare doing this with the SQL equivalent version of the query!)Maybe unusual but not completely alone.
I use a mix but none are complete:
- Notes on iPad lets me start (or continue) a note by tapping the Pencil on the locked screen. Brilliant. Hard to link notes though.
- Pencil planner is totally brilliant It integrates calendar and lets me write on top of it and in a close to magic way it shows day notes in week view and the other way around. It is easy to use but not dumbed down. Lacks photos for now though but I can live with that for now.
- Joplin. For anything that doesn't contain handwritten or images.
by the way, bibliothecula markdown works with embedding image links. You add them as a file on the note document and attach it's UUID in the markdown text.
For Chinese, Arabic and so on you'll need to a custom tokenizer which may or may not be available on your target platform.
Other than that, there's also tokenizing (splitting text into words) that's also unicode defined and stemming (reducing tokens to a base stem like "likes"->"lik-" in English)
What good rises out of including extra paperwork with each number, essentially saying "attention please, a random number", or "this is a timestamp, a monotonic counter and a MAC address", like if it was a newfangled smartphone? The RFC also says that there is no mechanism for validating a UUID, save for checking if its timestamp part is in the future, for versions that employ such a portion. Why can't any 128 bit value be accepted on the receiving end?
But, since that only holds true if people follow the rules, it's worth standardizing on either v1 or v4, and so it's good to be able to distinguish between the two for validation purposes.
It would be nice if there was cgi script that could serve a sorted and paged index with links to the html, md or gmi to a web or/and gemini browser.
Bear organizes with "live folders" based on tags added to the document, powered by sqlite.