Storing knowledge in a single long plain text file
breckyunits.com
breckyunits.com
Overall it seems like your ideas would make more sense if you used more widely adopted language to describe it. "Concepts" are records, "measurements" are fields.
I agree with you. More details will come out over time but I wanted to keep yesterday's paper a single page.
> You mention in the citations that "Hierarchies are painless in our system through nested parsers, parser inheritance, parser mixins, and nested measurements." Nothing else in the article gives any hint as to what those things are or how your system implements them except nested measurements. It's unclear at all what a parser is in your system.
Below is a link to a web IDE we built. You can see parsers (on the left), and concepts (on the right). Nested parsers and parser inheritance are demonstrated. Mixins is not currently in that branch yet. Ignore the "cells" stuff at top (that turned out to be an unneeded division between lines parsers and word parsers).
https://jtree.treenotation.org/designer#url%20https%3A%2F%2F...
> Overall it seems like your ideas would make more sense if you used more widely adopted language to describe it. "Concepts" are records, "measurements" are fields.
Yes, concepts often map to records or rows. Measures to fields or columns. Measurements to the cells in a spreadsheet.
There are reasons for my terminology, that should become clearer over time.
:)
We are aiming at roughly the same problem. Our implementation has solved some important details.
It does support multiple pages but you can use just one.
It has a nifty feature in that you can divide the single file into virtual parts. They just have alternate backgrounds to tell them apart. And each virtual part can have a type for syntax highlighting (plain text, markdown or a programming language).
I've been using it for a few months now and it's my primary note taking / knowledge recording thing.
Even though it's web based, on Chrome you can save notes on disk so it works like a desktop app.
Each note is a plain text file so you can edit them in any text editor.
If you put notes on a shared drive (Dropbox, OneDrive, Google Drive etc.) you can work on notes on multiple computers.
It's also open-source: https://github.com/kjk/edna
I love the math block. Is there a way to reference a variable elsewhere, or fetch data online? Then you could build a little personal dashboard with it.
I was thinking about making math more like a mode i.e. make it available in every block type, as opposed to it's own block type.
Then it would be active in plain text, markdown and even code blocks.
As to data fetching - falls a bit outside of scope.
Mostly it supports multiple notes and it's a web app, not a desktop app.
I could build a desktop app but it would not offer almost any advantages given that Edna can also save notes on disk (that's how I use it).
You can use Chrome's "Install" feature to make it look act like a native app (it opens in it's own window and acts independently of the browser).
I wanted multiple notes and I didn't get why it was made as a desktop app first given that all functionality to implement it is available in a browser (well, Chrome).
So I forked it and added those features.
Been using it daily so it was worth it.
So Edna either stores notes in localStorage or in a directory of your choosing on disk.
In Edna you can right-click for context menu to switch between localStorage and disk.
If you ask: "how do the browser APIs work", you can look at https://github.com/kjk/edna/blob/main/src/fileutil.js
Basically, there's `window.showDirectoryPicker()` to ask user for permission to access directory (either read only or read write). And then using that directory handle you can read list of files, read / write files or create new files.
History not repeating but rhyming, I suppose...
Anyway, thanks for this. I've just added it to my bookmarks.
I've been thinking about more featureful markdown note taker that would support images and more.
I've started on such a thing but stalled. It's way more work. The good thing about Edna is that I spent less than a month adding the features I wanted to Heynote fork.
The current version is at https://notedapp.dev/ but don't use it for actual notes.
This seems nearly identical to Heynote[0], which was also on HN[1]. Even the example blocks share some content with that used as an example in the screenshot on the Heynote homepage (and I think in the app too)
[0]: https://heynote.com/[1]: https://news.ycombinator.com/item?id=38733968
Initial note storage is in localStorage. To switch to disk: right-click for context menu, `Notes storage` / `Move notes from browser to directory`.
Then choose a directory on disk and we will do one time migration from localStorage => disk.
You can then switch to another directory (some apps call it a "workspace"). Because why not.
Encryption is probably the next feature I'll add because I want to store secrets in my notes and I'll feel better if those notes are encrypted.
More docs: https://edna.arslexis.io/help
Multiple notes is pretty big addition. I loved the concept and implementation of blocks in Heynote but a single note was a deal breaker for me.
I've also added some UI like right-click context menu for discoverability, ability to enable spell checking.
And I'm really trying to optimize for speed of use, including speed of switching between notes.
For example you can assign Alt + 0 .. Alt + 9 as note quick access shortcuts.
By default I create 3 notes: scratchpad, daily journal and inbox and they get Alt + 1, Alt + 2, Alt + 3 quick access shortcuts but you can assign them to any page you want.
What I know about Codemirror I mostly learned by reading other people's code so I suggest that.
Specifically code of silverbullet: https://github.com/silverbulletmd/silverbullet/tree/main/web... (and a few other directories there).
It implements very advanced Markdown mode, lots of code to learn from.
Is “>” a special character together with space and new lines? He calls it a trick, why?
How do I add data with spaces and new lines?
Is “Parser” a keyword that you postfix to names of values? He writes “idParser” and then has a value in each observation that is named “id”
In our ScrollSet implementation, a measure definition (what you call a "data definition") is a subset of a parser. You will know something is a measure definition when you see a line starting with a word with a "Parser" postfix, and nested inside that definition is a line like "extends abstractMeasureParser".
Below is a link to a web IDE we built. You can see all of the measure definitions currently powering PLDB on the left. On the right, you can see a concept ("more data", in your terms).
https://jtree.treenotation.org/designer#url%20https%3A%2F%2F...
> He calls it a trick, why?
The current term of art is "Offi-side rule" (https://en.wikipedia.org/wiki/Off-side_rule). I never liked that term. I call it the indentation trick. But I am referring to the Offside_rule.
All jokes aside, I think the equivalent for this would be markdown not xml/csv/json/yaml.
I suffer from the same flaw as the author, a tendency towards grandiosity and fervor in describing my good ideas. So I'm in a good position to advise that he knock it off: people don't like that, and it will keep them from using your stuff even if it's good.
Which it might be, actually. The extreme simplicity of the foundation is laudable.
If the idea is truly good, the products built on the idea should do just fine.
My guess is that the attacks you draw will skip any basis in technical merit and land directly on the tone, proceeding on an emotional basis. We have an n=1 here with plenty of that behavior on display.
You'd like to believe that someone proposing Tree Notation for a project wouldn't be dismissed with "isn't that, like, the YAML for TimeCube guy?". But this is, in large part, how the world actually functions.
I don't mind the pushback.
If it wasn't for the pushback against Tree Notation, I never would have started PLDB. ("Learn to research properly", one commenter once said. And he was right. I think PLDB is the proper way to do research).
It's much nicer to get pushback than crickets. That means people are generously giving their time to consider the ideas.
Crickets is the worst. I should know, I mostly get crickets.
Alas, vscode will choke on it.
I have a project where a thin wrapper loads from a giant markdown file into a sandboxed iframe. That way you could paste code from an unknown source into it and play with the output and paste private data into it and it wouldn’t be encoded into a URL sent to a server, as making network requests and following links are blocked.
https://codeberg.org/ristretto/pages
notebook.md is huge, output in the project website, link to source in the README.
> For pragmatic reasons, it is best to split your data into 1 file per concept and combine concept files at runtime.And files are just key/values anyway.
A file is a very abstracted concept and it technically can mean a lot of different things depending on the file system.
However, it is a very good abstraction that's nearly universal and practically there is little to no reason not to use them to organize things.
"The M1 laptop was the first consumer machine I tried where the performance of this system wasn't abysmal." - https://breckyunits.com/building-a-treebase-with-6.5-million...
Thank you!
It reminds me of that scene from The Shining where the character writes the same sentence over and over again.
With AI I think interfaces will further bifurcate between "users" and "creators" and pretty much all of our "desktop" ui paradigms will be consigned to history in favor of structured collaborative text interfaces.
To comment on the appearance though since it seems a focus none the less... I appreciate the ideal of syntax sparseness, but in this case I feel like it loses visual salience in plaintext after looking at some of the .scroll files. It's difficult to recognize the shape and proportion of what the content will be when rendered. The applied meta content lacks visual differentiation in plaintext from the content itself. I don't think total spareness should be the sole goal here; Markdown for example isn't strong in plaintext just because it is syntactically sparse, but because it is sparse in tandem with not supporting applying extensible meta content to content - but this does.
This is not some new concept, however. I stumbled upon this concept two years ago with some dude promoting a "Vault" architecture, where you use a single "notion.so" table to store all your data. You create views from this data to separate topics. You'll then be able to "centralize" all you notion stuff in a single file; all while being able to link any two topics or more together.
What hit me is that I can export the notion table to CSV and then this can be fed into an AI pipeline that might be able to predict my tasks better (like code). Only problem was, a couple of months into this and the notion interface became completely unusable.
This can be done with a regular database. Though the views/interfaces to interact are not that easy to create. I didn't find an alternative (I tried airtable too)
https://notes.dsebastien.net/30+Areas/33+Permanent+notes/33....
then you need some syntax for the strictures of your use case.
and /etc is reborn.
How do you handle writes? It seems like an interrupted write process could corrupt a section of text, which could be difficult to recover from.
Given the example at "breckyunits.com", I don't see hashing information associated with each item.
Are you depending on git to prevent such errors from corrupting individual items? If so, then I would be concerned about gits propensity for data corruption [1, 2, 3, 4].
I wonder if adding some ZFS-like hashing and integrity checks would be helpful. Then, as it's one big file, it seems to act like a TAR archive [5], where you append to the end, but have to scan through the previous content to find what you want. If that's the case, then it may be viable to do copy-on-write [6], where information is never modified, but instead referenced with a key, and later modifications supersede older versions.
(Again apologies if this is redundant, I just had the thought and had to get it down. XD)
[1] https://superuser.com/questions/1253830/does-git-prevent-dat... [2] https://superuser.com/questions/1635797/what-if-git-reposito... [3] https://stackoverflow.com/questions/tagged/corruption?tab=Fr... [4] https://www.reddit.com/r/git/comments/oq9wph/power_outage_in... [5] https://en.wikipedia.org/wiki/Tar_(computing) [6] https://en.wikipedia.org/wiki/Copy-on-write
This ... Is namespace hell, and if you squint at the caching problem, it's actually an indexing problem, which is also related to this.
Caching itself is quite easy, just ask the designers of speculative execution at Intel :)
I guess I can TIAS but is it documented anywhere ??
Thanks! (https://github.com/breck7/pldb/commit/83ba14454ed80fa682c85d...)
There are at least two (2) existing prior art implementation that have done this for years, only better (as in, with better tools):
- recutils: https://www.gnu.org/software/recutils/
- ndb: https://9fans.github.io/plan9port/man/man7/ndb.html
Please, developers of everywhere, I beg you: please learn what came before you before reinventing the wheel, only triangular this time. Please take the time to appreciate that if it's so obvious maybe it's because of your ignorance and not your genius.
I am reading through the source and will have more to say soon. If anyone has any links to massive plain text datasets based on these (or other similar tools), I would appreciate more pointers.
I can tell you now (subject to change), based on my preliminary read through the source is that the two systems you mentioned missed some highly important details that I have presented in my paper, with order of magnitude impacts. Not to discredit them at all, rather I think my work gives them credit, in that they were on the right track, and we just have the benefit of some recent innovations, and get to stand on their shoulders (and the shoulders of others).
Edit:
I have updated the paper with a reference to Recutils. Thanks rakoo! https://github.com/breck7/breckyunits.com/commit/71b706d296e...
The added text:
GNU Recutils^recutils deserves credit as the closest precursor to our system. If Recutils were to adopt some designs from our system it would be capable of supporting larger databases.
https://www.gnu.org/software/recutils/
^recutils: GNU Recutils: Jose E. Marchesi
https://www.gnu.org/software/recutils/
- Recutils and our system have debatable syntactic differences, but our system solves a few clear problems described in the Recutils docs:
- "difficult to manage hierarchies". Hierarchies are painless in our system through nested parsers, parser inheritance, parser mixins, and nested measurements.
- "tedious to manually encode...several lines". No encoding is needed in our system thanks to the indentation trick.
- In Recutils comments are "completely ignored by processing tools and can only be seen by looking at the recfile itself". Our system supports first class comments which are bound to measurements using the indentation trick.
- "It is difficult to manually maintain the integrity of data stored in the data base." In our system advances parsers provides unlimited capabilities for maintaining data integrity.If I'm lucky, I'll wake up tomorrow to someone else pointing out another precursor I overlooked.
I don't know if "please do a thorough literature review before showing me things" is the right sentiment here.
My message is an echo to earlier comments of earlier posts that talked about a similar point: nothing is ever new, everything has already been done before. If we tell ourselves we're engineers, we should be studying what came before in order to prove that the new thing is indeed better.
That being said, recutils is the standard method of recording data in GNU, and ndb is the standard method of configuring stuff in Plan 9, a system that any proponent of UNIX mindset should know about. I'm not exactly talking about obscure stuff here.
If the author was following a paper-like presentation, the author somehow skipped the section listing relevant prior work. This is something every single journal enforces, as researching prior work is the very first step any author does when they come up with something.
My confidence is shaken considering the sparse “prior art” section links to Apple M1 as an example of “fast file systems”.
https://github.com/breck7/breckyunits.com/commit/61792237c0b...
A number of things have to best fast for this system to be enjoyable to use (at scale) and before the M1 no personal machine I ever tried came close.
That's perfectly fine, but that's besides the whole point.
The point is that between coming up with something and implementing it, there should be a step to check if anyone already did something similar.
The whole point of researching prior work is to a) don't waste time reinventing the wheel, b) leverage prior work to improve your own ideas, c) make better use of your time by doing meaningful contributions instead of taking a risk on whether you're ripping off someone else's work.
That's the absolute basic standard on scientific publishing, for example. If you pick up any paper at all, you'll notice that right after the introduction and summary you get a bibliographical review listing any relevant work that your peers already contributed. When anyone submits a paper, the reviewers can and outright do reject your submission if it fails to adequately contextualize the paper with regards to prior art and related work. One of the points is to ensure the author is not wasting everyone's time with a novel approach to the wheel.
More importantly, if an author fails to know what's already there, how can they tell their idea is any good?
Except that the title ("A New Way to Store Knowledge") is leaning heavily on NEW.
;)
Thank you for the feedback. I've updated the paper with some more links.
The language in which the measures are written in (currently called Grammar. I will like rename it to something like Parssers) is quite advanced.
The improvements over Recutils, the closest precursor I am aware of, have now been added.
The PLDB ScrollSet is now about 500,000 cells of information. Each cell is strongly typed and fully auditable by git. There is a high amount of signal in that dataset. It is an intelligent set of weights, and continually getting more intelligent. And it is read at runtime as a single plain text file and compiled to a single CSV (or tsv, json, etc).
All from using the system documented in the paper (and the advanced language for Parsers).
If you can point me to a similar database or similar scale anywhere in the world (plain text base, >10e5 size, git backed, strongly typed, hierarchical and graphical), I would be grateful as I might learn something.
I don't see it as a this-is-all-a-joke thing though, more tongue in cheek.
also I think one-big-text-file has a certain simplicity, like everything-is-a-file on unix (or more properly plan9)
You know it is when it starts like this: "...All tabular knowledge can be stored in a single long plain text file. The only syntax characters needed are spaces and newlines."
That's fundamentally the simplest way of storing text. And it's nothing new, yet people have long ignored that simplicity for much more complicated ways of storing text.