GNU Recutils
labs.tomasino.org
labs.tomasino.org
† I don't know, maybe they've already done this work and ruled out anything bad from the dozens of unique crashes AFL trivially finds? I wouldn't want to pretend that I've done any serious inspection here.
I vaguely remember doing this with wget, there was a way to make it think the terminal's width is (unsigned)-4, then when printing the download status to stdout, it clears a buffer with a memset(ptr, ' ', -4). Of course -4 in this context is a huge number. It overwrote its whole self until segfault. (this issue was fixed, btw)
great learning experience, for anyone who knows enough C to understand what they're looking at.
There you go: https://github.com/google/AFL/blob/master/docs/QuickStartGui...
Sometimes the manuals seam clear, but when you actually want to run the program you discover that it needs a library, or that the directories must have some specific names, or ...
Recutils has MUCH more to offer beyond the basic intro I gave here. It has wonderful org-mode integration for you emacs people.
Here's a recfile of my read books for reference. I generated this from my Goodreads export csv and a few recutils calls: https://ttm.sh/Equ.rec
Or maybe link some resources you found most helpful?
Thank you!
setxkbmap us -option compose:rwin &
Then just press [Right windows key] + [' ] , [a] in order to type in an 'á'.`guix search`[^] outputs data in recutils format, so if you are searching for a database driver for Python, but want to filter out "python2" variants, and ignore uninteresting fields such as versions or dependencies, you can do:
$ guix search python mysql | recsel -q 'python-' -p name,synopsis,homepage
name: python-mysqlclient
synopsis: MySQLdb is an interface to the popular MySQL database server for Python
homepage: https://github.com/PyMySQL/mysqlclient-python
name: python-pymysql
synopsis: Pure-Python MySQL driver
homepage: https://github.com/PyMySQL/PyMySQL/
name: python-peewee
synopsis: Small object-relational mapping utility
homepage: https://github.com/coleifer/peewee/
Without recsel, the output is 100 lines long, with lots of duplication between the Python 2 and 3 variants.[^] GNU Guix is a package manager that works on top of any GNU/Linux distribution.
OP/blog author: Great post, thank you for sharing your experiences
I've been sloshing in the text soup of Tiddlywiki Sqlite3(csv) VimWiki OrgMode Mediawiki Freemind(xml) and now it looks like Rec is the next ingredient to experiment with.
FZF and moreso Ripgrep https://github.com/BurntSushi/ripgrep has been really great to add to the mix.
In fact, you can already do the querying part using https://stedolan.github.io/jq/, I believe you can make modifications with it as well, but a different front end a la recins/recdel would make that a bit more convenient.
[{"title": "Opening",
"lede": "What happens when we \"open\" a file?"},
{"title": "ltrace",
"lede": "It turns out that ltrace uses the \"ptrace\" system call."}]
or title: Opening
lede: What happens when we "open" a file?
title: ltrace
lede: It turns out that ltrace uses the "ptrace" system call.
? When you get a parsing failure after you edit it (or after a sector on your SD card gets a read error), which one do you think will be easier to fix?Binary tends to be used for big datasets recorded by instruments that are left in the field, unattended, for months to years. Since every byte counts, these instruments cram information in very tightly. The binary nature of the files makes them a pain to deal with, but it also confers an advantage: the files are very seldom corrupted by a person who thinks they are benignly viewing the data.
The text files, on the other hand, sometimes get extra junk inserted because someone in the data analysis pipeline thinks it's OK to look at information in MSword or MSexcel.
Sometimes, opaque binary data formats are superior, in terms of data integrity.
I thought of this, while reading about recutils and thinking of a contrast with sqlite. Recutils looks great, but if I started sharing data in that format, I bet it wouldn't be long before derived versions of the files had become corrupted, as someone edited with MSword.
In this forum, people will sniff at people who use MSword, and I have done so, myself. But, it's a simple fact that some people who are good at one thing are not good at another. Some of my colleagues who use MSword for every silly thing (e.g. seminar announcements) are actually very good at their subject matter (e.g. the science talked about in those seminars).
I have a tool called yaml-to-sqlite ( https://github.com/simonw/yaml-to-sqlite ) which converts a YAML file into a SQLite database, which I can then use with Datasette ( https://github.com/simonw/datasette )
My biggest project with it so far has been my site https://www.niche-museums.com/ - a guide to small and niche museums. The museums themselves live in a single ~100KB YAML file in GitHub: https://github.com/simonw/museums/blob/master/museums.yaml
I have a CI script which builds that YAML file into a SQLite database and deploys it + Datasette + custom templates to https://www.niche-museums.com/
I've been running the site like this for a few months now and I really like it. I love having my content in source control, I find editing the YAML to be reasonably pleasant (I even edit it on my iPhone sometimes using the Working Copy app) and any YAML errors are caught by CI before they are deployed.
GNU Recutils is a nice metalanguage for data. You can churn out CSV from it which in turn can be \copy loaded into PostgreSQL. Maybe its too many steps, but I found it an important format for creating the logical model of representing data without predisposing it to any other technology or encoding.
Unfortunately its org-mode integration with emacs is broken when using spacemacs-- to far for me to fix as I'm no Elisp wizard.
That's actually the main selling point I am seeing here.
Is there are any better format (apart from yaml maybe?) to collaborate on datasets, that can also immediately be used with code?
Does anyone have any more resources (besides the manual of course [https://www.gnu.org/software/recutils/manual ]) they can recommend about Recutils and Recfiles?
This, or SQLite could be really useful to embed data into pages.
[0]: yeah, I've had to dig beneath the surface of Confluence lately and Confluence has this weird property of immediately making me motivated to try writing a decent wiki
At this point I've recently read someone from FSF(?) saying this is only MongoDB and others misinterpretation so at this point I am just utterly confused.
Also, when somebody uses AGPL that usually means: we found the scariest license we could find while still calling it open source, but, we have a commercial license to sell you.
However I couldn't find any licensing option. Does this mean it isn't just a way to sell commercial licenses?
I'm completely honest here. I really don't get it, but then again it took a while before I really understood the GPL as well so I'm ready to be enlightened :-)
I went ahead to study the FSF FAQ but they don't really answer it completely as far as I can see. The clear cut answers are:
- if you combine your program with an AGPL program then your program has to become AGPL as well, just like the GPL.
- if you use an AGPL program unmodified it doesn't seem like you have to distribute kt to users who use it over the network
But as far as I can see the FAQ doesn't say what happens if the AGPL program reaches out to other applications to get data. For some reason I always though anything that was touched by the AGPL program, either over the network or otherwise would have to become AGPL.
If that isn't the case - and there is more and more to suggest that, then I think FSF should point that out clearly.
Then I forget about that idea 10 minutes later when a new episode of my anime watchlist gets dubbed. I wonder how much that has cost me.
I've been using emacs for a couple of years now and every couple of months I find something new and huge buried in it that surprises the heck out of me... how can these things hide for so long!
From your list, I'm guessing that you like anime that makes you think, and aren't turned off by story arcs that are overall dark and depressing.
The first thing that comes to mind is "Death Parade". It's currently available on Funimation, not sure where else.
A few more suggestions are "Elfen Lied" (although it can be a bit gory) and "Black Lagoon" (not super dark, but very entertaining imho).
Note: I only watch English dubs because I can't follow the action while also reading subtitles. If you're open to subtitled anime, I'm sure someone else could give you a much bigger list of suggestions.
I did watch the subs for the latest Attack on Titan because the dubs weren't out yet. But I was already invested enough in getting answers to that shows many questions as soon as possible.
It's a shame there's not much Deathnote level stuff. I'm so invested in that particular execution of a cat/mouse detective story with well defined rules and Sherlock vs. Moriarty levels of intelligence that _actually convince_ you that the characters are geniuses. Watchmen also comes to mind.
Any recommendations for non-anime that fits that genre? Novels, shows, comic books, etc?
For example, the film "A Beautiful Mind" is about character who's clearly a genius, but I don't recall the script showing signs of its writer having rare intelligence.
Whereas stories like Deathnote have some plot twists that make me think the author is highly intelligent, regardless of the IQ of most of the characters. I'm guessing I'd put a story into this category if there were plot revelations at the end which in retrospect I could have figured out my self, but didn't until the moment chosen by the author.
Paranoia Agent: https://myanimelist.net/anime/323/Mousou_Dairinin?q=paranoia (this is amazing)
Samurai Champloo: https://myanimelist.net/anime/205/Samurai_Champloo
And of course, GITS: https://myanimelist.net/anime/467/Koukaku_Kidoutai__Stand_Al...
Ah, then you'll love Monster: https://myanimelist.net/anime/19/Monster
Like you might expect, this is quite poor. CRC32 for authenticity, mac-then-encrypt, no binding between keys and values, using low-entropy passwords directly as AES keys, and a fairly trivial looking read overflow in the decrypt function. That's just two minutes looking at one source file.
$ sudo guix processes | \
recsel -p ClientPID,ClientCommand -e 'LockHeld ~ "perl"'
ClientPID: 19419
ClientCommand: cuirass --cache-directory /var/cache/cuirass …There are so many usecases where this would fit much better than traditional approaches.
Thank you for sharing!
recsel contacts.rec | while readrec
do
if [ $Checked = "no" ]
then
mail -s "You are being checked." ${Email[0]} < email.txt
recset -e "Email = '$Email'" -f Checked -S yes contacts.rec
sleep 1
fi
done
0. https://www.gnu.org/software/recutils/manual/Bash-Builtins.h...If I want human readable files I think I would opt for just a bunch of markdown files in a directory structure before I go the recutils route. If I want a lightweight database I think I would go for SQLite like others have already mentioned. In what situation do you really run into that you need human readable/editable referential integrity?
You can do this with SQLite but it would require CSV export:
The idea is to take full advantage of plain text, just like, markdown! You could even use git for versioning.
Just like markdown provides a quick way in that to renderize something pretty, but is still readable in plain text, recfiles do the same for data, it has data types, keys, integrity, and is easily queryable with its tools, and also easily read in plain text.
Sure, you could use some type CSV, but that is not pretty to read as recfiles. The point is the same as markdown, usable as just text, but with good tooling around it.
with markdown files in a dirs, you would have to provide you own tooling.
Let me be clear, I get that you don't have db operations with a bunch of markdown files with a directory structure, but on small scale data repositories do you really need that? For example, if I have a bunch of recipes in my "personal database of markdown files" I can quickly find the chili recipe I'm looking for by going to the directory "Soups & Salads > Chili > Slow Cooker Chili Recipe" or something along those lines. Or I could have grepped for "slow cooker chili". Either way I'm going to find that recipe with no extra tooling on my part.
Where do rec files features add value? Plus, if I use rec files now it seems I have to build my own formatting because I can't rely on markdown editors to build it automatically for me. Or is there a way to specify formatting for rec files?
You could definitely use markdown within fields of a recfile. There would be some multi-line syntax you wouldn't be able to use, because it would conflict (e.g. '+' unordered lists), but other than that you could mark up text with Markdown, no worries.
[0] https://www.gnu.org/software/recutils/manual/Templates.html#...
1. Reference data (countries, car makes and models, zip codes, tax rates, etc)
2. “Human scale” databases - ie stuff that a small group of people are manually curating
O(n) would probably be fine for either of those.
An alternative is using JSON files with tools like jq, but Recutils looks much more powerful.
EDIT: nice, Python bindings https://github.com/maninya/python-recutils
[0] https://savannah.gnu.org/projects/recutils/
[1] https://git.savannah.gnu.org/cgit/recutils.git/tree/python
Why?
http://www.strozzi.it/cgi-bin/CSA/tw7/I/en_US/nosql/Home%20P...
your database is a human-readable text file that you can grep/awk/sed freely, and a line-oriented structure makes it perfect for version control systems.
But in fact it's not great for grep/awk/sed/, etc. For those tools to work well, you'll find you need to keep each record on its own line. BEGIN {
RS=""
FS="\n"
}
{
for( i=1; i<=NF; i++ ) {
nfields = split($i, afields, ": ")
if( nfields != 2 ) {printf( "Bad fields count: %s %s %s\n", NR, i, $i) | "/dev/stderr"; exit 1 )}
field(afield[1]) = afield[2]
}
}
(Untested, but should be generally accurate.)That parses records based on blank lines, into fields based on lines, splitting the fields into individual data recoreds based on ": " as a regex.
For more, see the GNU Awk User's Guide:
https://www.gnu.org/software/gawk/manual/html_node/Multiple-...
Never heard of recutils before but the on-disk format looks compelling. There has been limited (no?) takeup however which makes me fear recutils are intrinsically broken
It's weird that the FAQ page linked there [1] doesn't seem to link back in any way to recutils' own page. The only way to get there [2] seems to be to click on the "Software" header link and Ctrl-F for recutils.
[1] https://www.gnu.org/software/recutils/faq.html#whyturtles [2] https://www.gnu.org/software/recutils/
I was so focused on "human readable" == "text editor" that I hadn't thought about GUIs in the more "G" (graphical) sense.
Does some sort of general data compiler/decompiler exist? Please don't say Binary XML.
Not trying to be funny, genuinely wondering.
Example: we operate a pipline that pulls datasets from multiple government agencies and merges them. The sources are updated periodically and updates frequently introduce new inconsistences. To reconciliate, we use auxiliary datasets that record tweaks to each source, such as entries to add, mutate or delete.
In our implementation of the pipeline the source datasets, the tweak tables, and the resulting merged dataset are kept in Sqlite for ease of processing. The pipeline writes out dumps of each table during processing, even for intermediate stages. When running the pipeline, I can scan the diff and decide whether the changes are reasonable or tweaks must be introduced or removed. Once I'm satisifed with the result, the dumps are then committed to record the current known-good state.
When somebody wants to know why one record in the result is the way it is, I can determine how it changed in the source data, how it was tweaked, and what the result was before and after. It's really easy to produce diffs between revisions. Code-reviews are conducted over changes in pipeline code, validation logic, source and result all in one step.
If you read the documentation 'Purpose' [1], it says, the issues with databases, like SQLite:
- The stored data is not directly human readable.
- The stored data is definitely not directly writable by humans.
- They are program dependent.
- They are not easily managed by version control systems.
I just want to say that I also love sqlite, but I cannot avoid worrying that when the asteroid strikes and the ocean men invade, I will not be able to program a sqlite driver on a homemade 8 bit computer.[1] https://www.gnu.org/software/recutils/manual/Purpose.html#Pu...
I like the rectools for educational purposes, but beyond this for any however small purpose project I'd rather switch to sqlite pretty soon.
Especially if you absolutely need 100% reliable compatibility between many different devices.
Even more so if the format need to be part of a standard and/or a contract.
- It's almost a standard due to licensing.
- Tiny, really tiny.
It would be simple to implement an append-only writer in Recutils. I don't know where I'd start if I wanted that for Sqlite. I don't think the format allows it.
Also, why are you ignoring the other points?
But if you're working with a modest dataset where volume and performance aren't a primary concern, working straight in text is infinitely more convenient. There's a reason people do so much bulk data processing in CSV, JSON, or other easily manipulated formats.
https://www.ibm.com/developerworks/library/x-matters17/x-mat...
https://web.archive.org/web/20160510001507/http://dan.egnor....
https://web.archive.org/web/20100113205649/http://dan.egnor....
https://packages.debian.org/search?keywords=xml2
We used it heavily for triaging into log files and such.