Show HN: JSONlite – A simple, serverless, zero-configuration JSON document store
github.com
github.com
Instead, there are two tiny tests, and I found a couple of issues in 30 seconds of looking at the code.
EDIT: I should prove my point.
1) Race condition on calculating uuid (obviously won't be too serious for a good uuid implementation)
2) No check that file is written successfully.
3) No check when json is invalid, just silently swallowed, or filesystem full.
4) Makes one file a json file, will scale terribly past a few thousand json files.
Not sure what you mean exactly by 1 and 4 though.
The usual workaround to that is to create subdirectories based on the first few characters of the filename.
Since you're using bash, you could do something like:
$ echo 1df8be33-4392-471a-99af-3df967b87cb6 |sed -e 's/^\(.\)\(.\)/\1\/\2\/\1\2/'
1/d/1df8be33-4392-471a-99af-3df967b87cb6
Edit: You may want more than 2 levels of directories, or directories with 2 character names instead of 1, etc. Apache's mod_cache defaults to 2 levels of 2 character names, but the filenames are base 64, so more possibilities than hex.With 4, I find once a directory has about 100,000 files in it, things get bad, from the simple (ls * won't work) to the nastier (git starts using huge amounts of space, you'll hit github's size limit even though all your files are quite small).
You could just as do stuff like "int x=1; assert(x == 1);" and so on. And it would be just as futile.
Many systems are relying on the statistical improbability of 32-hexdigits uuid collision (name-dropping one - Disque).
For instance, if you run this database on Linux, the RAM overhead for each document is 1K just for the file handle, see the comment at the bottom here http://lxr.free-electrons.com/source/fs/file_table.c
New version is 1.1.0 (https://github.com/nodesocket/jsonlite/releases/tag/1.1.0) if you care.
* Serverless typically implies something accessible from more than the host you're on
* It's not zero configuration, there's at least 1 configurable parameter (albeit with a sensible default)
* I'm actually not even clear why this is a document store limited to JSON, other than you're piping it though a JSON python module.
Having done some things like this in the past, as you continue, you'll probably want to create subdirectories based on the first N characters of the UUID, but I'm not sure you wouldn't get considerably more value of just putting your json blobs into an AWS Dynamo table.
What is the correct terminology to highlight this architectural distinction?
While I am sure there are use cases for this, it is not correct to say this is "serverless".
By that standard you can call literally ANYTHING "serverless".
For starters, this coffee I'm drinking right now is "serverless". It does its job well, completely without a server!
The idea is breakup uuid keys into sub directories based on the first couple of letters to prevent file system performance issues. Seems like the right usage of the word from a dictionary perspective.
Example: ./jasondir/aa/bb/cc/aabbccdd
You won't like what happens when you put 100k files in one directory.
You can also run into issues like "Argument list too long". ARG_MAX is larger on linux than it used to be, but it's pretty short on older kernels. I assume similar issues might exist on other operating systems.
mkdir uuid; cd uuid
uuid -v 4 -n 1000000 |\
while read uuid; do
touch $uuid;
done
And ran out of inodes, but: time ls uuid|wc -l
425621
real 0m1.796s
user 0m1.552s
sys 0m0.240s
Sure, it's not exactly stellar performance for a linear scan of ~400k keys - but it's not terrible (for various values and expectations of terrible).This is in a hyper-v vm on a Surface 4 pro/i5.
The fact that the "uuid" program can quickly generate uuids make me wonder if maybe one approach would be to generate uuids to (a) fifo(s), and then let db thread/processes read uuids from the other end?
Even at 400k files, you see longer wait times if you've aliased ls to ls --color (pretty common), or use something like ls -F. Either runs stat() on every file.
Then, somewhere in the 1m+ range, it gets unusable.
As this is a standard Ubuntu install, ls is indeed aliased to "ls --color" - but afaik ls as standard detects pipes, and turns off color (so you don't get a lot of control characters if you do "ls --color | sort > file.txt". Unless you use --color=always if I recall correctly.
Just wanted to differentiate between "select * from documents" and "select count(*) from documents" being slow.
What are the use cases of this project?
What is the added value of that project? does it come up with something new or interesting? a query language for data? an efficient data storage? no, it just store json files in directories. Sure the file system is already a database, but that project doesn't add anything to the file system.
By the way sqlite just got support for JSON :
https://www.sqlite.org/json1.html
off-topic : I'd like to see more books on database implementation for beginners. Of all the crap load of CS books that come out year after year this is a matter with very little literature, as most papers on the subject are research papers. Seems like an excellent topic to teach distributed programming.
- SQLite: https://sqlite.org/serverless.html
- UnQlite: https://www.unqlite.org/features.html#serverless
Serverless is used to highlight this specific property of them. The OP's database is NOT embedded, but it is serverless.
SQLite stores all its data (every table and all entries in tables) in one single file. JSONlite does store every JSON entry in one particular file, as much I saw. So, JSONlite uses the file system (one directory) as data storage -- in contrast to SQLite!
Huh, what does that mean?
id=$(uuidgen)
echo '{"hi":"mom"}' | jq . > id
cat id
I really don't get why this made it to HN first page, especially when plenty of other show HN demonstrate way more efforts than writing a short bashscript yet never make it. I suspect it has to do with the number of buzzwords that were inserted into the title.
I would be really happy to learn what are the use-cases for this tool, since using it via bash is actually an overkill.