All of these embedded NoSQL databases seem to be missing critical features. One such feature for my use case is database compaction. Last I checked, an LMDB database file can never shrink. Full compaction of LevelDB is slow and complicated (as I understand it essentially breaks the levels optimization which is the whole point of the thing.) SQLite meanwhile supports fast incremental vacuum, and it can be triggered manually or automatically.
SQLite just has everything. Plus the reliability is unmatched. Even if you just need a single table that maps blob keys to blob values, I would still recommend SQLite over any NoSQL database today.
Linux has VFS cache which is very robust and efficient.
Remember that it is typical for programs to read /etc/nsswitch.conf, /etc/resolve.conf and tons of others at startup time - the filesystem is the datasource in Unix tradition, so the machinery is very well optimized.
Not to mention all the other problems with this. The filesystem has a complete lack of higher-level features: no transactions, no snapshots, no indexing beyond filenames, no easy robustness guarantees (doing fsync() properly is a lot more complicated than it appears.) Honestly for modern apps the filesystem is just terrible at storing any internal mutable app data.
Once you start writing code to store auxiliary indices, synchronize writes, or pack multiple records per file, well at that point you're just implementing your own database. This might make sense if, say, you have a special way of compressing your data (like git). But generally you're better off using a real embedded database.
https://unix.stackexchange.com/questions/197570/is-it-possib...
https://ext4.wiki.kernel.org/index.php/Ext4_Disk_Layout#Inli...
https://www.gnu.org/software/libc/manual/html_node/Directory...
# zfs create -V 100G -b 4096 tank/test && mkfs.ext4 -v /dev/zvol/tank/test && mount /dev/zvol/tank/test /mnt && cd /mnt
# df -k .
Filesystem 1K-blocks Used Available Use% Mounted on
/dev/zd16 102626232 24 97366944 1% /mnt
# for x in `seq 1000000`; do echo $x >$x; done # create 1M tiny files
# df -k .
Filesystem 1K-blocks Used Available Use% Mounted on
/dev/zd16 102626232 4022348 93344620 5% /mnt
# bc -l
(97366944-93344620)*1024/1000000 # free space diff per file
4118.859776
Plus you're wasting an inode per record - a limited resource in ext4, increasing which requires reformatting. You'd probably run out of inodes much sooner than out of space.Re inodes, this is a good point too. These definitely reduce the size of db that fs works nicely for.
If you just need a r/w store for some jsons in a single file, why not sqlite? You can put arbitrary-length blobs into it. Some sql will be involved but you can hide it in a wrapper class tailored to your application with a few dozen lines of code or so.
And it's quite suitable for this purpose: https://www.sqlite.org/fasterthanfs.html
There's no strict definition of nosql so everyone can choose their own. My personal take (in broad terms) follows:
No, that's not a feature of nosql. Nosql means not relational, which in turn means no guarantees about the relationship between two objects, i.e. no atomicity of access across multiple objects (for either read or write operations). A consequence of this lack of atomicity is that it's easy to store different objects in different places, thus opening up opportunities for horizontal scalability. Caveat: those opportunities can be taken away by other choices you make. If you decide to offer and enforce transactions, you are bringing atomicity back into the system, and thus making horizontal scalability hard again. Or you may decide you want a nosql-in-a-file.
Here is a NoSql database based on the SQLite backend: https://github.com/rochus-keller/Udb.
I use it in many of my apps, e.g. https://github.com/rochus-keller/CrossLine. It's lean and fast, and supports objects, indices, hierarchical "globals" like ANSI-M and transactions.
An interesting one I ran into recently is Datalevin, a Datalog DB on top of LMDB for Clojure: https://github.com/juji-io/datalevin
It's really awesome, and the team behind it are super responsive and helpful.
It's small yet capable. If you are familiar with MongoDB, you will feel right at home.
It's great for .NET developers as it's written in C# but since it's Netstandard 1.3 compatible, you can presumably run it under Ubuntu or Mac OS or wherever else the new .NET 5 runtime works. I've got a C# app running on ARM64 the other day - just saying.
I wrote about my experience playing with LiteDB here - https://tomaskohl.com/code/2020-04-07/trying-out-litedb/. It's not an in-depth look at all, just a few notes from the field, so to speak.
[0]: https://en.m.wikipedia.org/wiki/Lightning_Memory-Mapped_Data... [1]: https://dbmx.net/tkrzw/
GDBM and BerkeleyDB are the "grey beard" references, but their 80's & 90's heritage shows.
With proper concurrency control, it can work very well even for multi process applications.
However, it is still in heavy development and a bit of a moving target even if the developers are currently heading toward stabilization of the file format.
Tango Kilo?
In memory, high performance, no schema. Get the object to journal to disk and you are almost there!