Diskhash – Disk-based, persistent hash tables
metarabbit.wordpress.com
metarabbit.wordpress.com
Secondly, the other nice bit is that you can use structures larger than the physical RAM on the system. It's kind of like getting a system with enormous amounts of slow-ish RAM.
Thirdly, you can build a whole bunch of these for different purposes and just open the ones you need, and not bother with the rest, again like having even more RAM.
Finally, you can multi-process trivially on these as the OS sorts out the mess and suddenly it's like having a slightly slow multi-million dollar multi-tb shared memory supercomputer from 5-8 years ago at your disposal. Works great on Beowulf clusters if your disk storage is shared on fast links and can handle the IOPS.
The number of applications you can do with a 5-8 year old super computer are vast. I've seen 20tB classifiers built on more or less commodity hardware (<$100k for a cluster of cheap hardware and one big disk) that were previously thought to require a $5 million dollar supercomputer just because of the large shared memory. You can probably build an equivalent machine for <$10k these days off of Newegg and some decent NAS boxes.
Right, but one of the not-so-nice-things is that you can't do the I/O asynchronously, so you can end up with poor performance depending on the access pattern. (I guess you can, with another thread running in the background touching pages, but it's more of a pain and you're not really guaranteed the data will stay in memory until you use them.) [Edit: Actually I guess if you touch by writing to those pages rather than just reading from them then they'll have to stay in memory... though do note that I'm assuming no swap here.]
Beautiful. There are so many times in my career that I've had to use poorly hobbled together wrappers for other languages that I now really appreciate it when someone takes their time to make their API "Fit" with the language.
It was introduced in Perl 5.0, so it's been around for about a quarter century, and it's generic in the sense that there are tons of backends that use the tie interface. BerkeleyDB and gdbm and flat files and probably any DBD SQL database are among the backing store options. But, again, a lot of people think it should be avoided in the general case, even though it seems like a really nice simple solution to a common problem (I've got these hashes/arrays/whatever, and I want to save them on disk so they're persistent and shareable).
This implementation can't delete keys (yet?) which seems to limit its utility for general purpose tasks. I also wonder about reliability. When I first started reading I assumed it'd be used for small things, but he's working with a 32GB file shared across multiple processes. That's a lot of opportunity for data corruption.
> Perl tie, which is kinda considered an anti-pattern in modern Perl ... a lot of people think it should be avoided in the general case, even though it seems like a really nice simple solution to a common problem
Why? Any links to blog posts?
> When I first started reading I assumed it'd be used for small things, but he's working with a 32GB file shared across multiple processes. That's a lot of opportunity for data corruption.
I've used something similar and data corruption because of size is low on my list of concerns. If your storage layer fails in this manner silently, you have much bigger problems.
Here's one blog post about it (still opinion):
http://www.skrenta.com/2007/05/tie_considered_harmful.html
And, this article from the Modern Perl site itself has mixed opinions on the subject, but comes down on the side of not using it, in the general case, but knowing how and when to use it can be really valuable (Tie is the last subject covered on the page):
http://modernperlbooks.com/books/modern_perl_2014/11-what-to...
I'm not trying to say there's no place for it. Tie in Perl has been used for really cool stuff.
cdb is limited to 2GB, yes.
https://github.com/attic-labs/noms/tree/master/go/nbs
... but the idea is developed to support fast mutations as well as other related datastructures (sets, lists, etc).
From the second paragraph:
My usage is mostly to build the hashtable once and then reuse it many times. Several design choices reflect this bias and so does performance. Building the hash table can take a while. A big (roughly 1 billion entries) table took almost 1 hour to build. This compares to about 10 minutes for building a Python hashtable of the same size.
If you have a tiny application, doing a tiny thing, do you need redis? Do you need sqlite? do you need to write some custom-text-file-ma-bob? nah. I feel like simple simple little tools like this are in short supply.
If it's more dynamic than that I mean I guess it's cool to have a hash in a backing store that's dead simple.
But at sizes big enough and with enough dynamism that gperf or something similar isn't appropriate I think you should be considering dbm or SQLite, sort of by definition.
Use redis if you can. Use something like diskhash if you cannot afford the extra overhead of sqlite/redis.
In my case, the alternative we were using before is an in-memory hash table built at startup from a text-based representation of the data. It was pretty good and worked very well up to millions of entries, but at the current scale we work with, it takes least 10 minutes at startup were necessary to build up the thing and it used ~200GB.
We are a publicly funded research organization, so even our "work code" is open source.
I was so confused.
Unfortunately, on my initial tests, it just was not fast enough.