fatcache - Memcache on SSD
github.com
github.com
I was curious why they would bother, but it seems this isn't quite accurate.
What happens is they first use 32 bits from the SHA-1 hash to find the hash bucket, then they scan for the full SHA-1 of the key. They do not check for actual SHA-1 collisions.
edit: Also on the subject of hashes, the readme suggests switching to MD5 as a possible way to reduce entry size. That is unnecessary; SHA-1 can be truncated to whatever size you're comfortable with.
Michael Cornwell: Anatomy of a Solid-State Drive
Communications of the ACM, Vol. 55 No. 12, Pages 59-63
http://cacm.acm.org/magazines/2012/12/157869-anatomy-of-a-so...
Alternative link in case the first is paywalled for people not on a university campus: http://queue.acm.org/detail.cfm?id=2385276
It goes into quite some detail on how SSD storage works on a system level and how it differs from hard disks.
Since this is a cache I really dig skipping any kind of cleanup/compaction step for deleted/expired keys.
I played around with a similar thing except as a K/V store and the performance and density is pretty amazing. With a 64 byte key and 1.5k value (compressed from 2k) I was getting 85k inserts/sec and several hundred thousands reads/sec with a quad-core Sandy Bridge i5 and a 128 gigabyte Crucial M4 on SATAII.
In my experience this is not true. A 1 TB SSD, even commodity hardware, is pretty expensive. And that's not even talking about high performance SSD's like Fusion IO, where you're now talking about 10K+. Where do you see 64 GB of RAM costing that much?
https://www.computerworld.com/s/article/9235277/Micron_unvei...
Fusion IO is probably as expensive as RAM, so only makes sense for apps that really need persistence, ie databases.
Is that check really necessary?
To have 1 in a trillion chance of having accidental SHA-1 collision they'd have to store 1.7*10^18 keys, and mere key index of that would require 54000 petabytes of RAM.
By then, maybe the founders of this system won't even be alive.
Accidental sha-1 collision is probably not a problem, but in a few years [1] it will be possible to crate sha-1 collisions and use that as an attack. It looks difficult, but supposes that with the correct string an attacker can retrieve the cached information of another user, for example sha1("joedoe:creditcard")=sha1("atacker:hc!?!=u?ee&f%g#jo").
I don't know if they are using randomization, because the collision can be used (in a few years) as a DOS atack [2]
[1] http://www.schneier.com/blog/archives/2012/10/when_will_we_s...
For instance MD5 collisions are really easy to create but for preimage attacks on MD5 there is still no better approach than just doing brute force.
There have been attempts to use an SSD as a swap layer to implement SSD-backed
memory. This method degrades write performance and SSD lifetime with many small,
random writes. Similar issues occur when an SSD is simply mmaped.
To minimize the number of small, random writes, fatcache treats the SSD as a
log-structured object store. All writes are aggregated in memory and written to
the end of the circular log in batches - usually multiples of 1 MB. ptr = mmap(..., len, ..)
/* do stuff with ptr */
msync(ptr, len)
No sane OS will pay attention to an mmapped region while it isn't under memory pressure, so dirty pages are effectively buffered until you explicitly tell the OS to start writeback. msync(ptr1, len1)
msync(ptr2, len2)
msync(ptr3, len3)
Where [ptr1, ptr1+len1], [ptr2, ptr2+len2], ... are the chunks within a big mmap'ed region where changes occur and need to be written to disk for persistence.Or do I just msync the whole region then hope and pray that the OS will do the right thing?
The whole point of this piece of software is to be smart about how and when it flushes data so it can minimize impact on the write counter.
http://www.monomachines.com/shop/intimus-crypto-1000-hard-dr... or you can get a service to come out and do it on site.
You get a write amp of 1 until the drive is filled the first time. After that, it's a function of 1) how full the drive is (from the drive's point of view—this is why TRIM was invented) 2) the over provisioning factor 3) usage patterns, such as how much static data there is 4) how good the SSD's algorithms are 5) other (should be) minor factors, such as wear leveling
Source: I used to be a SSD architect.
https://github.com/twitter/fatcache/blob/master/src/fc_util....
However this one isn't kernel-based, so it won't help your NFS server or your postgresql engine. On the other hand it's much easier to build.
With that, how much faster is SSD over disk and then memory over SSD?
DRAM (~100 ns) is very nearly 1000x better latency than SSD (~70 us) which is only 100x faster than disk (~10 ms).
26 microseconds for OCZ Vertex 4 (reads)
http://thessdreview.com/our-reviews/ocz-vertex-4-128gb-ssd-r...
No more than 20 microseconds for Corsair Neutron GTX and Vertex 4 (writes):
You should be looking at companies like STEC or higher end Intel SSDs for server applications.
It's too bad most cloud providers consider SSD disks to be a 'premium' feature. I guess this would work fine on custom-configured hardware at places like softlayer and serverbeach.
Awesome, useful, and cool, but insane.
I like it.