RAM Is the New Disk
medium.com
medium.com
If you use Linux, the fastest way to test how much faster your application is off disk is to simply make a filesystem in RAM, and run the whole thing from there. Because library-chasing to build a chroot is a hassle, I would recommend simply putting a container on a RAM-backed block device, then installing your application on the container.
I have personally designed, built and managed large clusters of diskless machines and find that the mix of RAM-only and PXE[1] boot is an excellent one for maintaining state (and security) across well managed infrastructure. Disks be damned. For permanent storage, consider sharing a DRBD[2] cluster from dedicated nodes.
[1] https://en.wikipedia.org/wiki/Preboot_Execution_Environment
Here's my anecdote based on 16GB workstation with NVMe SSD (Samsung 960 Pro):
Watching my project compile I occasionally open iotop in another terminal and don't see anything above occasional write flushes. To confirm, I did create a tmpfs volume and did not observe any improvement. `free` reported my buffers to be at ~4.7GB, which is basically all of my /bin, /usr and all of Golang sources+libs.
[edit] Not sure if ramdisks are pinned though.
Ramdisks will go to swap. A memory leak will force the entire ramdisk into swap, and reading it back into memory afterward is 10 to 100 times slower than reading normal files off of a disk.
Assuming that you have swap. I don't; I want my SSD to stay alive.
Hopefully it will target the program with the memory leak, but this is not guaranteed.
Swap is useful because you can shift unused memory onto disk. There are many programs that allocate (and write) a lot of memory that they never afterward use.
By having swap you make more room for cache in memory.
SSD doesn't matter here - this is not swap thrashing, but rather occasional writes.
DRBD is fine (used it for years) but it's not something that is one size fits all.
Rare for most services. Stuff like logs can be shuffled off elsewhere for a write, requiring no commit validation. Only DB/fileservers really require permanent storage with commit validation, writes are typically rare, and 100Gbps+ LAN on a PXE-based diskless cluster is not going to be introducing massive latency, especially if you prioritize the VLAN or link multiple ports. Reads are typically cheap and cacheable.
that can't be cached|copied to the your ram filesystem for the lifetime of your container
IMHO most services and their dependencies will come in well under 512MB, so that's a non-issue.
you can't scale with RAM backed storage without more phys memory
By definition, one could say the same about anything... although to be fair you could still scale via compression, sharding, or another established strategy.
when you run out of RAM...
In a managed scenario a service container or VM would terminate or a significant degradation in response time would be detected, it would be taken out of the service pool and stop having traffic routed to it, be restarted, then be re-introduced to the pool. Ditto extra CPU load, broken network policies, anomalous block IO, etc. Leaving modern service-level architecture aside, basic heartbeat-style IP monitoring with reliable node-level failover has existed in open source since the 90s. There's really no excuse to wing this stuff on production systems today.
it's not something that is one size fits all
Nothing fits all!
I was the third engineer at VoltDB and spent six years making that bet. It's not a good bet.
Maybe there are other factors, but if VoltDB could page out cold data to disk I think it would be at least 2x if not more successful. No one agreed with me so it never happened.
I saw so many use cases go out the door because hey you know what? RAM is expensive and it's cheaper to page out cold data. The scale where that cost starts to matter is not that big.
We did work on equally important things also, but we also split focus with IMO unimportant things.
A combination of me not having a seat at the table (literally was told this after a year or so) and IMO non-technical leadership driving focus by chasing what they thought were the important factors.
The company survives and does OK though.
In memory is fast and awesome, but it doesn't have to be as mind boggling expensive as it is. Why are we all making the same mistakes?
Not that I think that's ideal either though, having both in memory for hot used data, and the rest on disk is ideal. With an extremely easy to use setup that makes it essentially automatic, but with rules engines for finer grained tuning.
SQL Server similarly has the hekaton in-memory tables + columnstore indexes and the latest version allows combining both for in-memory columnstores.
The results of the columnstore data was pretty fast, and it's even faster in memory. Depends on what you're doing, and what the requirements are.
Was really impressed by MemSQL, and loved the wire compatibility with mysql, so don't take this as just a knock on MemSQL in anyway.
What do you mean?
But then I went to their docs to link you to the details, and it feels like they intentionally avoid stating clearly the problem.
Essentially, they allow committed transactions to hit memory and not disk. They allow you to configure it so that's not the case, but it isn't the default, and looking over the current documentation they certainly aren't clear about it like an open source project would be.
transaction-buffer needs to be set to 0 for durability, but the way the docs are explaining it is trying to confuse not being durable, as a different kind of durability.
I'm not interested in getting into a long discussion about this though, but it's difficult to explain the literal issue when they do such a marketing job of trying to hide the specifics.
Now I'm far less surprised a user wouldn't know this. Apologies for my forward initial statement.
The enterprise edition has HA so with data on 2 nodes for safety. Otherwise you lose the whole point of in-memory performance (for writes) if you're going to write every single bit to disk immediately.
They explain it clearly on the durability page: http://docs.memsql.com/docs/using-durability-and-recovery
So, given this insider knowledge of yours, can we make a prediction by what date predicively DRAM/SSD-NVMe prices may make in-Memory Database Startups lucrative again?
--
Offtopic:
I feel empathy for you, being overrun in decisions as an engineer in a field you're the expert in market and technology by decision makers can be heart-breaking.
EDIT:
removed irrelevant personal experience
What is also not shown is that cold data is everywhere. You need to have it, but paying to put it in RAM generates zero value for a profit seeking business. So if you don't page out cold data you effectively throw yourself out of the running for a huge swath of use cases.
For a small deployment sure it's dwarfed by engineering costs. But infrastructure per engineering head count is trending towards more infrastructure per head and infrastructure cost matters to more businesses.
The other thing is that data volumes are also increasing at a rate competitive with RAM is decreasing in price. This is because there are new opportunities to make money using more data and this is a trend you can't really beat. The more data you can have the more use cases and lines of business get invented.
This is not a scientific analysis it's just conjecture based on anecdata from my time in the industry.
From ~$1000.00/gb to $0.03/gb
[1]: https://www.backblaze.com/blog/hard-drive-cost-per-gigabyte/
While I dont believe in Infinite growth of data, I still think a RAM only DB isn't as good if we have SSD that is ridiculously fast. My thinking is that RAM / SSD should always be 1:5 or 1:10.
For interpreted languages like Python or Javascript, figuring out RAM storage and access patterns of data is very hard. So we probably need programming language mechanisms to help with understanding the locality patterns of our programs and probably tooling to help change it.
It depends on the libraries you use. Take a look: https://github.com/Const-me/CollectionMicrobench
As you see, in practice, a good linked list is same or slightly faster than std::vector. And it’s consistently 2-3 times faster than equally linked std::list.
That’s not just synthetic tests. Recently, I’ve got 2.5x performance improvement in my app just by switching from std::unordered_map to CAtlMap with the same keys/values.
Theoretically, C++/11 fixes that with stateful allocators. Practically, I’ve not seen good open source ones with the performance comparable to CAtlPlex that powers these ATL node-based collections. I’m not even sure it’s possible. STL is too standardized and too old. It might be there’s no room in its allocators API for sufficient level of integration between a collection and it’s backing stateful allocator.
BTW, this is a blind spot/elephant in the room of most *VM-based and functional languages nobody wants to talk about...
The commenters who are saying disk has a different price/performance trade-off that is still valuable are also right, but that applies to large data sets.
I don't even know what a large data set is anymore. I think my general definition is one you won't put into memory, whatever your threshold is for that.
Yes, totally agree with the conclusion in your link above, consumer SSD (NVMe or not, high end or cheap) doesn't worth a dime. Cheers!
Summary performance comparison by Storage Class: http://xitore.com/what-is-nvm-x
NVM-X RAM ≈ DDR4-3200 ↧Availability|↥Cost| ↥4TB@25,6 GB/s
Diablo Memory1 ↥Availability|↥Cost|↱256GB@10GB/s
EDIT, added source:
[1] http://www.diablo-technologies.com/memory1/http://pics.crucial.com/wcsstore/CrucialSAS/images/campaigns...
I have a feeling with multi terabyte SSD's at cheaper prices we'll be shuffling all our data back to "disk" again :).
What Id really like is for them to scale up to full fledged VMs once some usage or performance threshold was hit.
Really? Have developer costs actually increased in real terms in the last 10 years? Have your developer costs (if you're outside the VC/SV bubble) increased in real terms? And how much?
This seems like a terrible assumption.