When a disk cache performs better than an in-memory cache
productiverage.com
productiverage.com
I used to think this way, then I tried Rust. After C and C++, I never thought I'd want to give up the GC, but to avoid the issues that this article talks about, I wanted to go back to a systems language.
Now I want to use Rust for everything.
If you're not using new/delete in C++, you're probably not doing something very complicated. (And hey, that's not necessarily a bad thing!)
There are really two classes of C++ features, the basic use part which is fairly straightforward and nice to use. This is the API provided by the STL. Followed by the infrastructure stuff like templates, move semantics, SFINAE and other messy and non-obvious things. This latter part is much more complicated and still a mine field - but necessary for the STL to do what it does.
If you want to learn C++ get proficient in using the STL. The rest should only be learned once that is second nature.
Slight caveat.
> Step 3: Pooling large byte arrays used in serialisation/deserialisation
If you are using a explicit memory language, buffer pooling is still quite important - especially if you are doing any IO. You can avoid this to a certain degree by using fast malloc implementations.
https://internals.rust-lang.org/t/on-native-allocations-poin...
https://github.com/rust-lang/rfcs/blob/bdbe73d1be948dd925c6b...
In this situation on .NET, I wrote a resource manager that handed out handles to large byte arrays (in particular, ones that were too big to be collected outside of gen2, in the large object heap, which kicked in at 80kb at the time IIRC). The handles implemented IDisposable so taking care of handing back the byte array was no more or less tedious than any other resource you need to manage explicitly. The resource manager kept a hold of the arrays internally using a weak pointer so they could still be collected when gen2 collections actually happened, but allocating the buffers themselves would never cause gen2 collections in a steady state.
To turn that into a cache, you'd need another layer with keys, an eviction policy and an invalidation mechanism. I think it ought still be better than round-tripping to disk.
I wrote a different version of the resource manager that used P/Invoke helpers and unsafe code to allocate from unmanaged memory directly, but it didn't perform any better - it didn't relieve any pressure on the GC, which was 2% of CPU usage at full load in any case.
And that's often not the worst idea, because, when done correctly, this stuff never hits the disk when enough RAM is around.
Plus, life cycle management is done by the OS, not by you. It also tends to work better than "not at all" on memory pressure or if there isn't a lot of memory in the first place.
If you write a lot of them, then even if you have the memory, you may overrun the size limit of the write buffer and cause application stalls.
Writing to a ramdisk (e.g. tmpfs) is always an option, though.
I agree in a memory-rich environment it's far from the worst idea. However now you need to manage the files. You've just pushed the problem somewhere else.