Writing a Memory Allocator for Fast Serialization
idryman.org
idryman.org
Jiri Soukup's 2001 book on "Serialization and Persistent Objects: Turning Data Structures into Efficient Databases" consists of many techniques, including the serialisation of mmap'd pages.
https://books.google.co.in/books?id=DHDABAAAQBAJ&pg=PA74&dq=...
Tuning the allocator is not as straight forward as you may believe. If you have variable sized allocations the problem is fairly difficult... you essentially are forced to rewrite a worse version of ptmalloc, jemalloc, or tcmalloc. If your allocations are fixed, you're in a slightly rosier situation. However, you have to consider - how will you support deletions? Will you journal and garbage collect? Are you going to force variable latency? Are you going to implement atomic barriers on lockless structures? Now that I think of it... what is the cost of an atomic operation on a shared memory map? You will also need to concern yourself with cache hits/misses. In my experience it is somewhat difficult to predict what memory in your map is going to be in cache and what won't. If your data scatters... your performance is going to be fairly slow.
In some cases, where data is built up once, then reused read-only, you don't need to pay the reallocation cost. For example, the Borland C++ compiler used this approach for precompiled headers. The symbol table information was allocated in a single contiguous blob of memory, with pointer locations noted just like you'd note fixups when writing an object file. Then, when the precompiled header was loaded, the fixups would be iterated over and the difference in old load address and new load address would be added to every pointer location: and all the pointers work again!
The idea is in this conjunction of linkers, loaders and garbage collectors; all three are strongly related functionalities. A smart linker is a copying GC; a moving GC is almost isomorphic to an OS loader, except the source is memory rather than disk (or mmap); a loader is a runtime linker; etc.
Reference. https://github.com/cksystemsgroup/scalloc https://github.com/kuszmaul/SuperMalloc
* libsrt i64-i64 map (equivalent to std::map <int64_t, int64_t>: > 10M QPS
* libsrt string-string map (equivalent to std::map <std::string, std::string>: > 1M QPS (> 2M QPS if key size <= 19 bytes)
[1] Repository: https://github.com/faragon/libsrt
[2] Benchmarks: https://github.com/faragon/libsrt/blob/master/doc/benchmarks...
The biggest problem with interprocess is that out of the box it is not capable of handling application failures gracefully and transparently.
[1] Basically the container must not assume that the allocator::pointer type is an actual raw pointer as interprocess uses a custom offest pointer.
As there is exponential increase in the in the performance critical software which runs on the dedicated machines, what if we avoid the abstraction of Operating System for them and run using bare minimum, optimized system software?
What if we can pick the specific OS kernel modules + drivers required for our application/ machine needs, tune it for the app and deploy the whole stack (kernel modules + app) as one software?
Example: For a database to work, we need Networking module (accept and send request on a given port), Memory module (to access disk, memory, cache, etc), and Processor handling modules (to create a process/ thread, if we can call that)
Let's say a Dockerfile kind of thing, which specifies all the required modules/ driver for the given software. The modules can be compiled for the architecture and deployed.
Advantages I can see are:
1. Lesser abstraction 2. Full control on scheduling the software. Hence, lesser synchronization issues. 3. Kernel modules optimized for the application software.
All the above three things leads to much better performance.
I see the following problems with the approach:
1. Existing softwares(both user and system) doesn't suite well. 2. Increased development time (as system software needs to be tuned as well). 3. Not so many system software developers.
Isn't that's how the software architecture for the high performance software (which runs the dedicated machines) should be in the first place? Given that we are running so many of them now.
A complete OS (with all the modules and generically coded) looks fine for just the end users who uses variety of not-so-perf critical apps.
What do you think about it?
One interesting thing I've noticed about these kernels is their tendency for full-stack language integration a la Lisp machines and Lisp OSes like Genera. For example, Mirage integrates strongly with programs written in OCaML (the home page describes the project as a "library operating system") and HaLVM with Haskell. A quick search shows unikernels for Golang (Clive) and JS (Runtime.js) as well.
Shouldn't software running on dedicated machines need lesser management (scheduling), abstraction (virtual memory) and finer control (optimized for the architecture) over the hardware?
If you just don’t wanna use C++, fine. But this looks like a great fit for it to me. Avoid the STL, turn off the features you don’t need, etc.
The idea was that the individual elements were not always fully accessed when the application ran, so if I created them on demand from such a dense memory-mappable dump I could persist that instead of parsing every time.
The overhead of creating Python object was too high for that. But if you are using Python and are deserializing read-only dictionaries, http://discodb.readthedocs.io/en/latest/ does a subset of that -- if your app e.g. reads in 100,000 translations from a JSON file, Disco's serialization will let you just mmap them.
This problem is generally hard. See [Ensuring data reaches disk](https://lwn.net/Articles/457667/)