Storing C++ Objects in Distributed Memory
people.eecs.berkeley.edu
people.eecs.berkeley.edu
But what C++ is doing (and static languages with metaprogramming in general) is specializing and statically inlining everything at compile time so the assembly just takes the fastest path every time; no runtime cost.
I'm unsure which static languages do and do not support partial specialization. My understanding is that the current iteration of Rust doesn't have partial specialization, but maybe it has other features that could do what this post is doing?
I'd be interested in hearing people's perspectives on the metaprogramming power of different statically typed languages - people have expressed interest in me building libraries like this in Rust and a few other languages.
Buuuut I'm probably missing something.
Besides, there are techniques to avoid dynamic dispatch for generic specializations, at least in C#.
[0] this thread
[1] https://news.ycombinator.com/item?id=18650902
We had to build everything ourselves except for the STL like features of RW Tools.h++ and a few other of their libs.
1990s C++ compilers were ... interesting ... - especially for templates.
Have you used SOM as well?
We tried a few things with it. It was a good idea that took too much manual work. Much like CraptiveX/COM but without the toolworks.
I lost everything I had archived in a series of hardware failures and stupidities...
I've repeatedly asked the co-founders to let me have some of the stuff and they still refuse. They didn't think I should have had it to start with.
It really was simple in concept though. Basically a queue server on each machine, memory mapped files and containers/string class/bignum class etc using placement new. The hard part was finding machines that weren't steaming piles of crap. Compaq was the only company we found that had server class machines we could rely on. We also had to stick to IBMs C++ compiler due to all the bugs in Borland and Watcom at the time.
These sorts of things form a spectrum from low level message passing all the way to a full fledged database. Some things in between are abstractions like MapReduce which still gives a lot of control, but also gives you fault tolerance and lower configuration of your cluster to run different applications.
Another abstraction in this space that is just a bit higher than ordinary message passing are systems like Kafka and RabbitMQ which give additional guarantees over plain message passing using sockets as well as less configuration to set up/remove machines.
One thing a lot of the more successful abstractions in this space seem to have is a pretty clear line between which parts of the system are local and which are distributed. It seems like this system doesn’t have as clear of a demarcation which would make understanding performance much more difficult. Databases can also have this problem since they do a lot of low level performance tuning automatically, but at least they provide a very high level of abstraction for applications to work with. This seems like it could end up being confusing to tune without providing a nice high-level interface like databases.
Remote pointers end up being pretty nice for locality, since you can explicitly see what process a remote pointer is pointing to.
None of the backends that my remote pointer implementations depend on really offer fault tolerance, though. So you're right: if you're in a situation where you have nodes entering or leaving your cluster before your program ends, this model needs extending before it works for you. That's true for MPI-style programs, generally, though.
You can then avoid doing any copying or serialization by directly sending the entire memory region as a unit.
The main difficulties with this approach is in deciding how large to make these regions and also handling free over this region. The first part is often not too bad depending on the application, but handling deallocation and expanding/shrinking memory regions can be tricky. This allocation problem is much more difficult to dismiss since it is necessary to solve in order to avoid making the user decide how much memory to allocate initially which is impossible for many applications. For instance, when dealing with strings, it may be impossible to get a good bound on the size of these strings.
One problem is that languages don't support the use of standard container types inside a region with a variable base address (variable because it will change the next time you call mmap).
I think modern languages should support this concept, especially the ones that aim to be "systems" programming languages.
Funny you should mention that - something has kept me thinking that the modern operating systems should have a built-in support for the cross-node sharing of the virtual memory via mmap.
[1] https://256fd102-a-62cb3a1a-s-sites.googlegroups.com/site/mk...
Making remote accesses explicit rather than implicit makes it a little more obvious to code readers how grossly expensive they will be.
Today's multi-core systems are really distributed systems, and the cpu cache does not hide that well enough that OS software doesn't have to work around it.
For example, kernels have to send IPIs (interprocessor interrupts) to other CPU cores to get them to perform some actions; cores sharing a specific cache line, or memory address that happens to alias to the same line in cache can create a 'ping-pong' effect that can severely degrade performance; L3 cache latency can be non-uniform depending on how far away the cache is from that particular core. The black box is breaking down as core counts increase.
template <>
struct serialize<int> {
int serialize(const int& value) {
return value;
}
int deserialize(const int& value) {
return value;
}
};
error: return type specification for constructor invalid