> What happens when a mailbox exceeds it's limit? Does the data get dropped?
I thought there was some movement towards limits on mailboxes, but I can't find any documentation now, so I'm not sure if that happened? If not (or if you haven't configured it anyway), there is no explicit limit, your mailboxes can grow until you run out of memory; either by hitting a ulimit, or malloc fails, or maybe until your OS just kills processes (and probably the BEAM process, because it's biggest). In the first two cases, you'll get a nice crash dump from BEAM, but in all cases all messages are dropped, as BEAM is dead. Edit: i see there's a process_flag(max_heap_size, MaxHeapSize) to set the maximum size of the heap, and if process_flag(message_queue_data, on_heap), the default, is set, messages will eventually end up on the heap. But the maximum heap size is checked during Garbage Collection, but IIRC, GC can't be triggered when a message is added to the mailbox, only while the process is running, or if explicitly requested for the process (with erlang:garbage_collect/0 or /1); if your process ends up blocked for a long time (or possibly forever), it could still accumulate a large mailbox without being killed by the heap size limit.
You can (and should!) regularly call process_info(Pid, message_queue_len) to observe the message queue of all processes, and alert on large queues. You can then observe the messages themselves and consider appropriate response.
> Or, how to recover from a network segmentation? These proved somewhat challenging to reproduce and troubleshoot (as distributed problems can be).
Recovering from network segmentation is application dependent, and can often be tricky. Some applications can just reconnect and call it a day. Other applications may have accepted writes on both sides of the segmentation, and need some sort of reconciliation process. Mnesia has hooks for this, but I don't remember seeing any examples, and the default logic is to just continue segmented even after the segmentation is done; this is usually not what you want, but at least it's consistent? I think it should be fairly easy to simulate and trigger network segmentation, just kill drop packets between selected hosts; although you'll need more work if you want to simulate stuff like congestion between hosts or congestion on only some paths between hosts (LACP is very nice, but debugging congestion on only some paths isn't as nice).
On this particular issue, where I worked, we had a policy of flushing mailboxes that were too big (usually 1 million messages, which isn't the Erlang way, and wasn't in public OTP, but keeps a node running at least), and we wouldn't have tried to log all of the messages in a mailbox, because 1 million messages or whatever is way too many to log. Pretty printing with no limits is dangerous, even if it doesn't include a ton of references to the same big thing. We also didn't tend to use anonymous functions/closures, but that's just a happy accident: we were using Erlang before crash dumps had line numbers, and anonymous functions are hard to track down, so it's easier to give them a real name and use that instead. Of course, there's some places where closures are way more convenient than explicitly passing Terms to Funs, so it's not that we never used them, just they were rare, and unlikely to show up many times in a single logging statement, like in this case.