If you're really into telephony history, the Internet Archive has "The History Of Engineering and Science in the Bell System" (3 volumes) online.
If you have to build reliable distributed systems, it's worth understanding how this was done in the electromechanical era of telephony, where the component reliability was much worse than the system reliability. "Number 5 Crossbar"[1] is worth reading, but hard to follow if you have no idea how telephone switching worked and are unfamiliar with the terminology.
Number 5 Crossbar, in current terms, was a collection of microservices. There was a big, dumb switch fabric, and "markers" which told it what to connect. Other microservices included trunks, originating registers (which listen to incoming dial digits), senders (which sent dial digits to the next switch), billing punches (which recorded toll call data for later billing), translators (which held routing tables), and trouble recorders (which logged errors.) Central offices had at least two of each resource, for redundancy. Resources were "seized" as needed from resource pools, with a hardware timeout and alarms to prevent resource lockup. If something went wrong in setting up a call, it was retried once, using different resources. If it failed on the second try, the caller got a fast busy and there was an alarm and a trouble recorder dropped a trouble card. Markers did not have persistent state. They started each call with a reset. So they could not get stuck in a bad state.
In the entire history of the Bell System, no electromechanical switching office was ever down for more than 30 minutes for any reason other than a natural disaster or a fire. It's worth understanding how they did that.
[1] https://telephoneworld.org/mdocs-posts/number-5-crossbar-sys...