For us to measure our success in increasing availability, we first had to figure out a way to measure availability.
We came up with a multi-pronged approach. The first thing we did is figure out how to predict how much traffic there should be at any given time. This was basically using historical data to determine the shape of the curve and then adjusting it to fit current traffic.
Then we would figure out how far off of the predicted traffic we were and the was our downtime.
Since we also had control of the client experience in almost all cases, we were able to measure from the client side as well (and trust that data), and include that data in our measurements.
Where things got interesting was when say a large ISP was down. A bunch of people couldn't get to us, but that wasn't really our fault (or was it?). At the end of the day the users didn't care if it was our fault or their ISP, so we counted that against ourselves.
All of this is to say that yes, it's really hard to figure out uptime, especially for distributed systems where almost every failure is a partial failure.
But at the end of the day your users don't care which microservice or network segment was at fault, they only care that they couldn't use the product.