First, you prune tests to only those affected by the change. There's basically nothing (other than the build system itself) at the root of all code because of the huge number of languages and platforms in the repo.
But you still have library code that's depended on so much that you can't easily run all the tests for each change. So you run a train system so that you can run tests for a whole bunch of changes (those on the same train) at once. Developers have to schedule changes to get on the next train. Other code analyzes the failures from the train to try to apportion blame correctly.
Then, because getting on trains can be cumbersome and tests you care about can fail because of other changes, you build random test sampling and smoke test subsets so that you can get some immediate results to attach to a review, etc.
It roughly works, as well as anything can at that scale.