However the "Replica selection" section seems to shed some detail, albeit somewhat indirectly. From what I can gather a probe consists of N metrics, which are gathered by the backend servers upon request from the load balancers.
In the paper they used two metrics, requests in flight (RIF) and measured latency for the most recent requests.
I assume the backend server maintains a RIF counter and a circular list of the last N requests, which it uses to compute the average latency of recent requests (so skipping old requests in the list presumably). They mention that responding to a probe should be fast and O(1).
At least that's my understanding after glossing through the paper.