You’d think there’d be a very simple solution to this—one that I believe Google
already used for a long time, but just never generalized.
That approach: “proof of human work.” Google owns ReCAPTCHA, and every time you do a ReCAPTCHA for Google, you’re doing a little one-time proof-of-humanity for them. But it’s also a proof-of-work; and proofs-of-work that cannot be automated are aggregatable.
In other words, the fact that someone with Google account X solved a ReCAPTCHA, doesn’t just tell you something about who that account is lately. It should add to a sort of “human-proof credit score” for the account, where Google’s systems are more willing to put faith in the user because of all the times they’ve proven themselves human already.
And, for some scenarios, Google does use the aggregate proof ReCAPTCHA represents this way. This is why you’ll never see the Google Search “stop searching so fast” message when accessing Search through Chrome synced to a well-used Google account; why you’ll get a ReCAPTCHA portal from them instead if you’re not logged in (you’re being asked to build the credit score of your IP/session); and why you’ll be denied upfront if you perform botlike behavior through Tor (where there’s nothing that can be correlated to give you a persistent credit score.)
Now, such a “highly-proven” Google account could still be heuristically detected elsewhere in Google’s systems as being responsible for botlike behavior (e.g. spamming); but, when such a highly-proven account is flagged, it should go in for manual review. Because — as you say — this is incredibly rare! So this process doesn’t need to scale through automation, the way regular Google processes do. It can be high-touch.
But right now, it’s not. (Or they’re just not even using the high-proof-of-humanity metadata on the account during this determination.) Either way, that’s kind of silly.