Gmail is opening and caching URLs within emails without user intervention (2019)
support.google.com
support.google.com
/validate/email/1d00a5c2648c211befd33f5a8a7cbfab
The token is cryptographically strong and disappears after access. It can't be guessed and no one but the email account holder should click it, but I am seeing the URL accessed multiple times from multiple IPs, so I investigated.
Turns out, if the user provides a Gmail or Gsuite email account during registration, Google clicks the link. I was curious if others on HN had encountered this and how they dealt with it.
A malicious page that knows the url used in the email could open the url from the email in a popup. The js will execute in the popup, and do the POST request. It doesn't matter how much csrf protection you have on the POST step, if anyone can trigger it with no user interaction just by opening some page with a GET request.
Sorry but this whole scenario is just ridiculous. If somebody can access your email it is already game over. It doesn't matter what web technology you are using at the that point, user interaction or not.
If a "malicious page" knows the URL it doesn't matter at all because it means it is capable of arbitrary code execution and that point it could just exfiltrate the URL to somebody or a Chromium instance to perform the user interaction. Actually if the page can open a popup I think it could also execute JavaScript within the context of the page and perform the user interaction right there.
Scenario a) No authentication-y bits in the url. User goes to the url, site checks if the user is already logged in via a cookie. If so, does the POST request.
Typically in this case the urls are easily guessable, so that's an easy CSRF. In principle they could be made per-user (some sort of HMAC on a user+timestamp). In practise, I think its fairly common for websites not to do that in this sort of situation.
scenario b) The url contains some sort of nonce, or signed assertion that automatically logs the user in. I think this is fairly common in email urls, because web developers want people to be able to take an action just from clicking the link in the email, even if the device they read email on is not the same as the device they normally use to interact with the application. I also think this is the scenario that applies to this discussion, since it was started around talking about the issues caused by google auto following links, and in scenario A, google auto-following links would not be an issue.
Of course, in principle, its possible that the authentication bits in the url, just authenticate that action, and don't generally log the user in. In practise I think its really common to just generally log the user in, since most sites want people to stay on the site once the user does anything, and not just immediately exit after the email action is completed.
These urls are typically not guessable. A small percentage of users do tend to smatter these across the internet (e.g. https://urlscan.io/), but ignoring that, this is a login-CSRF. That is, an attacker can generate their own such url, and force their victim to log into the attacker's account.
The impact of a login-csrf tend to be very application specific. Sometimes its kind of minor, but I have definitely seen cases in major websites where a login-csrf can lead to a full account take-over of the victim's account.
> Actually if the page can open a popup I think it could also execute JavaScript within the context of the page and perform the user interaction right there.
This is only true if the pop-up has the same origin as the site that opened it. Otherwise there is just a very limited API (Basically, postMessage(). Also both sides can change the current url of the other side, which is a bit nuts). Also there is now a new http header, Cross-Origin-Opener-Policy that affects this.
Why not send a short random code by email for the user to then copy into the sign-up form they were in the process of filling in?
What I do think would be reasonable is having a well-labeled link which takes you to a confirmation form: someone can follow it easily and choose to submit it with far less friction and it leaves standard web semantics intact.
All of that may affect the sign-up rate.
Ragequitting is one way to exit a process, but just not going to the next step from distraction is surely more common.
The original discussion was about clicking links vs reading and entering the code in sign-up confirmations. The former takes less steps and is easier to complete. Power users with unusual habits might disagree. But if they complete the sign-up anyway, it makes more sense to focus on regular users.
Switching to no confirm obviously changed this to 0% drop off of people who clicked the link, but the number of people who clicked was the same.
It was curious to me why people wouldn’t go through with the confirmation step, but never learned why. We just learned that for some reason more people click once instead of twice.
So it doesn’t matter to me if they were bots or not.
For example, 100 users clicked on the first link, 80 completed, and had normal account activity (clicking on stuff, uploading and downloading things, etc).
100 users clicked on the second link and then had normal account activity.
Maybe they were all bots, but they seemed human based on the “normal activity.”
Extra steps are hard and boring and people don’t want to do them.
I consider myself a savvy user and I want to click a link. Not click a link, then look up a code from the email, then paste, then click submit.
I’d live with having to manually click “I’m sure I want to unsubscribe” or something.
This is most annoying when the site wants me to type in my email address to unsubscribe. I have lots and lots of different email addresses that funnel into a single one. When the site doesn’t put my address in the “To” field, I dont know who they sent to.
Services should be respectful of users time.
Entering emailed or texted codes is becoming more common with 2FA for banking, PayPal etc. anyway so I think most people are going to broadly manage.
UX is important, and I think saying “suck it users, I’m going to use GET the way I think is write” is not a positive way of thinking about it.
I think the problem is just the mechanics of POST not being allowed in an email, so if there’s a way to POST from just clicking on a link I think we should use it. But there’s not, so having a GET that triggers something is the least bad thing. I like it better than javascript and forms in email. And better than autosubmitting, hidden forms on load.
One of them mentioned that you can continue keeping things as a 1 click solution with the token in the URL, but instead of doing the destructive action upon visiting the link -- instead you would get sent to a page with a form where the token is put into a hidden field that gets auto-submit as a POST request with Javascript.
This way from your POV it's a 1 click solution. You only waste a second waiting for the redirect and if the user doesn't have Javascript enabled you can <noscript> the field as being an input field which is pre-filled out based on the value from the URL (this can be done server side).
Now everyone is happy, unless gmail is going to go as far as auto-following redirects with JS enabled.
- You GET /reset/abc123
- Your server responds back with a page that has a form
- There's a hidden field with the token
- Javascript kicks in and on page load executes the form as a POST request
- Your server responds to that POST request and does whatever it needs to do
All of that is kicked off by gmail visiting /reset/abc123, and now it comes down to whether or not gmail's pre-visiting code will run the JS on the page. If not, then the above workflow fixes this issue, if it does then you're in the same position as avoiding all of this and having a GET /reset/abc123 perform the destructive action.I've become increasingly suspicious of this practice
if my email address is in the URL, why don't you autofill that email box for me? if it's not in the URL, why aren't you fetching it from your database using my unique hash in the URL? do you even keep any records of email subscription preferences? am I just signing up for more shitty spam by giving you my email, again? am I just being marked as 'active' i.e. fresh meat, somewhere in the spammiverse?
these questions become more poignant and the suspicion more fiery when, low and behold, it turns out you are still subscribed
I don't fill them in any more, I just block the sender
don't get me started on the "it may take up to 28 days for our systems to register your desertion" bollocks -- I'm not working a contractual notice period, or running a lap of dishonour. it's a bitflip to 'false' in the 'is_pesterable' column, else a respectful deletion. it takes microseconds, not weeks!
I typically never type in any information because I assume the site doesn’t know and wants to know.
I’d rather just set up a kill rule on my end than risk getting my email on one more list.
Kinda hard to pre-scan a URL if you can’t provide the password for it.
This is not just a Gmail thing. Most corporate mail filters visit a link and scan for malware as a feature.
Make the emailed verification/reset link (GET request) idempotent (1 and >1 request has the same effect).
Have the link just present an interface for the user to take the next step. In the next step make a POST request that actually commences your verification/reset process.
In all likelihood you'll want expiry logic (let's say it's 30 minutes) - if you store the token with a created_at timestamp on the server you can have your verification/reset process check that now < (created_at + 30 minutes)
If expired, provide a UI for the user to request a fresh verification/reset email.
A better solution is to assume that some middleman (email server or client) will always try to access links in the email. Instead send the user a code and have them manually enter it on the linked page.
I'm not a web developer. Out of curiosity, why is that?
These assumptions are so baked into web software that while assuming a GET request won't do anything zany or overly stateful is probably fine, assuming the same for a POST request should probably be considered negligent.
74.51.221.37 - - [19/Aug/2021:22:05:16 +0000] "GET /validate/email/1d00a5c2648c211befd33f5a8a7cbfab HTTP/1.1" 404 0 "" "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/92.0.4515.107 Safari/537.36"
$ dig -x 74.51.221.37 +short
cache.google.com.
So a link that is only valid once would be affected. Restricting the validity by time is a good way to solve this while still maintaining decent security.
We solved it by having a screen with a confirmation button , then later we added javascript to show a loader page over the button and click the button automatically.
See https://blog.healthchecks.io/2019/12/preventing-office-365-a... for someone who did have this experience.
People don't tend to report problems like "my account was activated sooner than I expected"
The verification process serves only you, the administrator. To everyone else it's a tedious obstacle.
Nobody will reach out to you to say "I made it through the registration process just fine but it was slightly less burdensome than I expected, is everything OK?"
If you want to know if Google is hitting your activation URLs, check your access logs. Your users will almost certainly not realize it happened. Even if they do notice it, there is no impact on them and no motivation to inform you. You would have to be extremely lucky to hear about it from a user.
And I'll instead refrain from using sites that inappropriately provide bare get URLs that are really state-mutating booby traps in disguise.
If everything send to gmail is opened upon arrival and cached, you know nothing about when or if the recipient actually opened the email.
What is being described here is likely being done for some other purpose.
> Instead of serving images directly from their original external host servers, Gmail will now serve all images through Google’s own secure proxy servers.
Https://gmail.googleblog.com/2013/12/images-now-showing.html
(Because waiting for Gmail to load on a laptop is painful, whereas on my phone is shows up as a push notification within seconds)
https://www.w3.org/TR/capability-urls/
It is also a good way to communicate between two parties that don't want to have user account in any service, we constantly request input from B2B customers by providing forms with a capability URLs. An no, we don't want to use an identity provider. Maybe good ones like auth0. Amazon Cognito is pretty decent in my opinion, but Amazon is also big tech. Industrial espionage is something real for that matter.
We have mail providers that respect privacy, just saying... I don't understand the love for Gmail at all, especially when you use a mail client, which I would heavily recommend to everyone.
Ironically a lot of security scanner also do follow links. Understandable, but I just hope they don't plaster the logs too much...
On the other hand. It validates the email address more quickly, so you could even refresh/poll when it's verified automatically
- `sec-fetch-dest` header is present (HUMAN)
- `accept` header is present (HUMAN)
- `from` header is bingbot(at)microsoft.com (AUTOMATED)
- `user-agent` header includes BingPreview (AUTOMATED)
HTH
I have my browser configured to retrieve the page only and no additional requests for CSS, images, or javascript.
https://example/com/token?forBots
https://example.com/token
https://example.com/token?forBots
Hopefully any automated systems will open the first or last link first, so that you can save the request info and filter based on that. In case requests come out of order, you can always add a small delay to the "human" link before responding.I haven't yet gotten to implementing any of the authentication on my current project, so I might be missing something really basic.
The next best thing is to set a cookie when requesting the magic link, but the downside (or upside?) is that it will be valid only for the browser it was requested with.
But in those cases there may be an automatic POST after you travel to the link, so it wouldn't be triggered by gmail looking up the url.
For this use-case, it seems like even an automated link click would be a good signal of a deliverable email address.
If you're building your web app in development, chances are your links will have localhost as their hostname which wouldn't trigger a visit from Google. You may also end up having an in memory fake email server to not even send the email in dev too (lots of web frameworks have solutions for this).
Checking the user-agent might work but I'm not a fan of this method because now it sets you up with having to keep a list of all known agents for every email client / service that might pre-visit URLs.
74.51.221.37 - - [19/Aug/2021:22:05:16 +0000] "GET /validate/email/1d00a5c2648c211befd33f5a8a7cbfab HTTP/1.1" 404 0 "" "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/92.0.4515.107 Safari/537.36"
I suppose you could somehow block cache.google.com but I suspect Microsoft and others do similar things.
I sent a link to a large file over Viber and immediately some ip connected and started downloading. Stopped at 350mb of around 3.5gb. I get that they want to show thumbnails or whatnot, but they just don't discriminate between content types.
I'd be fine if the fine print in the EULA provides a guarantee that the feature scans content solely for generating previews, and that M$ keeps no copy of it, etc....But, I'm sure I'd go blind looking for such text in the EULA.
<link rel="prefetch" href="/actual_validate/email/1d00a5c2648c211befd33f5a8a7cbfab?prefetch=1">
<script>
location.href = "/actual_validate/email/1d00a5c2648c211befd33f5a8a7cbfab?js=1";
</script>
<noscript>
<a href="/actual_validate/email/1d00a5c2648c211befd33f5a8a7cbfab?js=0" rel="nofollow" class="btn btn-primary" role="button">Click to confirm your account</a>
</noscript>I assume so, because of an old trick where query strings are used for ad-hoc cache control as in /style.css?1629472765
Google will also pre-load all the images in your email too.
You shouldn't take any write action to your database just based on a URL being visited. Take them to the verification page and ask them to sign in or submit a form with the token pre-filled.
PLEASE disable automatic loading in Gmail settings. Don't let the idiots use unethical, stalkerish e-mail read receipts.
Not saying Google is virtuous here -- it only serves to enforce their advertising monopoly -- but I don't see how the image caching in itself is a bad thing.
Then it’s all noise.
According to some articles I've read, the marketers can still name the images unique per user.
So when Google's caches query for it, they still know it's you.
I will keep "always load images" off as usual in Gmail.
https://arstechnica.com/information-technology/2013/12/dear-...
I'm not sure if that's the case though.
"Why did you click that link? But, I didn't."
Anecdata: just running curl on one of those test URLs will trigger a failure and can result in a long discussion with HR and IT.
> > Many phishing test as a service companies will report clicks vs. people who actually interact with the page.
> Would you mind naming some?
KnowBe4 is one such company. Their emails are also easy to spot because they'll have an X-PHISH-TEST email header.
1. https://explain.depesz.com/ great site - I highly recommend it for getting into postgres performance analysis.
That is patently not true, otherwise you would be dealing with utter chaos as you interacted with the internet. If, as the OP claims, Gmail actually _is_ doing this, then that is worrying but it's not the general case.
Google pre-loads and caches images, which many people consider problematic, but they're not pre-fetching URLs.
Refer to these two incredibly recent posts to understand why:
1. https://news.ycombinator.com/item?id=28192269 - How to prevent email spoofing, using an unholy combination of silly standards
2. https://news.ycombinator.com/item?id=28194477 - Email Authenticity 101: DKIM, Dmarc, and SPF
Microsoft also scans links sent in encrypted Skype messages. https://arstechnica.com/information-technology/2013/05/think...
I don't think that's true. It should be pretty trivial to know whether a click came from a user or google.
So, your question is how to evade security alerts for actions with potentially significant side effects?
Well there's already a big ol' button in an email that says, "click me to register". The end user doesn't really care about the implementation. If one pops up a security alert (the POST form) and one doesn't (the simple link), how do you think everyone implements that big ol' button, 100% of the time, for 100% of everything?
I wish email didn't work this way, but as far as I understand, this is the lay of the land. If there's a better way, I'll be happy to implement it in the system(s) that I have control over.
I'm really asking for engagement within the community with help solving this sticky problem (if it wasn't clear). If link caching is this prevalent, what to do about it, for things like registering via email?
We would get spikes of thousands of requests per second from Microsoft IP addresses, which after some googling were linked to their threat detection.
Just send email in tranches.
If the marketing department is sending too many e-mails for the web server to handle, then the volume is probably out of proportion to the company and what they are doing is probably just spam.
I have a common name gmail address, and I get verification emails all the time that I never open. If websites keep emailing me after that, then I rightfully mark them as spam.
Or even simpler, you make it so the user has to click a button to POST the request. You've had to do this for years, now.
I would assume though, that Gmail is smart enough to go, "oh hey, looks like a verification link, maybe I shouldn't touch it"
There's a similar but different problem with the reader-supplied "unsubscribe" button. This usually uses information found in the header of the email message - "List-Unsubscribe", but guess what also gets prefetched sometimes? Enter RFC8058 and, "List-Unsubscribe-Post"and another email kludge to throw on the pile,
If it's OK for your mail service to open one secret link, where does it stop? Is it also OK for them to spider the content they can reach from that link? Now they are potentially gaining access to all kinds of possibly sensitive information that they would not have been able to reach except for spying on your email. And if that's not OK, why was it OK for them to open the secret link in the first place?
"Secret link" is an oxymoronical concept. Resource identifiers are exactly that: identifiers. They're not private names, and any design that relies on keeping them secret is inherently flawed. If it's accessible on the openly resolvable web, then the content needs to be treated as if it's public. If your use calls for authentication or authorization, then actually use an authentication or authorization system.
Yes, the public could guess a 128-bit random value and log in - but that's no different from the ability of the public to guess your password, or your session cookie, or your SSL session state, or whatever. Every authentication mechanism is based on "There is a high-entropy value, and nobody but the authorized user has it." It makes no difference from a theoretical standpoint - i.e., in terms of whether it's "actually" an authentication system" - whether the high-entropy value is sent to the server as part of the URL or via a header or via POST data.
(It clearly makes a difference from a practical standpoint, because in order to have a secret link, the link must actually be kept secret. But that's no different from, like, the need to not expose your cookies to third-party requests or whatever.)
Then, when you click the link, if you don't have https, anything between receiver and site also gets a copy of the link. And there are proxys, add injecting ISPs, etc.
Which "random sites" see emails between sender and receiver?
Yes, proxies and ad-injecting ISPs can see the contents of plaintext HTTP. But that's hardly a reason to say that logging into a website with a password or presenting a cookie doesn't count as an authentication system!
That's fine. No one is saying that. They're saying that URLs aren't an authentication (or authorization) system.
I'm curious to know how those advocating a position similar to this think something like a password reset facility on a website should work. We all know security-sensitive systems should rely on alternative methods of authentication anyway, but for those of us living in the real world where billions of people access millions of systems via websites using their email address as ID/fallback, what else would you do that does not rely on trusting emails to be acceptably secret for at least a few minutes?
Let's put it in a statement instead of the form of a question: it does not follow to respond to the quoted part ("URLs aren't an authentication (or authorization) system") with remarks about "those of us living in the real world where billions of people access millions of systems via websites using their email address as ID/fallback[...]".
You (both of you) are confusing the subject here: URLs vs reaching back to drag emails and their privacy into focus. They're different fucking things! Stop responding to comments about one with responses that deal in the other!
Billions of people access websites using their email addresses? Granted! Now say something about URLs if that's what your quibble is and shut up about the emails that the URLs were sent in and whether or not those emails are private. The comments about email are misdirection at worst, and a sign of unclear thinking (and a hazard to confuse others) at best.
Is the argument "Information in an email should be treated as public"? Then how do you validate users on signup in the first place? How do you ensure that someone owns an email address that they claim to own?
Is the argument "Information sent over HTTPS should be treated as public"? Then why does the argument not apply to passwords or cookies?
If the argument is something else, what is it?
This not difficult at all. Playing dumb isn't clever, it's just obnoxious.
The argument, stated amply before, is that URLs are not private.
Email being a private medium or not is orthogonal.
Resource identifiers, on the web[1], are not private names—not even by virtue of the fact they were communicated over a private channel—and they need to be treated as public, full stop. URLs are not private names, simply because of what they are.
> It makes no difference from a theoretical standpoint [...] whether the high-entropy value is sent to the server as part of the URL or via a header or via POST data.
It makes no difference from an information theoretic standpoint. There is no reason, however, to narrowly consider the information content and its entropy and declare that you are done. From an information architecture standpoint, there is a difference.
> But that's no different from, like, the need to not expose your cookies
It is different, for the reasons above.
(Every entropy-based cryptographic protocol also begins with observations how hard it is to do something in practice, and is then founded on exploiting those side effects. To describe a system and then wave away concerns that it is merely unfit "from a practical standpoint" makes it a failure of a design. It is fundamentally at odds with not just the evaluation criteria that protocols fit for use are measured against, but from which they are born.)
For instance, I could argue "Fingerprints are not passwords, and they need to be treated as public, because of what they are" - because I can finish that sentence with "and what they are is a pattern that's left on every single random thing you touch, and is also immutable and impossible to rotate."
What's the analogous thing for URLs?
It's almost certainly true that you do, you're just being dishonest. (The alternative is worse.)
> What is this thing that URLs are which makes them public?
You mean other than being identifiers (universal identifiers, at that)? It's like you've never used or encountered someone else articulating an argument that incorporates (or would be appropriate to incorporate) the phrase "by definition" before.
If your security protocols are compromised by the card catalog or the Rolodex being invented—compromised not by knowing the contents of a given resource, but by knowing the correct way to refer to or otherwise describe the identity of that resource—then you don't really have very good security in your protocols (particularly in a world where those things have already been invented).
> What's the analogous thing for URLs?
What? What a bizarre request.
The next time someone asks you to make your case in terms of bad analogies just because they can't get away from using them themselves, you can go ahead and say, "No, thanks. I'll pass."
The invention of the card catalog and the Rolodex does not compromise anything, because the card catalog and the Rolodex simply catalogue information that is public, but in a poorly-accessible format. No card catalog can find the name of an unpublished, self-printed book that is sitting in my house. No Rolodex can determine the extension of the direct line at my work. Since I am claiming the URL in question is not public in the first place, I am claiming that it would not end up catalogued.
Can you explain, clearly, how this URL would end up catalogued? "By definition" is not an argument.
> I am disputing that you are interpreting the definition correctly.
And I question whether you've actually made an attempt to grok the subject as a matter of definition, rather than substituting your synthesis (based on an experiential mental model of the subject derived from firsthand inference) in place of what URLs actually are actually supposed to be. (Meaning the playing dumb comment would be apropos here as well.)
> Can you explain, clearly, how this URL would end up catalogued?
The mechanics of how don't have to be explained, because that's how definitions work (whether you accept it or not). Explaining how is not a pre-requisite to what.
But if you're really dying for some missing insight, how about pausing to demonstrate some awareness of the catalyst of this tedious exchange: that a company that was founded on the basis related to cataloguing documents and their public identifiers is (shocker) doing that, right before doing things to/with them—and this has led to people who built up a model of the world similar to yours getting upset because the mistaken assumptions that went into building that model conflict with they're now being told is happening.
The user experience with this is terrible if email is not set up on the device you want to log in from
But the fact is, many systems do work like that and many users do prefer it. I'm taking a pragmatic stance here because assuming the messy, unpredictable real world always follows some theoretical standards at a scale of billions of people and millions of organisations is very predictably going to give bad results in a lot of cases.
You visit a URL, and some JS POSTs to `https://[youraccount].us1.list-manage.com/unsubscribe/post` with a body containing your subscription and list IDs.
I'm not sure what prevents crawlers executing JavaScript on that page and triggering the unsubscribe action anyway, though, unless it's just that email crawlers don't execute JS.
Google could of course send the POST as well but then at least they're violating the HTTP standard.
74.51.221.37 - - [19/Aug/2021:22:05:16 +0000] "GET /validate/email/1d00a5c2648c211befd33f5a8a7cbfab HTTP/1.1" 404 0 "" "Mozilla/5.0 (Windows NT 10.0; Win64; x64) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/92.0.4515.107 Safari/537.36"
$ dig -x 74.51.221.37 +short
cache.google.com.
216.99.127.196 - - [20/Aug/2021:16:25:42 +0000] "GET /validate/email/2591b346e5b8b435bdde54d797fe23a9 HTTP/1.1" 200 811 "" "Mozilla/5.0 (Windows NT 6.1) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/92.0.4515.107 Safari/537.36"
To put succinctly what others are saying: it was never not a problem; it never should have been happening.
It specifically mentions in section 3.2 that mail receivers are not to crawl this URL without user consent:
> The mail receiver MUST NOT perform a POST on the HTTPS URI without user consent. When and how the user consent is obtained is not part of this specification.
I haven't seen any statistics on how widespread adoption of this RFC is among the major mail providers, though.
Bloody obvious in retrospect, but it took us an embarrassingly long time to realize that we were leaking mailing list subscribers because of these one-click unsub links.
It didn't take some companies long. They were just a bit more, uhhh, shady about the knowledge.
Theory: it preloads links in the background.
Practice: some old bb showed a (delete) link after each post and a (ban) link, among others, next to each user if you were logged in as administrator. All of these sent GET requests because some developer hadn't read that part of the standards, and there was no "are you sure?" prompt either.
What I think is happening here is that gmail is scanning the content of each link in an e-mail for some subset of {malware, fraud, phishing, child abuse, other bad stuff}. This is a feature if you're a non-techy user who clicks on phishing links, I suppose?
https://blog.moertel.com/posts/2005-05-06-google-web-acceler...
Use the nofollow value when other values don't apply, and you'd rather Google not associate your site with, or crawl the linked page from, your site.
It seems like crawling a nofollow anchor tag in an email breaks this rule. Am I reading it wrong, is there an exception for emails, or is Google being inconsistent?[1] https://developers.google.com/search/docs/advanced/guideline...
There is no actual rule stated in the quoted material, and it describes it aa a mechanism for specifying a preference for how Google handles the link when Google encounters the tag on the creator’s site, which a user email on Gmail...isn't, even approximately.
There are good use cases for links with single-use tokens in them. Companies sad and desperate enough to suck as much data from you as humanly possible (i.e. every single news letter with tracking links) ruin these use cases for everyone, assuming they are indeed the reason Google is implementing this feature.
Given that many links (like, again, from password reset emails) give the person who clicks the link instant access to your account, I'd say this behaviour goes further than just filtering out trackers. Google has no business opening my Slack account or entering the change password page for the services I use. The stalking prevention they apply to external images and such is fine in my opinion, but links to external web pages should be left alone. You never know when Google accidentally clicks a link that says "confirm order" or "unsubscribe" because its magical AI misinterpreted the contents of an email.
Google is following the standard. People who take actions based on GET requests are not. Sure, mistakes happen out of ignorance, but they should be fixed.
Maybe, but then several decades passed and now billions of people don't use online systems the same ways any more. It used to be that I could send a legitimate mail to a friend or family member and not worry that whatever mail system they use would refuse to deliver it to them because my system didn't jump through several not-quite-standard hoops that didn't exist when the email protocols were defined. Google seem fine with discarding the historical standards on which most of the Internet is built in that situation.
In any case, if a communications service is going to snoop on your private communications and take actions that would be impossible without spying on you, I think the burden is 100% on them not to screw anything up for anyone, ever.
All big email providers know this and know that GETs should not affect their unsubscribe/confirmation links. It's part of their job.
You're talking about an idealised, theoretical world. I'm talking about the real one, the same one where big mail providers routinely ignore valid standards themselves in their efforts to fight real world problems like spam and identity theft.
Again, if someone is going to help themselves to private information then the burden should be 100% on them not to screw anything up for anyone, whether or not that anyone was following any particular set of rules.
Google isn't the only vendor you need to worry about. Other mail agents, browsers, or proxies will sometimes prefetch URLs. On the web you often don't know what software your users are running or what it will do, and there's no way of knowing what they'll do in a year. Users have the right to run whatever mail software they want. But if you follow standards then you have a better chance.
If you don't, it's a risk. Sometimes vendors do go the extra mile to tolerate other people's buggy code, but not always.
You can decide to blame everyone else if you want, but if you use GET requests for user actions, something will likely break eventually.
(Also, consider mistaken clicks, which happen all the time on touch screens.)
That is literally the opposite of what good native email clients have been doing for a long time. They won't even open linked images and the like by default, to prevent tracking.
Some links include automatic login functionality. I definitely don't want Google logging in to my accounts.
Then don't use websites who provide links that provide insecure features.
As to masking your IP when you actually click on the link. How could that possibly work? Your IP is still definitely making its way to the tracking server upon clicking the link. There would be no mechanism for GMail to prevent that unless it rewrote the links to point at its caching server, which would pretty much break GMail.
Sure, pre-fetching would create a bit of noise because your IP would be mixed in with the Google IPs, but as you've quite rightly noted, filtering out clicks from AS15169 would eliminate that noise.
This might be new for Google, but some virus scanners have been prefetching links for decades.
My heart bleeds for the oh-so-poor marketing people whose data, consisting of unwitting human test subjects, has been poisoned.
I have little sympathy for the first poster: those kind of phishing tests are good if your goal is to train your users to think of the security group as an adversary but not much else. If clicking on one link compromises your security, you need to put the IT house in order first (hint: where’s the WebAuthn which completely?) and especially deal with the vendors who are training everyone to think that clicking on obfuscated links is routine.
I get a support request at least once a week from a user who "Doesn't understand why they can't verify their email".
Turns out, their Gmail account already verified it by clicking the link before they opened the email, and they didn't think to try signing in, because they couldn't even verify the email.
¯\_(ツ)_/¯
The POST with valid captcha solution testifies a human clicked purposefully.
Hard no. CAPTCHA is a blight on the web to anyone who values privacy or has accessibility issues.
anything new on this?
IIRC was a known thing and surely was already discussed on here somewhere then.