“Magic links” can end up in Bing search results, rendering them useless
medium.com
medium.com
It takes incredible arrogance to continue using them in order to "improve usability" given all the obvious and common cases where they completely destroy usability. The difficulty for a provider to verify they aren't sending you to a phishing or browser 0day page barely scratches the surface.
And I think that subset is much larger than some would expect.
I sadly can't put my finger on what's so compelling about this, just that my "oh that person should talk to a UX team lead!" meter just went plink
I love them and prefer them to creating yet another account with a password.
It's only annoying if the site is constantly timing you out so that every single visit you need to resend. Why not just use secure cookies to remember the user for say a week?
This... was impossible to do, because by long-pressing on iOS to get the Copy prompt, iOS also goes ahead and opens a preview of the link next to it
That's a yikes from me! So I can sign up on your service as anyone with an Outlook account, without verification?
We had a handy quick decline and accept button on there so they were auto declining things…
I didn’t hate email until I got into web development….
It was a super basic web form too. Probably the most html markup standard thing we have. Nothing strange about it that could have triggered some sort of strange behavior.
How about having a input field asking for the email again to double check?
- It's me, let me confirm my address
- I never signed up for this heap of diamonds
Whenever I (not even a bot) click or follow a link from my mailbox, by accident or on purpose, I don't expect that to validate an account for anyone else, but me, intentionally, using a password I know.
So while I think what Outlook does here is wrong, what these webpages do is simply a bug that should be fixed and shows a lack of understanding of HTTP.
I think that's a bit too much. Nothing in that suggests that they are breaking anything in the HTTP specification. You're right that GET requests has to be idempotent, but the exchange from the single-time use code you get in email with the API token, is most likely behind a non-GET request (like POST). The HTTP server responds to GET requests with the static assets (HTML/CSS/JS), but then the static assets has JavaScript that calls the POST endpoint for the exchange.
At least that's my guess. I agree it's a bug on their side, and they should fix it. But I think it's more of a UX issue than breaking the protocol.
We found at least one scanning service to be fetching the URL with the user agent of a browser, and executing JavaScript on the page.
A lot of our user interactions could be simpler, but this sort of behaviour led to many things being put behind a "go" button.
I only wonder how long before scanners and search engines start clicking these buttons to activate content on the page so they can scan/index it.
Idempotent != side-effect free
!Idempotent != Reversal of changes
1. You open the link once = your account is verified
2. You open the link twice = your account is still verified
???
1. You open the link once = you're logged in
2. You open the link twice = you're presented with an error that the magic link has already been used.
Request methods are considered "safe" if their defined semantics are essentially read-only; i.e., the client does not request, and does not expect, any state change on the origin server as a result of applying a safe method to a target resource. (RFC 7231, § 4.2.1)
Bing indexes it. This is my first major security incident and I have no idea how to fix this without making everything totally shitty for the users.
The original post is just complaining that the malware scanning is visiting the links.
They come to the following conclusion
>This effectively makes all one-time use links like login/pass-reset/etc useless.
Which we all know is not true because sites like onetimesecret.com allow for entering a separate password to prevent this sort of thing when it does happen.
It would be an interesting discussion to talk about what Microsoft's whitelisting process looks like, but the original article doesn't seem to understand what is going on well enough to drive the conversation in that direction.
It's about trying to get to the core of the issue, not just the random speculation going on in the article and in this comment thread.
The point I was making is that someone should research this instead of relying on wild speculation as the basis for the conversation.
All pages with one click links should have no index follow or no index no follow. Your seo consultant (if you have one) should have advised you on this.
I am not saying this excuses the privacy violation but just suggesting there are things we can do...
There's also RFC 8058: https://datatracker.ietf.org/doc/html/rfc8058
Dunning-Kruger level of understanding: Never trust the client for anything ever, it's unreliable, everything must be off client.
Never mind the client is literally the interface into your system, so it being compromised is already game over for an application where the user is most vulnerable party you wanted to protect...
Deep understanding: Trusting the client requires a well thought out security model.
If the client is hacked in this case, they already have full control over what the user sees, they can cut out your remote check.
Maybe a good balance would be to hash the root of the URLs and compare those, or use fuzzy hashing on page contents, just so that the backend isn't getting a bunch of private urls that might accidentally get logged somewhere.
Trades detecting stuff hidden behind redirects for less liability on your backend, something to possibly consider depending on functional requirements.
1. User opens Outlook and types in their email and password.
2. The app requests the user's password hash from the server and checks it.
3. Outlook tells the server auth was successful and gets a session token.
The client has a much bigger issue to worry about if the client-side malware scanning has been compromised. Malware could modify the UI/network calls such that your server-side scanning displays a positive result anyway.
You have to trust the client to display information to the user at some point. Link malware scanning that job can safely be delegated to the client. Authentication cannot.
And also use the local client to scan unknown links? They probably dont want outsiders to have access to this code.
Also, mail might not live on the same computer.
I should not be able to take over your life because I compromised your phone which has sms, TOTP app and email.
* use an interstitial page so that the actual activation is a POST request;
* send a confirmation code instead of a link
The password only makes this autentication less secure and it's not needed.
Recovery codes exist and are created at the time before recovery is necessary. But most people are going to lose their codes.
Who is going to remember what 3 things they picked out of 30 years ago?
The problem with that is that the logic is broken. Microsoft cannot possibly know all phishing sites, especially for smaller things. By obfuscating the link the user can no longer verify it by themselves without clicking, but Microsoft will say it's safe. So the user is left with a false sense of security and are worse off.
It only works for huge sites ( e.g. mytwitter.lol phishing for twitter and similar), but drastically lowers the chance of less high profile phishing being caught.
That isn't by itself an argument against a good automated system -- I definitely like not having to sift through most of that garbage, but catching the 0.01% should be a routine practice, not something that seems like an insurmountable burden.
If you do so, in Outlook, there will be a pop that shows "Original URL: XXX". This allows users to make a determination for themselves whether the link is safe or not.
I don't trust myself enough to be 100% sure I can decode an URLencoded misleading mess perfectly all the time.
They already hid urls in the username of the url, like www.google.com.unholymessherethatscrollsoutoftheurlbar @ malignantdomainnotgoogle.blah
I noticed this because I generate links with a signed token to ensure integrity and started receving invalid token crash reports in Sentry, always from BingBot..
To fix this I had to move the tokens from the query string into the URL itself to avoid BingBot changing it. eg.
http://mysite.io/do-action?token=shvgaaehr2rnyxhh-391-1 to http://mysite.io/do-action/shvgaaehr2rnyxhh-391-1/
Anyone else noticed this?
I've noticed this too, and I found (in my case anyway) that Bing/Outlook seems to Rot13 the keys of the query parameters - is this what you're seeing too?
I don't see the problem here, all services need to do is add a page that's says "welcome back, $Username, click here to log in!" that sends a POST request to do any serious confirmation without breaking any specifications.
Microsoft claims the visiting not is BingBot but it's probably just SmartScreen system checking for malicious links/downloads/etc. like many cloud integrated security products do these days.
I can set my browser to pretend I'm BingBot, you can't derive anything meaningful from the user agent. Unless you find your secret URLs in Bing's search results, your secret links aren't actually being monitored by a search engine.
For a bit of added "fun", Google will do the same, but if you add a page to robots.txt and set noindex then they won't process the noindex parameter and external indexing sources might still generates search results: https://developers.google.com/search/docs/advanced/crawling/...
This is a can of worms.
The recipient of the email, or their employer's IT department that is paying another company for mail services.
If you send me an email with a link then I do believe I have the right to send that link to a third party service that can validate that it's not malicious. If I decide to sign up for a mail service that promises to protect me from phishing emails, then I [0] expect said service to read the emails I receive and examine the links within them. I would be upset if the service used the info I share with them for purposes other than keeping me safe, though.
I readily admit that I have, at various points in my life, signed up for services without reading the entire TOS that I agreed to. I try to choose companies that I feel I can trust to not abuse me too much, and sometimes I avoid certain services because I don't trust the company behind them enough to respect my privacy.
[0] I acknowledge that not everyone is as knowledgeable as me, and many people might not realize that this is how the protection works. So if the argument is more education, I'm in favor.
That sounds like a narrow view of email. This has never ever been true. Corporate firewalls have always opened links, and many users use tracking blockers in their email provider that automatically opens incoming email and detect tracking cookies. I cannot stress strongly enough that you cannot rely on only one "user" clicking a link.
Email isn't WhatsApp, there are probably at least three or four parties who will scan the link in any way they like. If you control your side you can make sure there are only one or two parties scanning the email en route, but the number can never be guaranteed to be zero without workarounds.
Who gave the mail server permission to access the page? The person who set up email on the domain. If you don't trust the hostmaster, don't send email to that domain. What if it contains copyrighted materials? Well, you just shared a plaintext link with a whole bunch of people, depending on your local legislation you may be in trouble.
You can't even expect a link clicked once by a single user in a browser to only appear once on the server side. TCP connections get dropped and retried. This isn't some kind of philosophical interpretation of a mystical protocol spec, this happens in real life. If you use POST/PUT/whatever requests, the user agent will prompt the user if they really want to repeat a request; this protection has been built in for years. It's just how browsers work and how they've been working for decades.
If your recipient is behind a proxy, the link may be visited several times each hour for up to a month while the proxy refreshes its cache. This was a more prevalent problem back in the day, these days web proxies are mostly a thing of the past; however, proxies still exist, and if you don't pay attention to those things they will bite you in the ass.
In real life bugs happen. That's fine in these cases, web dev isn't exactly rocket science, bugs are tolerated and can be fixed. The bug here isn't the fact that links get visited twice, though: the bug here is that the developers who set up their magical links forgot about idempotency when they wrote their code, or they chose to ignore the problem because they never ran into it themselves. Either way, the responsibility to get it fixed isn't on anyone but the party violating the spec.
As a workaround, S/MIME or PGP should work around most of these problems as intermediate servers can't see what's going on. What the client's machine will do with the decrypted message is still up to interpretation, of course.
Alright, so imagine this: we have two endpoints GET "/page" and "POST /increment". Making a POST request to "/increment" increments a counter kept in memory and returns it's new value. The GET "/page" endpoint returns a HTML file, which contains JavaScript code that when executed, calls the "/increment" endpoint.
Are we now breaking the HTTP specification saying GET requests has to be idempotent if we visit "/page" in our browser? I think not, but this is sometimes how pages are implemented, which robots are gonna have to deal with, as otherwise many would consider it broken.
Don't get me wrong, I think it's a shitty implementation as well. But is it breaking the HTTP specification? Unlikely.
"In particular, the convention has been established that the GET and HEAD methods SHOULD NOT have the significance of taking an action other than retrieval. These methods ought to be considered "safe". This allows user agents to represent other methods, such as POST, PUT and DELETE, in a special way, so that the user is made aware of the fact that a possibly unsafe action is being requested.
Naturally, it is not possible to ensure that the server does not generate side-effects as a result of performing a GET request; in fact, some dynamic resources consider that a feature. The important distinction here is that the user did not request the side-effects, so therefore cannot be held accountable for them."
Personally, I think web crawlers like Bing shouldn't be executing javascript at all but front developers can't go without their client side rendering frameworks so search engines are more or less forced to.
As for a security mechanism, you want to emulate a browser as closely as possible to detect tricks like redirects from safe domains to attack domains and obfuscated URL crap. I'd expect any automated, non-interactive code to execute in a security analysis sandbox.
Is this breaking the standard? Who knows. What is a cloud antivirus but a web user agent running in a data center? The email protocol doesn't specify how the client should deal with links, the robots.txt only works for spiders, not for manually submitted URLs like those clicked in emails, and without a noindex tag you're going to see your page indexed by the mail provider company regardless of what your robots file says.
I think in theory your solution solves the spec breaking problem, but it doesn't solve the problem in practice because there are many other components for which there are no standards and defensive programming is required.
BingBot scans the link, gets a dummy page with 'clean' content, Microsoft delivers the email message to the user, user clicks through the link with actual browser, gets phishing / malware content...
It's not just MS. Lots of enterprise email security stuff works like this.
Same if you paste a link in Slack/FB/Discord/Twitter whatever, they will visit the page to create a preview. GET requests shouldn't have side effects.
I don't see that in the robots.txt https://shoprocket.io/robots.txt
User-agent: *
Disallow: /cdn-cgi/l/email-protection
Disallow: /login
Disallow: /register
Disallow: /404
Am I missing something?
(I made a lot of changes today when testing all, including "visit as Bingbot" from their webmaster tools with and without the URL blocked by robots.txt)
Never mind indexing them (ie publishing them at Bing.com), if URLs are disallowed in robots.txt then Bing shouldn't even be retrieving them, even if only to scan the content for malware!
Robots.txt is not a reliable way to exclude pages from search engine indexes. That is not what it is for. It is for controlling crawler behavior.
The only reliable way to exclude a URL from a search engine index is to serve “noindex” on that URL, either with a metatag or an HTTP header, or both.
I must confess I've been sceptial of robots.txt for a very long time (if I want to stop bots I serve them HTTP 403 Forbidden using .htaccess or similar).
Be that as it may, it appears I'm also confused about what robots.txt does and doesn't do.
Assuming you're correct: let's say I run EvilBot which scrapes sites and want to scrape your site example.com, but your robots.txt only allows Googlebot and disallows everyone else. Am I really OK to:
1. scrape the SERPs from google.com which mention your site ("site:example.com") then 2. using that list of URIs, use my EvilBot to scrape your site, without needing to touch or respect your robots.txt, since I got the list of URIs on your site from Google, not by scraping example.com directly?
If the crawler does then visit your site, it will see your robots.txt and (if well-behaved) obey it and not crawl the contents of the page at that URL. But this does not mean it will remove the URL itself from its index.
Again: robots.txt is intended to control crawler behavior, not search index visibility.
Google's page is a pretty good overview of this distinction:
https://developers.google.com/search/docs/advanced/robots/in...
I'm obviously not asking the question clearly, I'm wanting to stop bots from crawling (it's scraping that annoys me), not search engines from listing URIs.
If I want to completely stop a bot from crawling my site (in the sense of "retrieving my content"), won't robots.txt prevent that? Even in the case of the bot having obtained a valid list of my URIs but not the pages contents from a 3rd party source?
Lets say I email you a list of URIs on my site. My robots.txt forbids all crawlers. Are you allowed to give the list of URIs to your bot and retrieve the content?
I keep hearing this, but our newsletter system has been using GET unsubscribe links since at least 2007 (but probably longer), and we never found a wave of Gmail users unsubscribing, we still have a lot of them. I wonder if this is simply an urban legend, if Gmail tries to recognize unsubscribe links, or if there is something else going on.
The article is about email verification links, which is a pretty clear case where this can be dangerous, but tons of other links can get emailed without being intended for a wider audience.
Besides, the fact that Outlook shares anything related to the content of your email with the outside world is just completely unacceptable.
(Should private links be sent over unencrypted email? Probably not. But lots of stuff gets emailed that's not super secret and yet also not meant to be shared outside the company.)
> Warning: Don't use a robots.txt file as a means to hide your web pages from Google search results. If other pages point to your page with descriptive text, Google could still index the URL without visiting the page. If you want to block your page from search results, use another method such as password protection or noindex.
From https://developers.google.com/search/docs/advanced/robots/in...
I realize that you are just the messenger and not the progenitor of that policy, so not addressing this to you, but: that is ridiculous. robots.txt is basically useless.
https://www.techspot.com/news/80134-google-uses-receipts-sen...
Just a reminder any email left on any online service over six months in the USA is allowed to be read by any law enforcement agency without a warrant.
You'd think these services would have a six-month auto-delete feature but nope.
There's good reason why there was a email server in the basement, everyone should have their email server where at least a physical warrant is needed.
Makes me curious if only the free, online, Outlook does this. There's also paid O365 online Outlook and the fat client Outlook.
Our 365 instance now turns every link into this massive monolith of safelink checking URLs through Microsoft, making literally every email undeterminable if it is a phishing attempt or otherwise without turning to pasting it into one of many online 'decoders'...
Though that is optional and configurable.
Its a shame those links can't have an alttext to show the real link.
Admin has also turned on "You don't often get email for __" warnings that edit the email so that gets included in replies. Very useful when you could a new large cohort of student email correspondents each semester :(
Not adding that was my downfall I think.
A single use link, e.g. for resetting a password or confirming a subscription, will usually show a webpage with a form that does a POST. Once that POST has been performed, the single use link is used up.
Single use links will mostly have a one-time secret that should not be leaked. Mails that contain such links or any sensitive information should be encrypted.
If your goal is to hide your secrets from Microsoft, then send the email encrypted or don't send it to Microsoft's servers at all. This is practically impossible, it at least impractical in most cases. You can't control the hosting provider and software of your customers.
We are monitoring you for your protection.
[0] https://www.washingtonpost.com/technology/2022/06/08/elon-mu...
Don't try this at home, kids. :)
Yeah yeah GET is idempotent and I shouldn't do that blah blah. That's not the point.
Or to prevent some site from being indexed by flooding it with invalid links?
Curious what HN readers think. Is this secure? Sufficient?
https://developers.google.com/search/blog/2011/11/get-post-a...
Automatically triggered POST is not sufficient to keep the bots at bay.
They seem to be implying that only automatically triggered POST are acceptable, but that was also >10 years ago.
With the way things are going, it might be that any on-page confirmation buttons won't be sufficient to keep the bots at bay. Maybe it's time to fight back, check the user-agent, and serve the bots a CAPTCHA?
Now stop complaining and look at the new ads in the Start Menu.
Now, having my private links indexed by Bing is a bit too much!? I sincerely hope OP is mistaken and Bingbot is actually the outlook scanner.
"Block request denied We found that the URL submitted for block is important for Bing users and hence cannot be blocked through Bing Webmaster Tools.
We recommend that the best way to block URLs in this scenario is to add NOINDEX meta-tag to the HTML header of the page."
I sent an email with a unique link in it to my @Outlook.com account 6hrs ago, and there have been no visits to the link. The email is in my inbox (though I have not opened it).
Does this only happen on opening the email (in the Outlook web ui)?
Many email clients generate link previews so that they can display a thumbnail of the webpage. Would seem to be necessary to filter out those referrers from the validate link
Imagine you send someone a "private" link to a file...Bing sees that and indexes it for the world to see. Not cool.
If I have a well-tested crawler sitting here, I don't see why I'd write a brand new one just to fetch previews in outlook...
Email address domain and enclosed URL domain are the same parent domain.
Also, somewhat a different topic, people that run their own mail servers might also still use Outlook.
no what might link to sensitive data leak is fools who store sensitive data on unprotected links
Are there any cloud vendors that don't follow this approach?
https://www.dropbox.com/s/vucien2ns8jktga/denim%20bodywarmer...
I doubt many people realise this when they email "private" links...
Whenever I buy a flight, google puts the date on "my" calendar.
Just lets not pretend Microfsoft is especially bad at this, ok?