Twitter is being investigated over data collection in its link-shortening system
fortune.com
fortune.com
Of course this is a completely broken way to implement a link shortener since it won't work with non-browser tools such as curl. I tried a t.co URL with curl and it returns a Location: header, which means they're doing user agent sniffing. If you need to use user agent sniffing to make something practical, it's generally a good sign you shouldn't be doing that thing.
$ curl -A "Mozilla/5.0 (Macintosh; Intel Mac OS X 10_14) AppleWebKit/605.1.15 (KHTML, like Gecko) Version/12.0 Safari/605.1.15" https://t.co/88MpPkUoJg
<head><noscript><META http-equiv="refresh" content="0;URL=https://bbc.in/2yDY0F5"></noscript><title>https://bbc.in/2yDY0F5</title></head><script>window.opener = null; location.replace("https:\/\/bbc.in\/2yDY0F5")</script>
I had no idea they were doing it that way. How gross.Arguments that implicitly assume everyone receives the same data from a server are frighteningly common. This is extra strange when it happens on forums like HN that also regularly assume the same server might be A/B testing or providing "targeted" advertising - or prices - that is unique for most users.
Any discussion about data from an unknown server should always include some sort of checksum. Without verification everyone is receiving the same data, statements about a server's responses don't mean much.
set-cookie: muc=4673c8f0-5aef-45eb-8e4b-ab06bc59944c; Expires=Wed, 14 Oct 2020 10:10:19 GMT; Domain=t.co
I have always preferred to use it (location.replace) within the same site. Also it allows to better control browser cache policies.
Although it has been almost a decade since, I doubt much has changed.
A 301/302 redirect works just fine for this.
Additionally, they're not actually part of web technology due to Twitter's ToS...
I run a web crawling company (http://www.datastreamer.io/) and we license data to other companies based on what we crawl.
This really opens up some weird situations for us...
If a URL is copied and shared OUTSIDE of Twitter but behind a t.co URL you can't access it without agreeing to their ToS even though the link might be to the nytimes or some other service.
I was initially upset about the GDPR but I'm starting to see the light of day here.
You can't have your cake and eat it too. You can't both be on the Internet but then put up an insane ToS claiming you have rights that restrain Internet users.
It's like standing on the street corner and yelling and then saying everyone around you owes you royalties because they're hearing your copyrighted speech.
Can you explain that further because you pretty much to have to have a ToS if you don't want to get sued to death for any moderately sized website? The WWW is not a complete free for all.
(I don't think rel='noreferrer' is fully supported by all browsers)
I used wget to get a t.co and original link (from the sibling comment) and diff showed no differences in the fetched pages.
--edit--
So HN is not a discussion site then?
AFAICT this is purely to allow for the 'pseudo injection' of the third-party JS, presumably for tracking purposes...
Only question I'd have is why they can't read the cookie server-side instead, but I'm guessing there are cookies on other domains that their JS is looking for? Haven't done web stuff in a few years so I'm behind on CORS-ish pros/cons/knowledge.
Nah, the browser doesn't let you do that. This SO answer suggests it's to pass the Referrer header so that the destination site knows the user came from Twitter:
That wouldn't explain the User-Agent sniffing (curl gets a proper HTTP redirect).
If a problematic link is shared, it can be pulled from the platform without "doing a gigantic grep"
On bitly, you can add a + to the end, to get to the stats page for that link; it also gives the destination.
On the goo.gl links, add .info to the end.
This seems equal or less effort than making a url shortener.
Those websites are not funded by respecting their users privacy. Although I think you mean Discord and not Discourse?
https://moz.com/blog/inside-googles-ved-parameter
If you're looking in your browser's status bar when you hover the link, Google is manually displaying the end destination URL. The link doesn't actually go there directly.
This is using Firefox while not signed in to a Google account. If you're really getting direct links, perhaps if you're signed in or using Chrome they give you real links and track you by other means instead.
Did they fix that recently? Because Iam sure that wasnt always the case.
For EU citizens, collecting data either requires a very strong reason (like not being to operate the service otherwise), or opt-in.
You can absolutely operate an URL shortening service without massive data collection, which means they need to get an opt-in for data storage from every EU citizen clicking on such a link, otherwise they are in huge trouble with GDPR.
So yes, I can absolutely be surprised that they don't seem to care about the law.
https://www.hirokomatsushita.com/
as it would come out as:
That's the cover story.
Ah yes, I remember when Tinyurl first came into play - people were extremely hesitant to click anything behind one because so often it was a goatse.
from a Don Norman design-of-everyday-things perspective the design is completely non-discoverable https://en.wikipedia.org/wiki/Affordance#As_perceived_action...
Also note that a ".info" suffix might sometimes be easier to type. [1][2]
Too bad most URL shorteners don't support them. :(
[0]: http://goo.gl/vulnz+
https://developers.googleblog.com/2018/03/transitioning-goog...
(although evidence of this happening in practice hasn't crossed my radar, but it's probably because I just don't click those links in the first place)
How does link shortening do that?
Also -- anyone who views a copy/pasted version of this content won't get this protection.
It's not more secure, but it's not less secure and it doesn't break the web. It also shouldn't add an appreciable amount of complexity, given that most of the heavy lifting to sanitize, parse, and format UGC content already happens on the server. E.g. if you're already turning UGC snippets into an AST on the server so that you can cleanly syndicate them in different formats, having the AST generate some js around URLs isn't a big lift.
I still don’t understand why you think url shorteners break the web.
How do you know where the links resolve to once FB goes out of business?
Given the fact that there are still lots of people whose entire job is translating 6,000-year-old grocery receipts from Sumeria, it's not at all unlikely that tweets being written today will be still be widely studied and considered important 10,000 years from now. But those short links are unlikely to resolve for even the next 20 years.
Also, adding js should no longer add more attack surface now that we have things like subresource integrity in addition to CSPs.
Though I agree it's not ideal.
Shortened links become trackable by a third-party (less secure), obfuscate the real URL (less secure), and can be brute forced easier: https://www.schneier.com/blog/archives/2016/04/security_risk...
[0]: https://twitter.com/8x5clPW2/status/1043236568394280961
My Pi-Hole blocks twitters analytics endpoint so I get an annoying name resolution failure when clicking t.co links
Who is David Rees? Glad you asked...
https://motherboard.vice.com/en_us/article/vvvve8/motherboar...
http://spaaaccccce.com/Gotta_go_to_space_Theres_a_star_There... (link to HN homepage)
Full URL since HN abbreviates it:
http://spaaaccccce.com/Gotta_go_to_space_Theres_
a_star_Theres_another_one_Star_Star_star_star_
Star_Space_Are_we_in_space_Oh_oh_oh_This_is_space_
Im_in_spacehttps://web.archive.org/web/20181015144639/http://fortune.co...
What does this even mean? It's a weirdly formatted sentence that makes it sound like Twitter has the magical capability of determining your location... just like everyone else on the internet can with a geoip database.
https://addons.mozilla.org/en-US/firefox/addon/another-twitt...
it's also interesting to think about why Google shut down goo.gl, in light of this and the Google+ story.