Git.io deprecation: Active links to be maintained in a read-only state
github.blog
github.blog
Git.io deprecation - https://news.ycombinator.com/item?id=31162829 - April 2022 (106 comments)
(Plus they're expressly intended to last "forever" and are maintained by one of the few organizations with the multi-century history to credibly make that claim (Havard Law School))
You might have access through a registered institution (e.g., libraries and universities get it for free). Otherwise, the rates for individuals are hidden on their blog.[2]
If you really want to archive and shorten a link at once, https://archive.today (often seen on HN for bypassing paywalls) is another option, but its unknown ownership makes its long-term existence slightly questionable. For general web archival, the Internet Archive (https://web.archive.org) looks like it'll be around for a little longer. But since its crawler and server have their own quirks, I'd still submit links to multiple services for good measure.
[1] https://blogs.harvard.edu/perma/2019/01/07/introducing-indiv...
[2] https://blogs.harvard.edu/perma/2021/09/15/usage-plans-updat...
-----
Speaking of web archive reliability, it looks like WebCite might have recently died for good, after a decade of funding issues and intermittent downtime[3][4]—the first major web archive to fall?
> As we continue our analysis, we may remove individual links that point to spammy, malicious or 404 links.
Please be careful with 404 errors. They can be the result of temporary server misconfiguration.
But one could set up an html on github pages that redirects to the 'real' URL.
As an example, I run an online store, Google checks the landing page urls for ads once a day. If you happen to have a temporary outage (surprisingly common on Heroku) at that exact moment your ads will be deactivated and not run for up to 24hrs!
I’m somewhat suspicious it also dumps your ad back into a learning phase.
No, wrong. Don't use a URL shortening service. Don't use any of them. Especially not in anything long-term.
Connecting a domain name to a piece of software running on any of your devices should require nothing more than a quick OAuth flow to your domain provider to connect a tunnel to the application. You shouldn't need to understand DNS, TLS, HTTP, ports, ip addresses, NAT, CGNAT, etc in order to self-host simple programs like URL shorteners quickly and securely.
No code version:
aws --profile $PROFILE s3api put-object \
--bucket $BUCKET \
--key $KEY \
--website-redirect-location $URLMaybe your internal systems are less prone to such complexity effects that turn "internal only" tools into "customer facing products" accidentally all the time, of course. Just that it's an interesting risk, well worth stating, and it is fun seeing all comments on shutdown notices of exactly one such "internal only" tool that accidentally became a "customer facing product" saying "just build it yourself". "Just build it yourself" doesn't entirely protect you from it being your "internal only" tool that's the next trending HN deprecation notice headline.
The "it's pretty trivial to do your own" piece is trivial if you already have gone through the effort of setting up an AWS account and finding a globally unique bucket name across all S3 users. That option by the way requires entering a CC, personal details, and requires you to pay expensive rates compared with shorteners which are typically free. Not to mention that the resultant URL is still comparatively long
Most people won't do that.
FWIW, most larger companies I've worked at have their own internal URL shortener that they run for all their internal docs and whatnot. That doesn't solve the public use case though.
Sure, most people won't do that. But the user base of github is not most people, it's developers, a large fraction of which work at a company which has an AWS account or equivalent, and I questioned specifically "companies" not "people" using an external service for this.
I've been running https://T.LY URL Shortener for the past 2 years and I'm the creator of the URL Shortener Extension (https://t.ly/extension) that has over 400K users. Let me know if any questions around url shorteners.
Just yesterday I received a text from a shortcode with a "bby.me/XXXX" URL saying my Best Buy order was canceled. Turns out it was a legit order I placed nearly 6 months ago and was on backorder that I forgot about.
I'd love for a shortening service to offer an "unfurl" tool where I could paste in the URL, and it shows me what the redirected site is without clicking it and potentially activating their clickthrough tracking.
For Bitly, at least, you can add a plus sign to the end of a short URL to preview what the destination URL is. I don't know if other URL shorteners have something similar.
I played around with it and there was some seriously sketchy interstitial stuff going on.
Sometimes the content on the link doesn't exist anymore, sometimes the eg. newssite changed their url scheme, so the content still exists, but not on the same url, sometimes the media changed their story, so linking to the "same" article, now paints a different picture, and of course, sometimes old domains get hijacked by spammers... and more related, sometimes everything works except the url shortener.
Honestly, if I ever link anything that will be read after a long time, I only use archive.org/ links, and I cross my finger, that the snapshot will still exist in X years.
I don't understand this practice. It seems more proper to use the original URL and let readers themselves use archive.org when the link becomes dead.
Using the nonprofit's archives when you don't need them seems abusive. On average it should be simpler for the original server to render the page than for the wayback machine to fetch a particular snapshot. What's worse is that I don't think you're the only one needlessly using their archive like this, and wonder if, as a collective, uses like this significantly increase their operational costs.
Something happens, you link to cnn/fox/whatever, and discuss the news story or their comment or whatever, and two hours later, the whole article is rewritten, the headline is different and the subtone is different. This would be ok, if there was a "history" section at the beginning of the article, saying "this article was changed on date X and Y and Z, click here to see previous versions", but usually it's not.
an example widely publicized last year:
https://web.archive.org/web/20210525043631/https://www.nytim...
Then people get angry for newyork times using anti-semitic attacks for left-right politics, and soon after:
https://web.archive.org/web/20210525173242/https://www.nytim...
Different headline all together, same url.
Now if i said that ny times is using those attacs for politics, and linked the article directly, you (visiting the second version of the article) wouldn't find the headline "problematic", and you'd think i was just some far-whatever nutjob conspiracy theorist... by linking to the archive, you see the headline that I saw.
Unless you need to change where the URL shortened link points to, don't do this, provide the full, canonical link instead.
Is it common in papers to attach a copy of a reference as an addendum? That would make sense if the sources cited are web pages.
Thinking about it, Wikipedia should maintain their own index and archive of cited sources, instead of or in addition to relying on archive.org.
Why? Wikipedia and the Internet Archive are partners, each working to their comparative advantage. If Wikipedia has archival needs, they ask the Internet Archive to facilitate those needs.
And the comments elsewhere on this thread should remind us all that "a lot can happen in forever"... My interpretation: eventually, everything will be deprecated. There may be some web sites (github.com, for example) that are "too big" (or "too important") to fail or disappear; but they eventually will.
I feel like there's a different "law" here, involving persistence and memory and history and recordability on the internet -- people want it, but it doesn't come for free, and there isn't necessarily anyone with a profit motive or business plan to make it happen.
I can never remember all the laws
EDIT: also for vanity "rememberable" urls
But for "vanity", the best vanity URL is a domain name. Buy it, own it.
foo.com/sites/blog/signup-for-my-conference-2022-seo-friendly-url
vs
shortenedurl.com/conf-2022
URL shorteners offering vanity urls solved people not needing to deal with the hassle of purchasing a vanity domain via a registrar, setting up DNS records, and waiting for propagation.
URL shorteners came later, and solved the problem of people needing to deal with “more steps” in order to have short memorable urls. Achieving memorable urls in “less steps” is one of the major selling points of url shorteners.
However, url shorteners did not solve short memorable urls — DNS had already solved that. But, as you’ve pointed out, acquiring a domain and setting up DNS records takes more work than simply using a url shortener service.
But really, opaque short URLs are significantly worse than the full URL: you have to go to extra effort to create them, and extra effort to retrieve from them, and you’re adding an additional point of failure, and you can no longer just read the URL.
People have many reasons they want a shorter URL. Twitter is only one of them.
Ok, maybe you mean mobile only: i have a wifi only android phone that I don't use email on, none of those options work.
There are worse things that could happen than shutting down too. Someone with control of an old shortener domain could set up spoof versions of the site that a link used to forward to and use it to harvest passwords or get people to install malware.
For example, if I go to some random product page on Amazon, the link looks something like this:
https://www.amazon.com/Really-Long-Product-Name/dp/PRODUCTID/?_encoding=UTF8&pd_rd_w=garbage&pf_rd_p=long-garbage-value-that-doesnt-matter-to-the-user&pf_rd_r=MORELONGGARBAGE&pd_rd_r=another-really-long-garbage-value&pd_rd_wg=garbage&ref_=even_more_garbage&th=1
That link looks ridiculous and isn't suited for sharing in many situations. For most people, this is the link they really want: https://www.amazon.com/dp/PRODUCTID/
I know there are tools out there that can shorten URLs like that. I used to have a browser extension which did exactly that. It doesn't work for everything, but at least your link wont die because a third-party service shut down or messed up.Because of that it ultimately needs to a site-specific database/algorithm, perhaps with a fallback to the default behaviour like simply cleaning up the most common garbage like (_encoding/usg/etc). I suspect it's possible to use some sort of machine learning to guess the meaningful parts of the URL path/query/fragments, but even for that we need some human curation for the training set. I wish we could collaborate on a shared database/library for that, have sketched some ideas/applications/prior art here: https://beepb00p.xyz/exobrain/projects/cannon.html
I started thinking about it since I have a similar problem in Promnesia (https://github.com/karlicoss/promnesia#readme), a knowledge management tool I'm working on. Ideally I want to normalise URLS, so they address the exact bit of information, and nothing more.
Additionally, CleanURLs to the rescue! https://github.com/ClearURLs
load the page with the original URL
for each part of the URL:
remove that part of the URL
load the page with the modified URL
if the page rendered differently:
put that part back
You'd need to incorporate an ad blocker, otherwise changing ads on each reload could screw it up. Of course, you'd also probably want to hard-code in logic for popular websites like Amazon to avoid wasting time with reloads.And I know there are limitations to the concept. It wont work for pages which require authentication or any kind of session data. It wont work for pages which are intentionally dynamic. It wont work for sites which cannot be accessed by the service. It wont be useful for pages who simply have long URLs. But it probably covers the many common use cases for a URL shortener, and you could always fall back on traditional shortening methods (redirecting) when it doesn't work.
So it basically automates detecting useful bits for a particular URL, but it's kind of time consuming and flaky. It could be very helpful to populate the 'rules' database though, and then this database could be shared with other people so they don't have to scrape.
I guess when I said ML (or preferably some fuzzy algorithm/heuristic), I was referring to generifying rules so they also work on the sites not in the rules database. If humans can detect garbage in the URL looking at a few examples, the computer can too :)
Affiliate links used to catch some heat on HN but https://amzn.to seems to be showing up in more and more comments without too much blowback. It wasn't mentioned much that affiliates could view everything people bought while under their cookie, but I'm not sure if that's still possible.
So perhaps it's open to git.io/your-bank-here-etc
Does it support github pages? Because then it's clearly a massive problem.
To bring down the point, here's the source code for that link shortener: https://github.com/technoweenie/guillotine. It was last updated in 2015. If you don't believe that this is the source code, this was directly linked to in its announcement: https://github.blog/2011-11-10-git-io-github-url-shortener/. If the extreme lag (in an attempt to save the links) is indicative if its backbone, it's probably just running in a single server, which is very likely to be horribly outdated. It might not be due to the cost of removing the malicious links, it might be that Git.io as a platform has a cost in of itself (even excluding common things like domain and hosting costs).
You just proved the point.
https://github.blog/2022-04-15-security-alert-stolen-oauth-u...
“The applications maintained by these integrators were used by GitHub users, including GitHub itself.”
* Technically the links are not publicly listed, which might jeopardise some obnsscure but technically-available repository, but it doesn't store private data.
From the original post: “due to the security of the links redirected with the current git.io infrastructure”
What does “security of the links” even mean? Disclosure? Tampering?
[0] https://github.blog/changelog/2022-01-11-git-io-no-longer-ac...
Isn't this how it worked?
Which is the stated reason why they're killing it off. It was being abused and Github/Microsoft didn't want to take the time to maintain/moderate it.
For example https://github.com/torvalds/linux/blob/9e02977bfa/kernel/dma...
I don't find this ugly, it's even very human readable. There's all info in the url that you even can use in the future if GitHub ever goes away.
Actually, I would use this argument that people should refrain from using url shorteners in papers for this very reason: It is anti-reproducible.
At some point I did a fairly automated from Apache but otherwise the redirects have been zero maintenance.
Notable, they I have broken since I own them!
For instance that's commit 9e02977bfad006af328add9434c8bffa40e053bb.
What happens when someone creates a commit and that happens to have 9e02977bfa-e-something?
I'm not even sure what github does here. Does it just refuse to select a commit or give you a list of candidates and you have to check the date (given that your link presumably was unambiguous when you wrote it)?
But the chance of a collision between a specific commit and any other commit goes up linearly, so for the Linux repository it's about 1 million / 1 trillion so 1 in a million.
The chance of a collision between a any two commits goes up quadratically (while it's small), and becomes ~50% when you reach the square-root of 1 trillion, i.e. 1 million. So there is a good chance the Linux repository already contains colliding 10 digit hashes, and if they don't exist yet, it'll likely happen during the next couple of years.
But if you pick an arbitrary commit to cite in a research paper, it's still 1 in a million that that particular commit has a collision. And this is for a repo which has to be in the 99.999th percentile of number of commits.
By now it's up to at least 7 (going by my clone of 5.12-rc1 that I had lying around), and it's becoming more likely.
Those collisions happened with objects other than commits (of which there are more), but that's by no means guaranteed.
You could probably cut the commit hash to 4 characters and still not have collisions for this application (ie. 1 / 65536 chance of collision @ 4 characters), or even fewer
The chances of a hash collision by chance with 10 characters are basically impossible (unless sha1 is broken further and it's a malicious repository).
The master branch of the linux kernel has a little bit over 1 million commits (measured with git log --format=oneline | wc -l).
With "git log --format=oneline | cut -b 1-4 | uniq -c | sort -n | tail -n30" you can verify that it has 25 prefix collisions with a 4 character prefix. With 5 it's still 4 collisions, with 6 there are none.
The master branch is irrelevant - github links to commits don't include the branch (and the way git works, commits don't really have "a branch"):
https://github.com/torvalds/linux/commit/ac632c504d0b881d7cf...
>With "git log --format=oneline | cut -b 1-4 | uniq -c | sort -n | tail -n30" you can verify that it has 25 prefix collisions with a 4 character prefix. With 5 it's still 4 collisions, with 6 there are none.
I'm sorry to say, your script doesn't work. You need to `sort` before `uniq` as that only counts adjacent duplicates.
The linux kernel as of 5.12-rc1 had 25 prefix collisions for the prefix "ffeb" alone (and 33 for "e3f2"). Here are the full shas:
ffeb03cfe2b49b73da7b325a31714003761fc6d5 ffebecd9d49542046c5ecbb410af01e016636e19 ffeb1e9e897b8d36b197275592d121c96d3bdb95 ffebbecaaa86f7cde4a6a813bed14f9d56e7c373 ffeb595d84811dde16a28b33d8a7cf26d51d51b3 ffebbaedc8616cffe648202e364dce6a045d65a2 ffebe74b7c95a41d2d0ac70a44d410e0efa37ad8 ffebc8c0344934db710afc76e3bfda11a823bb3d ffebb83b34f843aadcbd3e03c4e449da14d0870d ffeb6437f018f072a330ad5911036fd020b35ac3 ffeb883e5662e94b14948078e85812261277ad67 ffebf5f391dfa9da3e086abad3eef7d3e5300249 ffebfc364dcaa5dea1a589d42207834b028df789 ffeb13aab68e2d0082cbb147dc765beb092f83f4 ffebeb46dd34736c90ffbca1ccb0bef8f4827c44 ffeb40515971c9860ff671bb074689db15e18831 ffebad7948ee0e9c619ae6e87d99437d907fc7e3 ffeb501c6cba803eefc46b570feccffe61a6d883 ffeb33d20c6217bb8f0ab46d3f1396021c00c24f ffeb80fc30acbf6bd51cb47a1815f621a9d017dc ffeb414a59291d5891f09727beb793c109f19f08 ffebedb7ab3f7964a70a1771547b26af38a189d2 ffebabe0bf0de9ee500d4605d6acb71e1ee3b79f ffeb9ec72e18e16d0b0835d959cdf01650758638 ffeb874b2b893aea7d10b0b088e06a7b1ded2a3e
with 5 characters there are 8 for b91e1.
With 6 characters there are 5 for 120bda:
120bdafaece72056e48d97809c5abe172824a7f6 120bdac7376a36418eb1d55e0161dc0e660a45c3 120bdaa47cdd1ca37ce938c888bb08e33e6181a8 120bda35ff8514c937dac6d4e5c7dc6c01c699ac 120bda20c6f64b32e8bfbdd7b34feafaa5f5332e
and a total of 28606 colliding 6-digit prefixes.
There are 9 colliding 9-digit prefixes and no 10-digit prefix.
(found with `git rev-list --all --no-abbrev-commit | sort | cut -b 1-4 | uniq -c` and variations on that theme)
It's even more stupid considering that a 4 character hash has only 16*4 => 65536 different strings so there must be a lot of collisions in over a million different hashes.
As commit hashes ought to be globally unique—the same hash can appear in more than one repo but they will all have the same content and history—GitHub could offer a shorter form of the URL which omits the user and repo names, in exchange for a longer minimum hash length. This would require creating an index of all the commits reachable from public repos, unless they already track that.
> There's all info in the url that you even can use in the future if GitHub ever goes away.
GitHub could maintain a .git_io file (or whatever) in the repository with the short->long URL mapping, for instance. So, the info would survive GitHub.
* The file is available at the following URL: http://example.com <- good
* The file is available here <- bad [here = clickable link]
And then someone else can make a browser extension that automatically queries this look-up table when it sees the browser user has clicked a git.io link. It's like DNS, but different!
(Yeah the cleanest option would be for a trusted someone to take over hosting of the git.io service, but hey, welcome to the present...)