GitHub: Can no longer search code without being logged in
github.com
github.com
My guess is that the trade-off here is genuinely a question of if they spend 3-4x (I'm guessing, but I think this is likely a low-ball estimate) on their search infrastructure, v.s. having people angry at them for requiring login to use the feature.
I actually built my own code search tool for use with GitHub repos, but I've mostly stopped using that because the new GitHub code search is so useful by default: https://simonwillison.net/2020/Nov/28/datasette-ripgrep/
While a bad actor can certainly scrape/clone the same data... Given repos with thousands of files... Scanning by scraping vs searching vs cloning.... Search is certainly the cheapest for the attacker while the most expensive for GitHub
In fact, that's exactly how it was until in June GitHub started requiring login for searches even in just a single repo.
The end result is that it's no longer possible to participate in "GitHub hosted open source" without actually being a member of GitHub.
I'm not saying it's a deal that makes sense for everyone. But it's certainly a barter trade and is well within their prerogative to offer.
> That's a lot to give for "I want to browse a piece of code someone made for free".
You can still do this. All repositories are accessible without signing in.
This is just degrading in to conspiratorial FUD.
Or to put it another way: It's a surprise you need a Github account to use ("participate in") Github?
Since those are "fixed costs" per repo rather than per search, they'd now be much more expensive per search if they were only used by logged-out users.
Pure outside speculation of course.
In other words, GitHub is no longer suitable to host open source repositories because of all the closed "extras" it adds on top.
They're shady AF, but some features require an account not because they're moat water but for some other reason.
It's clear the "extinguish" phase of Microsoft's open source venture has started. There're plenty of very good alternatives, it's time to move on.
http://vger.kernel.org/vger-lists.html
Unless you count your ISP of course.
probably for the best.
Are these now also no longer "open source"?
No.
None of this is somehow a core requirement of "open source".
An account has always been required for active participation, e.g. committing directly to any repo, submitting a PR, opening an issue, etc.
Disclosure: I work at Google.
The other side is, that existing users will get pissed at github, because they can't even search anymore without logging in, and sometimes that's a pain (not their pc, public pc, incognito tab, the time needed to do the 2fa, etc.).
Github can still keep the cheap, fast basic search for users not logged in, but they didn't.
That's an assumption which github may not agree with.
1: You have to maintain two full independent search and indexing implementations
2: User confusion. Why are the search results different depending on if I'm logged in or logged out?
I mean, the person opening the issue escalated this into socioeconomic issues just because they don't know how to use grep... What else is there to say?
I suppose one could try to see the value in the argument and engage in good faith discussion but it's easier to flippantly dismiss people like this, I imagine.
There wasn't enough thought behind the issue, other than being outraged at their unsatisfied entitlement
Slightly more than your response, at any rate
They justify it with "engagement" and it kinda works to a point, since if you don't get pissed with the broken search and register an account, you're not considered a "user", and if you try to use the tool but don't want to register you're not "engaging" with it, I guess... But it makes the initial contact that much more miserable.
This makes no sense. The number of people who use GitHub code search but don’t already have a GitHub account is surely negligible
logged in user ad impression is 8 to 12 usd. anonymous user is .01cent per thousands.
Trying to think this through, my own best guess for what is going on is that there is some amount of traffic coming from bots/scripts and github would like to have all queries associated with an account so that they can block accounts? I'm not sure if this makes sense, but I think it's a reasonable possibility.
now under new ownership they closed search and "pray they don't alter the terms further"
ps: also everyone here guessing wrong. Microsoft is an advertising company. requiring you profile for a service is self explanatory.
Seems to be around 5pc revenue from a quick search.
Not that I want to defend Microsoft.
by putting login, they can throttle or block it entirely.
I suspect the change has more to do with security concerns (and having an audit trail) than performance.
I think the solution to hard coded credentials isn’t making search harder (security through obscurity == bad) but the other things GitHub is doing to detect and mitigate hard coded credentials.
Search is still possible using google and other methods, so any theoretical gains from forcing login are dangerous to count on as vulnerable projects are still vulnerable.
If anything, I think the solution to projects with hard coded credentials are to make the credentials easier to find and exploit so they fixed more quickly after being created. The most dangerous are ones they are hard to find so someone uses them for long periods of time without detection.
Some high-profile cases where credentials were leaked on public GitHub: Uber in 2014 and 2021 [1, 2] and Twitch in 2021.
> Search is still possible using google and other methods
Yes, you can search with Google or other sources. But the thing is, pretty much only GitHub has easy access to all the code present there, readily available to search within minutes of pushing.
You could try mirroring GitHub yourself, but you'd need enormous disk space and bandwidth, and would quickly hit rate limits. You also wouldn't be able to do fast full regex search like you can with GitHub's search, as you don't have their search infrastructure.
Hackers are aware of this and do make use of GitHub's search to identify possible leaked credentials. I recently experimented with uploading a dummy AWS credential pair to a public Git repo, and saw that numerous IPs started trying those credentials less than 5 minutes after I pushed.
> any theoretical gains from forcing login are dangerous to count on as vulnerable projects are still vulnerable
Indeed, requiring login is not going to _solve_ the problem. But it could lessen the impact of leaked credentials at large scale, by making it more difficult for automated systems to harvest them.
Perhaps more significantly, requiring login could give better audit trails in incidence response situations, as the logs would indicate which accounts were searching for secrets.
> I think the solution to projects with hard coded credentials are to make the credentials easier to find and exploit so they fixed more quickly after being created
Yes, this can help! There are several other companies that specialize in secret detection. I've also written Nosey Parker, a fast regex-based detection tool that has higher-precision rules than similar tools [4]. GitHub also has its own offering in Advanced Security to address this problem.
[1] https://www.reuters.com/article/uk-uber-tech-lyft-hacking-ex... [2] https://www.securonix.com/blog/securonix-threat-research-ube... [3] https://news.ycombinator.com/item?id=28770590 [4] https://github.com/praetorian-inc/noseyparker
They already did.
> . . . then I won’t like that.
There is no universe in the multiverse where Microsoft or one of its subsidiaries gives a single flying fuck. Something about those who forget history being doomed to repeat it . . . .
https://www.reddit.com/r/ChatGPT/comments/10j531u/i_made_cha...
And also, what is stopping the scrapers from creating a bunch of accounts to scrape with, and rotating them once they get banned?
> Give me repos with project.toml and a flake.nix at the root
The interface supports it, but searches come back empty. I end up leaving scripts running while I sleep which walk search results and narrow them by just checking to see if the files are there.
If you're out there Microsoft, please fix your search so I can stop being a bad citizen.
those accounts still need to be created, which is hard to automate
also, now GH can just ban more easily detectable accounts that scrape, vs anonymous visitors
If there is a 3rd explanation, I'd love to hear it.
They explicitly said that they "had" to require being logged in, so, yeah: https://github.com/orgs/community/discussions/77046#discussi... . Although they do assure us, however, that they're totally sorry for the inconvenience.
So are expected to believe that a hostile user who has zillion of ip addresses (to get around the rate limits) won't also be able to make a handful of accounts?
I understand perfectly why this has been done from a business standpoint as it invariably increases 'user engagement' by driving people to become literal users!
even searching a specific file name inside the repo has been locked behind the same way.
There are LOTS of reasons to host your open source project on GitHub
But if you look at any source forge repo right now, I think you'll see github as a long way to fall before it gets that bad.
No one has walled anything off in any way that is materially different than any other option in the space.
I would argue that Github has done a LOT of good in the space. Making good software, making it freely available. Keeping it reasonably open and accessible. Keeping it standards compliant where there are standards. Having API's for the rest. And in general, giving a huge amount of storage and compute away for free for open source projects.
A lot of these walled garden platforms have contributed a ton and github has made source code more easy to host, no doubt. At the same time, we need to ask them to do better and not allow them to concentrate power for when the inevitable enshittification begins.
I believe a long term effect will be a rise in the marketplace for proprietary data. Search engines, AI tools, anyone who needs the data will have to pay for the firehose or API access. (That, in turn, may cause antitrust issues if only the richest companies can afford access to that data.)
If something is truly proprietary, then the compensation should go to its owner, not to some random Git server operator or whatever. And the owner, not the server operator, should be setting the price and terms. If the server needs to make money, it can charge for the basic service itself.
On the other hand, if somebody created something to give it away, on a platform that at that time provided a channel effective for giving things away and advertised itself as suitable for giving things away, then the platform changing the rules midstream is morally piracy of that person's work.
Probably a very large fraction of the material on those platforms would never have been put there in the first place if the actual owners had expected random restrictions and arbitrary access charges.
If they want to radically rewrite the rules like that, then they need to not sell access to any preexisting content unless the original creator explicitly opts in. But we had Reddit, for instance, actively reinstating posts that had been mass-deleted by their authors specifically to prevent that kind of abuse.
It is time to move from centralized services to something more distributed and "micro transaction based". I know a lot of people in HN will dislike this butn All the ones you mentioned (Q&A forum, News link + commenting aggregator, Git Hosting + Approval Flow + Wiki + Ticketing system ) can and should be implemented in a completely distributed manner: (using Kademilla/Bittorrent technologies, IPFS and CryptoTokens for paying micro transactions).
Going 100% distributed is the only way these sort of things are going to stay "open" and free of corporate greed (we have seen time and time again, original founders may start the project with good intentions, but in the long term, the product gets sold and bean counters take the lead).
As for the actual important data (source code), you can simply clone the repo anonymously using HTTPS endpoints, and do your code search there (e.g. there are third-party websites like https://grep.app/ that can do this).
The issue with say Reddit/Twitter is that the source data (posts/comments/upvotes and tweets) is no longer available via APIs, making it impossible / hard to build third-party tools. With GitHub, the source data is still easily accessible, with search (the auxiliary layer) disabled for non-logged-in users. I think they are quite different situations and mixing them together is not useful.
I am not sure how much behavioral data GitHub can gather from logged in user, and how useful that is compared to the code that is there anyway. Maybe to figure out which parts of code are important? But that isn't really user-specific.
Many companies will black-hole you and force infinite captias despite solving them correctly to waste resources.
Virtually all of the traffic that was intercepted claimed to be modern Chrome or Safari or similar, which should be capable of "executing the latest bleeding edge javascript functions".
The primary reason why anyone gets shit from bot mitigation is IP reputation, this is far more important (and effective) than looking at browser characteristics.
But yes, there are multiple good explanations for why they would lock down the API.
What behavioral data can you glean from a code search like Github's? The context is very different than, for example, Google's, so is there really much useful data you can get here?
It's as simple as appending the repo URL, starting from github.com: e.g. sourcegraph.com/github.com/rust-lang/rust/ (you can try searching for unit on this repo)
I would not be surprised if they start restricting git clone next.
IIRC GitHub is already rate limited to slow down excessive cloning or API usage.
It's a perverse reality but it seems that in order to keep some ecosystems open one has to take actions which are resource-wasteful (though I would argue in the larger picture it will save resources)
A paid subscription could be crossing the line. Or maybe it would be fine as long as no profits go to the entity that provided you with the binary. Hard to tell.
I would, considering the amount of CI systems this would break. Say goodbye to large parts of NPM, go package management, Jenkins scripts, etc.
That would be pure chaos.
The underlying argument isn't even that good. Besides that repos themselves are _not_ gated behind a login and so an open repo is able to be accessed publicly, including cloning the repo, if the point is that the author wants his repos and code to be useful to the public, then if there is ever a need to interact beyond just downloading the code and searching, such as creating an issue or making a PR, that contributor would necessarily need to log in.
Cloning is usually fast but not with very large repositories, and you might not even have enough space on your current device to clone them.
By the way if a lot of people will end up cloning everytime instead of logging-in (after all logging in to GitHub is a pain with the 2fa) I don't think their systems will have less load.
And not everyone might want to have a GitHub account
Being on GitHub just makes it easier for them as they don't need to manually clone it from a different service.
(Full disclosure: I'm the Sourcegraph CTO)
I remember longing for the ability for github's search to rival git clone + grep, but I never expected a login wall to come with it. IIRC I expected the login wall to just be because the feature was in beta and would be removed when it became the primary search.
I can't think of any software I've ever used that I was this concerned about release schedule. Even if I was a user of your site, I might check it daily, but even that is doubtful.
Of course they can, but then they're gonna be chewing up their own disk space and bandwidth for anything after the initial hit. I think the real problem is that the bots hit the GitHub servers over, and over, and over again.
Why not either learn to use local code search tools and clone a repo, or just log in and search then? Are either of these things such a hurdle?
I sympathize that the person opening the post is frustrated by the chain of PEBKAC events that unfolded, but that's not really an issue with GitHub.
The issue opener is behaving like an entitled toddler by making his problems "societal" problems. Honestly, the lack of self awareness with some people is astounding.
https://sourcegraph.com/search has a really powerful code search can be a pretty powerful alternative and allows you to search codebases without logging in and across different code hosts.
Real leverage looks like all of us banding together to boycott Microsoft's most profitable income streams. Not GitHub. But cloud services like Azure, and Office, gaming, etc:
https://www.nasdaq.com/articles/these-2-revenue-streams-acco...
It's about making unilateral ensh*ttification decisions like this so expensive for parent companies that they become unthinkable. The board and investors sometimes need to be reminded how expensive these decisions really are.
We've all seen stock prices tank by half in one day since the Dot Bomb. A weak signal starts a big wave in this age of algorithmic trading.
A high-profile company publicly switching to AWS/GCP to avoid eating MS's sh@t sends a strong signal.
More importantly, after everything that MS has put us through over the decades, divesting is FUN! :-)
But that is repository search only, of course and not github-wide.
https://docs.github.com/en/codespaces/the-githubdev-web-base...
EDIT: No, I was wrong, you can search it without authentication.
First off, I have a GitHub account. But I'm often in contexts where signing into it is annoying, and intentionally so (it lets me push code!) I usually don't want it on my work machine, even though I search GitHub a whole lot from there. Some of it is of course related to my job directly, which you might plausibly make the argument for that my employer should somehow compensate them for a free service that indirectly makes them money, but a lot of it is literally my work computer being a second machine that I work on, with open source code running on it, and usually these searches are to contribute value back to the community–either because I want to file a bug against it later, or maybe even deciding whether it is worth my while to contribute code to it. And I do this all the time from other devices, too: I might have logged into GitHub on my phone once, but my iPad? My mom's computer? Being able to search GitHub is an excellent way for answer people's tech questions at any time, much like if Google asked you to sign in everywhere it would be a massive pain.
Second, and also quite important, is that I can't use GitHub links as a way of pointing people at code anymore. I frequently (check my history!) will post a comment like "yeah the Foo project uses bar API a bunch like this, [GitHub search link]". It's a very quick and very direct way of sharing this information. Of course I have no idea if the people on the other side are logged in or not, which means that if you put them behind a login wall I will slowly stop sharing these, because people will complain that they can't see what I sent them. It's the same way I hesitate to send people links to Twitter/X these days, because whether the content will be accessible is a coin flip. And I'm definitely not going to ask people to clone the repo (on what, their phones?) to see what I'm seeing, so I might as well link to Sourcegraph instead.
Original HN Discussion: https://news.ycombinator.com/item?id=22396824
Sourcegraph works too: https://sourcegraph.com/search
GitHub is still giving away resources for free, they've done it forever because it was their business model: attract as many people as possible through open source repositories, a fraction of which will then use their non-free services (and more recently, also to train AI models).
Now that they're so dominant (and owned by Microsoft) they can afford to worsen the services, but people have every right to be pissed about it and push developers to move their repositories.
What seems to be new new is that you can't search even a single repo.
You can just clone it and use your local tools.
If the code is hosted or mirrored elsewhere, and elsewhere uses GitWeb, you can use GitWeb's grep search.
Quite possibly Microsoft told it's subsidiary GitHub that they needed to spend less money and this is how they save resources while impacting the users they care about the least.
What I hope never happens is the actual enshittification of the site for logged in users.
every usa based project had to quit shipping any cryptography code. but who cares about history. who cares you can only participate in some open source project today if you have an account with usa companies? nobody cares. screw the Iranian engineering student. and hope the usa doesn't outlaw encryption again.
The only automated searches I could envisage to MAYBE be happening are those looking for vulnerabilities, and there are probably only a few dozen actors doing it. State actors seem actually more likely to use clones of everything.
Now that the entire programming world has just about hard coded GitHub into the very center of everything, it's time to start enshittifying.
I guess code search requires a lot of processing power so they'd rather know who's doing it in case they need to be throttled.
https://github.blog/2023-02-06-the-technology-behind-githubs...
What an asinine thing to do. My GitHub usage has dropped exactly to 0% for this very reason. I know I am not alone either since their org forums is filled with complaints about this.
Not going to use any strong words but I want to. Disgraceful UX choice.
> We are thinking about way to make this work in the new code search system, but don't have a date for when it will be possible.
If you wished to build a ML bot that taught switches/loops; instead of cloning every repo all you'd need to do is search for switches/loops within X lang.
Downloading a repo and than neither knowing the repo has the syntax you wish to learn is still going to be more resourceful than having a webpage thrown at you with all search results.
If people can search anonymously it gets a lot harder to datamine
See e.g. Troy Hunt doing this for HaveIBeenPwned: https://www.troyhunt.com/fighting-api-bots-with-cloudflares-...
I have had my own issues with Cloudflare Turnstile (https://news.ycombinator.com/item?id=38412057), but in principle, I don't see why can't this be done