Sourcehut will blacklist the Go module mirror
sourcehut.org
sourcehut.org
Go 1.19 added "go mod download -reuse", which lets it be told about the previous download result including the Git commit refs involved and their hashes. If the relevant parts of the server's advertised ref list is unchanged since the previous download, then the refresh will do nothing more than the ref list, which is very cheap.
The proxy.golang.org service has not yet been updated to use -reuse, but it is on our list of planned work for this year.
On the one hand Sourcehut claims this is a big problem for them, but on the other hand Sourcehut also has told us they don't want us to put in a special case to disable background refreshes (see the comment thread elsewhere on this page [1]).
The offer to disable background refreshes until a more complete fix can be deployed still stands, both to Sourcehut and to anyone else who is bothered by the current load. Feel free to post an issue at https://go.dev/issue/new or email me at rsc@golang.org if you would like to opt your server out of background refreshes.
Maybe it should be an opt-in list where the big providers (such as github) can be hit by an army of bots and everyone else is safe by default.
It smells like Go is on its way out.
I think Drew is right in that he shouldn't take a personalized Sourcehut-only exception because this doesn't address the core issue for any new small providers that pop up.
Between this and the response in the original thread that said, "For boring technical reasons, it would be a fair bit of extra work for us to read robots.txt," it gives the impression that the Go team doesn't care. Sometimes what we _need_ to do to be good netizens is a fair bit of boring technical work but it's essential.
Exactly. We already saw how this ended with Google vs. people running mail servers.
I cannot file an issue; as the article explains I was banned from the Go community without explanation or recourse; and the workaround is not satisfying for reasons I outlined in other HN comments and on GitHub. However, I would appreciate receiving a follow-up via email from someone knowledgeable on the matter, and so long as there is an open line of communication I can be much more patient. These things are easily solved when they're treated with mutual respect and collaboration between engineering teams, which has not been my experience so far. That said, I am looking forward to finally putting this issue behind us.
The question is who will swallow their pride first: Sourcehut or Google.
Hi ddevault, FWIW, in May 2022 on that #44577 issue [0] you had opened, it looks like someone on the core Go team commented there [1] recommending that you email the golang-dev mailing list or email them directly.
Separately, it looks like in July 2022, in one of the issues tracking the new friendlier -reuse flag, there was a mention [2] of the #44577 issue you had opened. In the normal course, that would have triggered an automatic update on your #44577 issue... but I suspect because that #44577 issue had been locked by one of the community gardeners as "too heated", that automatic update didn't happen. (Edit: It looks like it was locked due to a series of rapid comments from people unrelated to Sourcehut, including about “scummy behavior”).
Of course, communication on large / sprawling open source projects is never quite perfect, but that's a little extra color...
[0] https://github.com/golang/go/issues/44577
[1] https://github.com/golang/go/issues/44577#issuecomment-11378...
[2] https://github.com/golang/go/issues/53644#issuecomment-11751...
And given that they banned him for no reason, he is perfectly in the right to tell them that they should email him instead.
We have not been reading the same tickets and articles it seems
So yes. The issue they banned him from. Because reality's more complicated than flippant one liners.
The background refresh is meant to prefetch for that situation, to avoid putting that time on an actual user request. It's not perfect but it's far less disruptive than having to set GOPRIVATE.
...but it sounds like disabling background refreshes would have strictly better end-user performance than what the Sourcehut team had been planning as described in their blog post today (GOPRIVATE and whatnot)?
Why was the author of the post banned without notice from the Go issue tracker, removing what is apparently the only way to get on this list aside from emailing you directly?
Do you, personally, find any of this remotely acceptable?
...but as a place that could hold a rate limit recommendation it would be nice since it appears that the Git protocol doesn't really have the equivalent of a Cache-Control header.
But yes, it may be the best available solution in this case, even if I would argue that it isn't really it's main purpose.
What would be good is respecting `Cache-Control`, which unfortunately many RSS clients don't, and just pick a schedule and poll on it.
Eg: https://www.robotstxt.org/faq/kinds.html >"What's New" monitoring
A crawler has a list of resources it periodically checks to see if it changed, and if it did, indexes it for user requests.
Contrary to this totally-not-a-crawler, with its own database of existing resources, that periodically checks if anything changed, and if it did, caches content and builds chescksums.
(At least from an enviromental perspective.)
Easy to verify, get a report [1]: go install paepcke.de/fsdd/cmd/fsdd@latest && cd $GOMODCACHE && fsdd .
[1] Warning: Apple user with fixed restricted & expensive nvme space maybe very upset. Easy fixable via fsdd . --hard-link.
For some reason he did not do this and instead chose an option that causes breakage.
[1] https://github.com/golang/go/issues/44577#issuecomment-11378...
I am very concerned that "own assessment" of what is a DoS means that source code is expected to be hosted only on large platform or by large corporation which is another way to say that "the little guys don't matter".
Self hosting of source code should be an option and the proxy should be there to reduce the traffic load, not amplify or artificially increase that load despite the "level of DoS".
One thing Drew is asking for is to respect robots.txt to allow the operator to determine what a reasonable level is for that operator and not apply a github bias to it.
That’s (upper bound) 4Gib times 2500 per hour. That’s not nothing.
Also, how do you opt-out? Imagine a random developer in a startup, running a Gitlab instance and then pushing a Go module there and only to be left with inexplicable traffic pattern(and bill). I have no skin in the game but this default _does not_ sound reasonable to me, whichever way you slice it.
That's per Go repository. That's a non-trivial amount of egress data and probably adds up to thousands of dollars a month.
We must absolutely reduce our resource usage.
In this context, that seems like a lot. Of course the module mirror can't know about this context, but there are certainly a lot of scenarios where this is comparatively a lot of bandwidth. Not everyone is running beefy servers.
Seems like an exceedingly poor and unreasonable default, and it doesn't take much imagination to see how this could be improved fairly easily (e.g. scale to number of actual go gets would already be an improvement).
"All this would go away if other people would just do what they're told" is a pretty dystopian policy, but it seems to be a popular choice in Google.
Anyway, things get a bit kafka-esque when you realise that there is another company doing this WiFi thing and to opt out from that one you need a different SSID suffix. Since you can't have both, you end up with at least one company data mining you.
Why GDPR has not put a stop to this is baffling
I could be wrong but in comments to https://github.com/golang/go/issues/44577 issue there at least 4 hosters that forced to manually disable background refresh because of this exactly issue?
Holy cow Google! Wouldn't it behoove us to check if any changes occurred before downloading an entire repo?
A checkout with this can literally clone nothing but hash
git clone --depth=1 --filter=tree:0 --no-checkout https://xxxx/repo.git
cd repo
git logOf course, I have seen alleged examples of Go modules using tags like branches and force pushing them regularly, but that kind of horror sends shivers down my back, at least, and I don't understand why you'd build an ecosystem supporting that sort of nonsense and which needs to be this paranoid and do full repository clones just for caching tag contents. If anything: lock it down more by requiring tag signatures and throwing errors if a signed tag ever changes. So much of what I read about the Go module ecosystem sounds to me like they want supply chain failures.
I don't understand the Go ecosystem.
Yeah, on third-party code hosting platforms :). And maaaaaybe in some short-lived cache somewhere. I mean, why spend on storage and complicate your life with state management, when you can keep re-requesting the same thing from the third-party source?
Joking, of course, but only a bit. There is some indication Google's proxy actually stores the clones. It just seems to mindlessly, unconditionally refresh them. Kind of like companies whose CI redownload half of NPM on every build, and cry loudly when Github goes down for a few hours - except at Google scale.
But the whole thing is frankly a little rude.
Could you tell me more about what tool you used to host the repos? And I assume you noticed the traffic in your web logs?
That's some really entitled thinking on the part of the Go team at Google, and it's sad to see people stanning for them.
The entire point of the sumdb (go.sum), is to prevent the need for such a relationship. If Google (or any proxy you use) tries to return questionable packages, it will be detected by that system.
No, the error message you get is neutral about which side might be wrong - it says "verifying module: checksum mismatch" and "This download does NOT match the one reported by the checksum server." (I've seen it a lot because it also appears when module authors rebase, which a small but surprisingly high number do...)
That is exactly the detection of a poisoned module in the ecosystem. It would break builds, issues would get filed, and a new version would be released (and the malicious party may not be so lucky this time since it’s trust on anyone’s first use).
But I guess it's also fairly easy to test it: just serve a slightly different version to the google's go mirror (by the user agent), and see how long until somebody complains to you about it.
I think every company I know of with private Go modules (6-8 or so?) is running a module proxy, which will detect this. The several times we've detected this it's always been within 2-3 days of the upstream mistake. When I go to report a bug we're not always the first either.
I'd be interested to understand why that solution hasn't been implemented yet.
"it would be a fair bit of extra work for us to read robots.txt", so clearly tracking commit hashes would be even more work.
> I was banned from the Go issue tracker without explanation, and was unable to continue discussing the problem with Google.
is completely asinine. But it's also par for the course when it comes to interacting with Google. When is anyone going to hold them to account for their terrible customer service and community interaction?
Regardless, I don't really want to re-litigate it here. The main issue is that Google has been DoS'ing SourceHut's servers for two years, and I think we can all agree that there is no conduct violation for which DoS'ing your servers is a valid recourse.
Spoiler: Nothing.
edit: Possibly not as much nothing, see the replies to [0]. But the GH search kinda sucks for this.
> I was banned from the Go issue tracker for mysterious reasons [ In violation of Go’s own Code of Conduct, by the way, which requires that participants are notified moderator actions against them and given the opportunity to appeal. I happen to be well versed in Go’s CoC given that I was banned once before without notice — a ban which was later overturned on the grounds that the moderator was wrong in the first place. Great community, guys. ]
When a story has two sides and one party chooses to keep silent when accused, I tend to favor the accuser.
Can someone explain what this means? After all reading a small text file is plainly easy to do so its meaning must be something other than obvious.
Judging just from the linked post, the issue on which this was discussed, and this thread, it's feeling a lot like this proxy was some kind of proof-of-concept that escaped its cage and got elevated to production.
The other problem is that it's Google so their perception of "not much traffic" is "biblical floods" to other people.
You can run your own proxy service if you want. There's a large benefit to the go community as a whole for there to be a shared default proxy service.
This is a consequence of what IMO is another bad decision by the Go team: having packages be obtained from (and identified by) random github repos, instead of one or more central repositories like Maven Central for Java or crates.io for Rust. The proxy ends up being nothing more than an ad-hoc, informally-specified, bug-ridden re-implementation of half of a central repository, to paraphrase Greenspun's tenth rule.
E.g.: source-based packages on distributions where users may not be Go programmers, or even non-programmers, will compile and install Go software where some nested library dependency is on sr.ht. These packages will now fail and, sadly, this is going to cause widespread disruption. I think it'd be worse if those failures only happened occasionally, and not reliably repeatably.
I am wondering, because even if google's traffic is unreasonable, it might still be less than without the proxy.
CI configurations are notoriously inefficient with dependency fetching, so I would not be surprised if the actual client traffic is massive and might overwhelm sourcehut if all migrate to direct fetches.
Would it be simple to solve this with an additional layer of the same proxy? Currently, end users request a package from proxy.golang.org as per the default value of the GO_MOD_PROXY env var. Google runs many of these to handle the traffic, let’s say 1000. They all maintain a mirror of all packages that have been recently requested (note that I expect there’s more nuance here around shared lists of required packages, etc, but the structure should hold true)
The result is that every one of the 1000 proxy instances requests the source data from git.sr.ht every day.
Google could set GO_MOD_PROXY on the existing instances to internalproxy.golang.org. They could then run 100, or maybe 10 of these internal instances. This would drop the traffic to hit.sr.ht by one or two orders of magnitude.
I suspect it would require minimal if any change to application code. This might be accomplished entirely within the remit of a sysadmin (SRE?).
Any holes in my reasoning?
Isn't that exactly what "git fetch" is doing already ?
> For boring technical reasons, it would be a fair bit of extra work for us to read robots.txt […]
This is coming from one of the biggest, richest, most well-staffed companies on the planet. It’s too much work for them to read a robots.txt file like the rest of the world (and plenty of one-man teams) do before hammering a server with terabytes of requests.
If this is too much for them then no wonder they won’t implement smarter logic like differential data downloads or traffic synchronization among peer nodes.
Wouldn't the effect on sourcehut users be identical?
I've not seen any explanation about why the solution offered by the Go team was unacceptable. Its weird that that is completely left out of the blog post here.
Edit: Uh okay, if it's not user traffic then why wasn't the "don't background refresh" not an option?
Doing shallow clones, which are significantly cheaper.
Google is DDoSing them, by their service design. Why a full git clone, why not shallow? Why do they need to do a full git clone of a repository up to a hundred times an hour. It doesn't need that frequency of a refresh.
The likely answer is that the shared state to handle this isn't a trivial addition, it's a lot simpler to just build nodes that only maintain their own state. Instead of doing it on one node and sharing that state across the service, just have every node or small cluster of nodes do its own thing. You don't need to build shared state to run the service, so why bother? That's just needless complexity after all, and all you're costing is bandwidth, right?
That's barely okay laziness when you're interacting with your own stuff and have your own responsibility for scaling and consequences. Google notoriously doesn't let engineers know the cost of what they run, because engineers will over-optimise on the wrong things, but that also teaches them not to pay attention to things like the costs they inflict on other people.
It's unacceptable to act in this kind of fashion when you're accessing third parties. You have a responsibility as a consumer to consume in a sensible and considered fashion. Avoiding this means you're just not costing yourself money through your laziness, you're costing other people who don't have stupid deep pockets like Google.
This is just another way in which operating at big-tech-money scales blinds you to basic good practice (I say this as someone who has spent over a decade now working for big tech companies...)
Huh? I left a few months ago but there was a widely used and well known page for converting between various costs (compute, memory, engineer time, etc).
Thats a DDoS.
You know, like be good neighbors and respectful of other people's resources, maybe read robots.txt and not make excuses for why you are writing shitty stateless workers that spam the rest of the community.
I could build a system that did this in a week without any support from Google using existing open source tech. It's mind boggling that Google isn't honoring robots.txt, is requesting full clones, and isn't maintaining a local cache.
If you disregard the question who pays for a moment and only look at what makes sense for the "bits", the stateless architecture seems not so bad. Just a pity that in reality somebody else has to foot the bill.
I was not talking about how the nodes store data, but about a central cache. Purely architecture wise, it doesn't make sense to introduce a central storage that just mirrors Sourcehut (and all other Git repositories). Sourcehut is already that central storage. You would just create a detour.
It's also not an easy problem. If the cache nodes try to synchronize writes to the central cache, you are effectively linearizing the problem. Then you might as well just have the one central cache access Sourcehut etc. directly. But then of course you lose total throughput.
I guess the technically "correct" solution would be to put the cache right in front of the Sourcehut server.
If Google is going to blindly hammer something because they must have their Google scale shared nothing architecture pointed at some unfortunate website, then they should deploy a gitea instance that mirrors sr.ht to Google Cloud Storage, and hammer that.
It's unethical to foist the egress costs onto sr.ht when the solution is so so simple.
Some intern could get this going on GCP in their 20% time and then some manager could hook the billing up to the internal account.
It says in the post they'll check the UserAgent for the Go proxy string and return a 429 code.
I'm ok with saying "google should do better!" But the compromise solution from the Go team seems reasonable to solve the immediate issue in a way that doesn't harm end users. The author should at least address why they have chosen the more extreme solution.
EDIT: Moreover, sr.ht doing a workaround only for sr.ht, and lubar doing a workaround for lubar, etc... is not what Free Software is about. The point is that we're supposed to act as a community, for the betterment of the collective. Individualism is not a solution.
Furthermore, we try to look past the tip of our own nose when it comes to these kinds of problems. We often reject solutions which are offered to SourceHut and SourceHut alone. This isn't the first time this principle has run into problems with the Go team; to this day pkg.go.dev does not work properly with SourceHut instances hosted elsewhere than git.sr.ht, or even GitLab instances like salsa.debian.org, because they hard-code the list of domains rather than looking for better solutions -- even though they were advised of several.
The proxy has caused problems for many service providers, and agreeing to have SourceHut removed from the refresh would not solve the problem for anyone else, and thus would not solve the problem. Some of these providers have been able to get in touch with the Go team and received this offer, but the process is not easily discovered and is poorly defined, and, again, comes with these implied service considerations. In the spirit of the Debian free software guidelines, we don't accept these kinds of solutions:
> The rights attached to the program must not depend on the program's being part of a Debian system. If the program is extracted from Debian and used or distributed without Debian but otherwise within the terms of the program's license, all parties to whom the program is redistributed should have the same rights as those that are granted in conjunction with the Debian system.
Yes, being excluded from the refresh would reduce the traffic to our servers, likely with less impact for users. But it is clearly the wrong solution and we don't like wrong solutions. You would not be wrong to characterize this as somewhat ideologically motivated, but I think we've been very reasonable and the Go team has not -- at this point our move is, in my view, justified.
So, obviously he's supposed to know the even more obscure and annoying method of opting out.
https://groups.google.com/g/golang-nuts/c/6dKNSN0M_kg/m/EUzc...
Now a bit of personal history. The Go project was started, by Rob, Robert, and Ken, as a bottom-up project. I joined the project some 9 months later, on my own initiative, against my manager's preference. There was no mandate or suggestion from Google management or executives that Google should develop a programming language. For many years, including well after the open source release, I doubt any Google executives had more than a vague awareness of the existence of Go (I recall a time when Google's SVP of Engineering saw some of us in the cafeteria and congratulated us on a release; this was surprising since we hadn't released anything recently, and it soon came up that he thought we were working on the Dart language, not the Go language.)
Go is not only a "side project" at Google, but one of its most trivial side projects.
Yes, I do imagine that people who are really into Go are more likely than average to join or start Go shops, and then pick GCP over competitors because they have to start with something, and being Go people, Google stuff comes first to mind.
Lots of companies across lots of industries spend a lot of money to achieve more-less this fuzzy, delayed-action effect.
How many such people do you imagine there are? I'm active in the Go community, and I've been a cloud developer for the better part of a decade. It's never occurred to me to pick GCP over AWS because Google develops Go, nor have I ever heard anyone else espouse this temptation. I certainly can't imagine there are so many people out there for whom this is true that it recoups the cost that Google incurs developing Go.
Rather, I'm nearly certain that Google's value proposition RE Go is that developing and operating Go applications is marginally lower cost than for other languages, but that at Google's scale that "marginally lower cost" still dwarfs the cost of Google's sponsorship of the Go language.
AdWords is mainly written in Go. YoutTube is mainly written in Go. Just because they have strategic reasons for not directly monetizing Go doesn't make it a side project more than any other internal tooling.
It's core to their ability to pull in revenue now. If they were somehow immediately deprived access to Go, the company would go under. That's how you know it's not a side project.
Can you source these claims? Last I checked, YouTube was primarily written in Python, and I doubt that's changed dramatically in the intervening years given the size of YouTube. I assume there's some similar thing going on for AdWords.
> Just because they have strategic reasons for not directly monetizing Go doesn't make it a side project more than any other internal tooling.
Agreed, but all internal tooling is a side project pretty much by definition.
> It's core to their ability to pull in revenue now.
No, it's just the thing that they implemented some of their systems in. I'm a big Go fan, but they could absolutely ship software in other languages for a marginal increased operational overhead.
> If they were somehow immediately deprived access to Go, the company would go under. That's how you know it's not a side project.
I don't know what it means to be "deprived access to Go", but this is a pretty absurd definition of "side project" since it applies to just about everything Google does and a good chunk of the software Google depends on whether first party or third party (Google depends much more strongly on the Linux Kernel; that doesn't mean contributing to the Linux Kernel is Google's primary business activity). It seems you have a bizarre definition of "side project" which hinges on whether or not a business can back out of a given technology on a literal moment's notice irrespective of how likely it is that said technology becomes unavailable on that sort of timeline, and that these unusual semantics are at the root of our disagreement.
I love Go but not Google's stewardship of it. The tracking proxy, Russ' takeover / squash of the package management work, the weird silence / stonewalling on other community issues...
Drew has a valid complaint. I hate to hear he was banned from the issue tracker but that sounds about right.
As a sibling said - GOPRIVATE is probably a good solution without throwing the baby out with the bathwater.
Bourgon was frequently helpful and great, but also frequently rude, condensing, dismissive, and generally just unpleasant. I've seen this countless of times first-hand on Slack, Reddit, and Lobsters. I specifically stopped interacting with him long before he was banned. Whether he's a great programmer/contributor not isn't really important here.
I kind of hate how it's brought up here because I think he's not a bad bloke at all and I know the entire situation caused great personal hurt to him. But that doesn't change that he was the kind of "brilliant jerk" that would chase people out of the community with his behaviour, and that he was unreceptive to criticism of it (often getting pretty defensive/aggressive). No one liked how all of this turned out, but it did make the Go community a better place. Being helpful yesterday doesn't cancel out being a jerk today.
Same with Drew: he posted legitimate helpful issues. And he also ranted about how people were all a bunch of morons. I don't blame anyone for getting tired of that.
Do you have a source?
I don't have a full list of all posts at hand (some of which may be removed), but I've seen some other similar stuff as well; it's not an isolated incident. I was reading through the previous thread on this issue (goproxy sending loads of requests) and this one was posted as an example there.
ddevault on Feb 15, 2019
"EFAIL" is an alarmist puff piece written by morons to slander PGP and inflate their egos. The standards don't need to change to fix the problems it mentions. The proposals help... marginally. The problem is not and was never with OpenPGP, it's with poorly written email clients (e.g. all email clients).The linked comment was indeed out of line, and perhaps you feel justified in thinking that it should be sufficient grounds for a permanent expulsion from the community. I won't argue with that, fair enough. However, I don't think it's reasonable to use it as grounds to suggest that anyone should have their servers DoSed by Google with no recourse, and I think blocking Google is a reasonable move given two years of inaction from the Go team to resolve the issue.
2. This specific offer is not satisfactory: https://news.ycombinator.com/item?id=34313802
In the interest of not feeding the trolls, I think I can safely stop engaging with you on this thread. Or maybe on any thread -- you and I never seem to have a productive conversation on this website.
HN would be so very much more pleasant with ignore-lists.
Of course not; this entire thread isn't necessarily hugely on-topic here, but it got brought up, so ... well ... here we are. And in fairness, you did bring up your ban in the posted article.
> The linked comment was indeed out of line, and perhaps you feel justified in thinking that it should be sufficient grounds for a permanent expulsion from the community. I won't argue with that, fair enough.
No, I don't think anyone should be banned for a singular comment, no matter how egregious. Everyone deserves second chances, and third ones, even fourth ones maybe. There's some decent data from Stack Overflow that shows that after a ban many people keep posting and many don't get a second ban (i.e. their behaviour improves).
> I have good reason to believe that this incident is unrelated to the reason I am presently banned.
I think the thing is that it's part of a pattern. Usually the "final straw" isn't the worst incident, or even that bad of an incident in itself. Incidents like this aren't isolated and previous behaviour does tend to factor in: "oh, that's the same guy who called us a bunch of morons last year".
Wait, did some folks in the Go community write that EFAIL site that was referenced as a reason to drop OpenPGP? If so, that changes the context of the post a bit, but I didn't see anything indicating that was the case in the linked thread.
I don't see why not. Personalities fall on a broad spectrum. Still seems strange to me that the recent broad pushes for more inclusiveness, including neuro-atypicality, does not cover people that inconvenience you personally.
> that would chase people out of the community with his behaviour [...] but it did make the Go community a better place
I've noticed that claims like this are never backed by any evidence of this improvement, or evidence of people who have actually been chased away by rudeness. It no doubt causes great relief in the minds of those who dislike the exiled person, but it's always justified with a broader claim that "it's for the greater good".
I understand comments can be non-constructive, and that some people are more prone to it, but total exile is a big hammer that should be used more judiciously IMO.
I am one of them. I've seen other people claim the same. I did not keep a list, nor did I keep a list of all of his posts that I found egregious, and I don't really feel like spending a lot of time crawling through all posts to find them, so I guess this is all I have.
It's hard to get "hard evidence" for these kind of things in the first place. Most people just disengage and don't come back. The best I know of is "Assholes are Ruining Your Project"[1] from a few years back. It would be interesting to check similar numbers for Go and other projects. I'm not sure if it's easy to get these kind of numbers from e.g. Slack or Reddit though.
[1]: https://www.slideshare.net/dberkholz/assholes-are-killing-yo... / https://www.youtube.com/watch?v=-ZSli7QW4rg
> total exile is a big hammer that should be used more judiciously.
It wasn't the first time he was banned, but I'm not privy to the exact details on this. Was total "total exile" proportional? I don't know: obviously I didn't see everything. I just wanted to say he didn't "just" get banned over a minor thing, but after many years of problematic behaviour that had been raised plenty of times.
This is the kind of thing I'm asking about. Lots of numbers are trotted out but where's the actual data? Where's the methodology?
The blurb says, "This talk will teach you, using quantified data and academic research from the social sciences, about the dramatic impact assholes are having on your organization today and how you can begin to repair it."
Social science research has a dramatically poor replication rate, so on that basis alone I'm skeptical of the numbers even if he did interpret them correctly.
That said, I agree asshole behaviour has to be reigned in, but exile is pretty dramatic if you really think about it. It's super easy and I think that's why people do it, but that doesn't make it good option.
I'm saying that's still not a reasonable measure even in that case. Why not an exponential backoff, where the first measure is that they only get one post a day. If they want to be heard they have to be more careful in how they word things and they have more time to think about how it might be received. If they transgress again, then it's upped to every three days, then once a week, then once every other week, and so on. A total ban is the limit of this more nuanced process.
No doubt this feature doesn't exist, so I'm suggesting something like this should be added because I'm not at all a fan of bans. Even this is a stopgap measure used to manage assholes because we don't yet understand what's at the root of asshole behaviour.
Edit: to clarify, I mean the backoff/retry strategy is still not ideal, but an easy first attempt at trying to reframe this as a problem we can maybe address using programming abstractions to inhibit rather than facilitate communication. Most software is focused on reducing barriers to communication, which is why banning is the only recourse, but in cases like this you obviously want to raise barriers to communication in controlled ways so you don't have use the ban hammer.
It's not a perfect science, but that doesn't mean "do nothing" is the best option, or that we can't just use common sense for that matter. If someone joins a community space and their first interaction is being insulted then the chance that they will come back is lower than if they're not insulted. I don't think you need a whole lot of rigorous science to accept this basic point, just as we don't need a whole lot of rigorous science to accept that dogs can feel pain, have an emotional life, have different personalities, etc.
> That said, I agree asshole behaviour has to be reigned in, but exile is pretty dramatic if you really think about it. It's super easy and I think that's why people do it, but that doesn't make it good option.
It sure is dramatic! Like I said, I don't really have the full story on this, so it's very hard for me to judge if it's proportional. I don't think the decision was made lightly as everyone involved realized it's not J. Random Gopher but a fairly well-known person within the community.
Related story: in a community (unrelated to Go) I once sent a message to someone asking them not to insult people; pretty basic unambiguous "you can't call people idiots here" kind of stuff. They were also very helpful in other cases and I knew they were going to be sensitive about it, so I sent the kindest kid-gloves message I could come up with; no threats of any actions, just "hey, can you not do this here?" They just replied with "no, I will not change, fuck off". So ... I (temporarily) banned them. What else was I supposed to do at this point? Let them continue anyway even though it was clearly inappropriate? Anyone looking on might think "gosh, did you really have to ban them for those remarks? It wasn't that bad?" Not unreasonable, but ... they also weren't aware of the conversation I had with them, and their reply. No one made any remarks about it, but if they did, I wouldn't have commented on it because it's still a private conversation.
This is the kind of stuff we may be unaware of. In my first message I mentioned "unreceptive to criticism of it (often getting pretty defensive/aggressive)" for a reason. I don't know what happened behind the scenes, but from what I've seen in public cases where people commented on his behaviour I expect things didn't go swimmingly. It's one thing to screw up at times and at least acknowledge you screwed up, but it's quite another thing to be consistently dismissive about any concerns and outright reject the idea there is anything wrong with your behaviour. I expect that this attitude played a large factor in the decision.
I'm reasonably sure there had been at least Slack bans before though; this wasn't the first ban (I thought I mentioned this before, but looks like I forgot).
As a former Rust moderator, this, so much. So many people don't see this part, where you reach out to folks and spend long grueling hours trying to get them to correct their behavior, precisely because no non-psychopath wants to drop the ban hammer on anyone. (Unless it's for obvious spammers and drive-by trolls.)
And the people saying "well I'm not suggesting do nothing, but just use better tools." Well, yeah, great, let's use better tools. Who's going to get GitHub to implement them? Or whatever other platform you're using? Some platforms have better support for this kind of tooling than others, but GitHub's is (last time I checked) pretty bad and coarse. It is slowly getting better over time. It used to be virtually non-existent.
But in the mean time, the people actually in the trenches doing the hard work of moderation have to do something. If the platform doesn't have this sort of idealistic tooling that's easy to navel gaze about on HN, then they have to do the best with what they have.
Well it was used judiciously, considering there are not more than couple of people banned in Go spaces in a decade.
https://i.kym-cdn.com/photos/images/newsfeed/002/212/873/d5f...
In my (limited) experience, naming projects you left due to rudeness or bad behaviour tends to lead to that bad behaviour noticing your message and pestering you with questions about exactly why you left, and arguing that you are being unreasonable -- which is why I'm not naming those projects.
> Whether he's a great programmer/contributor not isn't really important here.
I'd make a distinction here between having a reputation as a great contributor and having something important and correct to say in a given exchange. No, a community shouldn't put up with a person with a great reputation (or elevated title or higher pay grade) if they are unpleasant and wrong. But if they're right and they're a little impatient or impulsive, the community has more to gain from listening and simply pointing out they don't need to be impatient and impulsive. Let him build up a reputation for being a jerk rather than just ban him.
It's still unproductive to say that in a community support space.
Who's the bigger moron, the moron or the guy holding a public grudge for over a year about how unfair it is they're not letting him in the channel full of morons?
Why not? Why shouldn't we offer more leeway to more valuable contributors?
The "classic" example of this is Ulrich Drepper, who maintained GNU libc for many years. Everyone agrees he's a great programmer. He's a better programmer than I am. But he was also ... difficult. More difficult than anyone else I've seen in a mainstream widely-used project. Many people didn't contribute purely because they just didn't want to deal with Drepper. Debian found it necessarily to fork GNU libc because of Drepper.
So even if we adopt a purely utilitarian attitude on this (and I don't think we should in the first place), I think it's still a bad idea to grant some people a license to be a jerk. In many cases you're not going to come out with better contributions and code.
I have no numbers to be able to confirm or deny this, but...
>Debian found it necessarily to fork GNU libc because of Drepper.
... Debian created eglibc because they needed glibc to support their use case and Drepper didn't. Even if Drepper had been the nicest person in the world, if he rejected patches to run glibc on non-x86 then forking was unavoidable.
> In many cases you're not going to come out with better contributions and code.
I think that you would most of the time.
Doubtful. Knowing that the Go team has a habit of ousting contributors because some feefees got hurt ensures that I'll never even consider trying to contribute.
And Drew talks "harshly". Says things like "shitty usage of the GPG command line tool by applications" where, "shitty" is considered by some crossing a line https://github.com/golang/go/issues/30141#issuecomment-46427...
Both are abrasive, but in my mind, this is a cultural issue.
I grew up in a judgmental, holier-than-thou religion that I rejected at an early age and then was ex-communicated from on my 18th birthday. I've lived this pattern of judgement and ostracisation. It's dangerous, closed minded, and wrong.
One of the reasons I work in software is because I'm a "strange" person. I say things people don't understand, my value system is radically different from my peers and "normies". I am a very different type of person. Software was a safe place for weird people, including the productive, but sometimes abrasive. It's a shame to make software not a place for a wide range of diverse people. In my difference, I don't want my cultural differences to be targeted by witch hunts, those looking for blood to prop up their own "moral superiority".
And I personally have no cultural issue with people who say things like, "shitty usage of the GPG command line tool by applications." They're welcome in my circles as I sincerely consider myself less judgmental than those calling for a ban, but perhaps in kindness I can urge people who say stuff like that to, over time, to be more kind and sensitive. Kindness is a skill that needs development. Not everyone is in the same place. It's certainly something I'm working on every day. And if not, that's okay too, not all of us are built with the same social skills, and that's okay. To _not_ be self-righteous requires long-suffering. Peter and Drew were worth more effort.
My first thought at the time was "oh gosh, it's not just me!" I think many people had a similar feeling. I don't think it's really a "witch hunt"; more a sigh of relief.
As far as I know, they don't. AFAIK the two cases discussed here are the only two notable ones (that is, people who are not outright trolls and the like).
Go package management was quite poor for many real use cases outside Google, so the community railed around many competing proposals, and when it appeared there was a clear winner, the Go team came out with their own solution instead.
This after kind of supporting the ongoing community efforts.
* Distributing build jobs across clusters
* Targeting any arbitrary platform, architecture combination
* Supporting multiple languages and toolchains
* Arbitrary build-time script execution (e.g., code generation)
I disagree. Lots of projects will never need support for multiple languages, code generation, targeting a vast array of platforms and architectures, etc and for those that do, there are workarounds that are a lot less painful than Bazel or Nix. I can go a looong ways with `go build` and some CI jobs before the pain of Bazel or Nix pay off. Basically, I don't think the "trivial projects vs everyone else" is a very good taxonomy, but rather it's "'approaching FAANG scale' vs everyone else".
If you're in the "approaching FAANG scale" group, then yeah, Bazel probably makes sense. For everyone else language-specific tools cobbled together with CI really is the least bad solution (and by a pretty big margin in my experience).
Package managment should not be created and maintained outside of the core team, maybe the process in which they decided the solution was not the best but the end result speaks for itself, go mod is very good.
But the Go team ignored it for _years_ until dep got a lot of traction and we had this big community meeting about it and live on the call Russ said he was going to work on minimal version selection and build a tool.
Sam Boyer had given a big talk about package management at GopherCon and on that call his only response when people asked what was going on was "no comment right now" or something like that. It was almost like it was his first time hearing of it.
It just seemed to be handled very poorly.
All parties involved are active on HN AFAIK so they can correct me. I am not trying to dig all the past back up just for funsies... maybe something will happen here and the Go core team will realize they need to work with Drew and sort it out. But history points to them doing what they want to do irrespective of the community.
Go has demonstrated that its governance model is capable of making better decisions than a pure democracy. This is not the case for most open source projects.
I left a job and came a few years later, and the Go code all still worked under 1.19.
Dep was very slow, tried to be very clever often resulting in it "cleverly" doing the wrong thing, and was generally a pain to deal with. At $dayjob we migrated to dep and then we had serious discussions if we should move back to glide (we didn't, as vgo was on the horizon by then). There were some plans to fix some of this IIRC, but they never really went beyond "plans".
It's always difficult if people spend a lot of time on a particular solution and then it turns out that's actually not what's desired, or if someone else thinks of an even better solution. All option suck here: chucking out people's work sucks (especially in open source volunteer-effort context), but accepting a bad solution "because people spent time on it" and then being stuck with that for years or decades to come sucks even more.
The communication could certainly have been better; a sort of "community committee" was given a repo on github.com/golang/dep and told "good luck". In hindsight the Go team should have watched development more closely and provided feedback sooner so that the direction could be adjusted. Actually, the entire setup probably wasn't a good idea in hindsight.
In general, I've been very happy with Go's design for the last decade with relatively few, relatively minor exceptions. I specifically have enjoyed that the Go team has resisted demands from the wider community to add every feature that every other programming language has support for. Biasing toward minimalism has made the language very easy to learn and consequently any Go programmer can jump into any Go project and start being productive in a matter of minutes, and any non-Go programmer can be productive in a matter of hours. Similarly, the tooling is novel in that it has sane defaults (static, native compilation; testing out of the box; profiling out of the box; build tooling out of the box; reproducible package management and hosting out of the box; documentation generation and hosting out of the box; etc).
And their weird entrenchment in "doing things the way plan9 did"
Hell, how you design a statically typed language and go "nah, we don't need sum types! That's too complex!"...
Users trying to clone their project would hit an almost certainly up to date Google cache and thus be happy and sr.ht save on pretty much all that traffic and thus be also happy.
> git clone requests with a GoModuleMirror User-Agent will receive a 429
So only one counter seems to be needed. Sure, it's doing work for an adversarial party, but it's not much extra work either.
That's precisely why it's so unfair to ask the sourcehut team to try to come up with a solution to this problem: it's the very design of the proxy that Google put together that is causing this issue in the first place. And the sourcehut team has no control over their design. And to add insult to injury: when the sourcehut team offers recommendations on how to improve their implementation, Google responds with "it's too much work".
Setting up an IP based rate limiting rule in nftables takes a few minutes at max, and I am not a professional sysadmin.
Google should definitely fix their shit, but when we see that a googler said "it would be a fair bit of extra work for us to read robots.txt", I understand the frustration on Sourcehut side, and banning an user agent is usually just a single line of configuration to add.
"Due to the relatively large traffic requirements of git clones, this represents about 70% of all outgoing network traffic from git.sr.ht. A single module can produce as much as 4 GiB of daily traffic from Google."
Of course I’m well aware this is not going to happen.
> YandexBot is a dickhead, too aggressive > User-agent: Yandex > Disallow: /
Hahaha.
User-agent: Amazonbot
Disallow: /
- peak queries are 2500 requests per hour
- some repos are 4gb in size
- let’s round and say 2000 requests at 1gb = 2000 gb/hour = 48000 gb/day
- AWS bandwidth at $0.02 / gb = $960 / day = $28,800 / month
So, one of the richest companies in the world is charging you nearly $30k monthly because they cannot be bothered to be polite. Would you be ok with that situation?
[0] https://github.com/golang/go/issues/44577#issuecomment-78531...
[1] https://github.com/golang/go/issues/44577#issuecomment-85107...
Something tells me there's another side to this story that we aren't hearing.
Extraordinary claims require extraordinary evidence.
https://news.ycombinator.com/item?id=3803568
https://news.ycombinator.com/item?id=15066518
https://news.ycombinator.com/item?id=19124324
https://news.ycombinator.com/item?id=20826618
https://news.ycombinator.com/item?id=20841586
https://news.ycombinator.com/item?id=21247759
https://news.ycombinator.com/item?id=24791357
https://news.ycombinator.com/item?id=24965432
https://news.ycombinator.com/item?id=26061935
https://news.ycombinator.com/item?id=26956077
https://news.ycombinator.com/item?id=26965146
Google is not a monolithic entity. The Go team operates pretty much independently from gmail and other stuff. Yes, there is a problem with Google randomly disabling accounts, but this doesn't really extend to Go, who make their own decisions on who to ban or not, and which are all manual. Just as you can comment on, say, github.com/facebook/zstd without a Facebook account, you can also comment on github.com/golang/go without a Google account.
It is, and claiming that a single corporate entity cannot manage itself isn't an excuse, it's an admission of further guilt.
I'm not sure what you mean by "act abusively" since the core of the product is storing and viewing source code. I mean, I could verify checksums if I needed to, and such a thing is common in workplaces I frequent.
What I asked before, and I will reiterate for an additional attempt to your comprehension, is how can Sourcehut store and view source code BETTER than Github can?
This is really unfortunate, especially as sr.ht isn't free (any longer). As far as I know, it's the only remaining managed source hosting/CI option for Mercurial users.
What a depressing way to start out 2023.
This isn't a bad thing, it was always meant to be financially supported by its users. The alternatives are ads or for Drew and co to run it at a loss indefinitely.