By asking nicely. Not by siphoning data out of my system.
Where is this world going ? Do i really have to treat every program as spyware and restrict its internet access ? And anyway, why would a compiler need internet access ?
By asking nicely. Not by siphoning data out of my system.
Where is this world going ? Do i really have to treat every program as spyware and restrict its internet access ? And anyway, why would a compiler need internet access ?
The problem is always how to randomly sample people and AFAIK there is no good way except randomly sample people (a.k.a telemetry). You can argue that developers shouldn't care that much about their software are being used and whether they are performing as expected, but asking nicely really is not a solution.
APPENDED: PLEASE NOTE I'm saying asking nicely is not a solution to get correct data from statistical standpoint. You can argure the data accuracy is not important in this context.
LLVM, GCC, rust, zig, D, etc don't have telementry and they seem to be doing fine. What makes go special? They aren't trying to be more efficient than rust. It isn't more widely used than GCC. I would argue they don't need this telemetry at all.
But that doesn't mean it's not useful, for Go but also gcc, or Zig, or D, or any other language.
A long list of use cases is given here: https://research.swtch.com/telemetry-uses
Apparently as soon as a company becomes big enough, the line between convenience and ethical becomes very blurry.
I never said anything about what should be done. I'm just saying it is useful, as the previous commenter said it's not.
Where I'd expect difference is precisely over the question of whether it's worth the cost & trust issues. That makes me wonder whether there might be some middle-ground here where some third-party like the Linux Foundation could run a telemetry service which is highly public in both its design and collected data, and whether that would be enough that many people would be comfortable with this kind of service without Google's business reputation entering into the equation.
Useful does not equal ethical or legal. A lot of things would be useful to me.
Here's what I think would be a closer analogy to what's actually being proposed: the water company installs a smart meter allowing them to read usage data more frequently. They randomly sample 2% of customers' data and prepare a histogram showing the top daily usage hours and total usage over a monthly period.
Probably the data is too wild to make any use of, but then all the folks who have a clue using the playground as a kind of proof-of-concept play area... would be great for analysis.
Anyways, I'm not sure how well it would work, regardless of if it should be a thing at all. The majority of stuff I do, and the folks at work, we all have build pipelines that control internet access, use cache systems, etc... so I'm not sure it even matters.
The main thing I think this would do is help them flesh out the non-mainstream configurations. For example, I've used Go on AIX but I'd be surprised if bug reports on that platform didn't hit basically every maintainer by surprise since statistically nobody tests on it regularly or even has access to do so. I'd think most of the value of systems like this would be finding out that the change they tested exhaustively on everything Google normally uses is also going to break 100% of some niche community.
Some people are posting here as if this is already decided -- AFAICT, that's not the case. It's not even a formal proposal yet, and the stated intent was to start a conversation around something concrete. (For context, this is standard for how the Go project approaches large topics, including for example I think there were something like ~8 very detailed generics design drafts from the core Go team over ~10 years).
It sounds like the Go team is going to take some time to look into some of the alternative approaches suggested in the feedback collected so far.
In any event, this is obviously a topic people are very passionate about, especially opt-in vs. opt-out, but I guess I would suggest not giving up hope quite yet.
[0] https://discourse.llvm.org/t/rfc-lldb-telemetry-metrics/6458...
https://www.oracle.com/java/technologies/javase/terms-java-u...
https://learn.microsoft.com/en-us/dotnet/core/tools/telemetr...
Also the idea of GCC is more used than Go and GCC does not have telemetry so GO doesn't need it is completly unrelated.
Most dev working on core tools want telemetry to improve those tools, it's invaluable data to make the right decisions.
Which Java distribution? Maybe the Oracle JDK/JRE? I highly doubt the Oracle JDK/JRE has serious usage numbers because of it's (quite commercial) license.
I wonder if there's some trend where corporate backed, large "open source/source available" projects being more open to telemetry compared to more "open community" projects.
For non-corporate projects, the general thought process might be it's fine to leave efficiency gains and improvements on the table if it means violating some foundational principles. While in a business setting, those efficiency gains could be very tempting since it can translate to more money, market share or promotions and that way of thinking gets applied to the open source projects as well.
I mentioned that because an argument could be made that they're the most popular, or they're trying to be the most efficient language and that's why they might "need" this telemetry data.
> Most dev working on core tools want telemetry
Yet somehow, almost all open source, popular, efficient tools don't have telemetry. Like the ones I listed + python + Linux + coreutils + bash.
Of course they want this. But there are mountains of successful projects that work very well without it, and every one serves as evidence that this Isn't as important as they let on. We shouldn't subject our personal computers to google tracking because they want more tracking data.
In case anyone wants to mention that it's Google's language and they can do what they want with it, this is a perfect example of why people aren't happy with Google's control of the language
it's owned by an ad company
It's just that, if the argument in my last paragraph did become widely accepted, it would lead to a pretty brutal society. (And, of course, it is becoming pretty widely accepted where software is concerned.)
Here's a long list of questions Russ Cox was interested in - notice how many of them would really want something like “what is your CI configuration?”, and consider that while they could sample GitHub Actions configuration that doesn't tell them anything about the much more varied set of projects who don't use that service:
Most of these things are in the category of bug prioritization. Why don't you... you know... talk to users?
If you want some statistics (oh sorry, histograms), maybe pull down the gigabytes and gigabytes of Go code residing on public repositories and places like GitHub and analyze all that. That'll tell you a ton about how people structure their projects, what architectures they target, etc.
eg, some options:
a) No telemetry. All users of the Go language get equal treatment/consideration by the developers ("best guess").
b) With telemetry. The users who opt out are thereby excluded from consideration ("punished").
It seems like over time, option b) would lead to problems for the Go language as the developers base their decisions on incomplete information, whilst believing that information is somehow more complete than option a).
Questions about bias in reported data are legitimate and a genuine topic of concern, but this:
> This may be by design, as it punishes those who dare to value their privacy.
Is just veering off in the conspiratorial.
The problem is an obvious one but doesn't seem to be considered. :(
> Is just veering off in the conspiratorial.
Punishing people who value their privacy seems like it'd be strategically useful to Google, an ad company with demonstrated intelligent but sometimes ethically challenged people. ;)
Generics were also controversial and it took multiple iterations and years to get them done.
I wouldn’t call it “suddenly”.
Also, why do they need the data “suddenly”? What do you suggest?
Reading through the use-cases at https://research.swtch.com/telemetry-uses one thing to consider is that these are the kinds of things people care about as a project matures — once you have a large number of people depending on your project, you have to worry more about unintended consequences as you evolve and since the user community has grown larger and the project more stable it's also harder to know whether the people you hear from regularly are representative. When a project is small or mostly used by enthusiasts you can be more confident that your understanding of what people are doing is relatively accurate but that just isn't true over time. For example, how many millions of Java developers never contact Oracle or the OpenJDK developers throughout their entire careers?
Well, to bad. Because if you don't, you've built a piece of malware, by definition. You need to ask, live with bias, and draw your conclusions accordingly.
Who cares?! It's about respect for individuals and respect for their privacy. Any kind of telemetry should be opt-in no matter how you spin the PR.
Allow me to help you. You ask first. If the answer is no, you don't do it. If you cannot master this skill, or refuse to, you are embarking on becoming one of the most problematic demographics on Earth.
There is no excuse for not asking, and honoring the answer. Sending data is not free, and I guarantee your telemetry is not that important, neither is your vaunted sampling. All you want is to know who your Users are, and the fact is, that is fundamentally gated behind them wanting to even bother telling you.
Just like answering polls makes more of a difference than voting.
The solution, at least for me, is to remove that software from my life. Personally I vowed a long time ago to never work for a company whose business is to track people. It's no big deal for me, but I understand at least some may have hard decisions to make.
If monitoring of the application isn't possible (And by possible I mean reliably statistically so opt-out not opt in) then that will drive the shift to centralized/web based applications even faster.
Ironically, those who are most wary of telemetry tend to be the same people who also appreciate being able to run local software over web based.
I write some paid local-only software, and I have no trouble not tracking users or usage. For example, here's the entire privacy policy for one of them:
> App does not collect any data.
> App works completely on-device. Messages do not leave your iPhone and are not shared with other apps.
> The essence of this privacy policy will not change.
--
For another software, the privacy policy includes details such as:
> App never uses the Internet, except to check for updates.
> Automatic update checks can be turned off in the application's preferences. Requests to check for and to download updates happen only through the domain domain.example. The updates server does not store or forward potentially identifying information, such as request IP addresses or user agent strings.
> License keys, obtained upon purchasing App, do not contain personal information, such as names or email addresses of license holders. License keys are validated locally; App does not use the Internet to validate license keys.
If I'm deciding between a desktop and a web client for my software, monitoring is a concern. If I suspect I'll either a) not get enough opt-ins for montoring or b) end up in a publicity shitstorm if I use opt-out, then I'll just opt to run the software on my server instead.
The amount of monitoring and tracking that is done on server side software is obivoulsy going to be a LOT more than what "anonymous usage data" would be for the client version. And that's a bit ironic (that it might have the opposite effect).
Obviously scenario isn't about this software specifically. A compiler isn't going to be "web based" tomorrow. It was more a comment about the software landscape and the privacy debate in general.
The lack of basic ethics on this topic is a great example of what people mean when they assert that software engineering isn't real engineering. Software should represent the interests of its prospective users, period. Inserting anti-features that directly go against the interest of the user to benefit the author is a violation of trust, and is blatantly unethical.
But sure, keep convincing yourself that if users don't like you violating them a bit, it's just naturally pragmatic to abuse them even worse. Either way whatever you create can't be considered trustworthy.
As for the original topic, the real problem is that it would be a rug pull - classic Google. If this were a new language built with surveillance in the compiler, nobody would use it. But if the maintainers decide to add in hostile features after the language has gained wide adoption, it will take a lot of churn to sort out, fork, sandbox, etc.
https://wiki.debian.org/PrivacyIssues https://www.qubes-os.org/
You can disable the Go proxy and go for a more direct source fetch with an env variable, but in the same way you can also disable the telemetry.
My feeling about Google is that they do not respect privacy, no matter the product. Invasion of privacy and tracking their customers are behaviors deeply ingrained in Google DNA.
Doing so for all other products has been incredibly profitable, so why not Go? It’s not like Go the language has a history of respecting the PL community (the resolution of Google stepping on the Go name was Google throwing their weight around and told the Go! author to pound sand), so why should anyone expect they respect their users?
If you're using an application from Google or Microsoft or any big tech company for that matter: yes.
> And anyway, why would a compiler need internet access ?
The go build tool will fetch dependencies for you. If you don't use dependencies or use a local network proxy instead of the Google one you can safely disable WAN access, but otherwise you'll end up with a broken tool.
someproject% go get ./... # or 'go mod download'
Next disable network access for the go command, or for the whole system.Then compile:
someproject% go buildThis leads to the compiler having access to the internet and possibly exercising its access. In a perfect world of separated concerns the compiler binary wouldn't have any networking code at all, necessitating a go-get before calling go-build.
I understand why the Golang team decided to go with this approach, but it does have side effects that people coming from different tooling (say, C) wouldn't expect.
Nix got there via reproducibility, but it's exactly what's needed to mitigate compilers with surveillance features or other backdoors.
It doesn't. It compiles without internet access. If it has internet access it can also send telemetry data. If the compiler refused to work in all cases where it couldn't send telemetry data, I'd see the point...
Languages like Go have enormous test code bases with unit tests to ensure that every corner of the language is compiled and executes correctly. We don't need gobs of telemetry to phone home about how people are using a compiler.
The tests are the map, the actual use of the language is the terrain. The map is assembled largely by blindly guessing what the terrain is likely to look like.
That isn't very scientific.
The result is a situation where users will scream at the language authors if they break something, but they'll refuse to share any information about what they're actually doing with the language, which is perverse.
There are numerous compilers for C++, one of the hariest languages around, that were developed with no telemetry and work just fine and cover the entire language spec. Some of these compilers like G++ and Clang++ manage to target many different architectures with very few issues.
The obsession with telemetry is a mix of laziness and I strongly suspect user-hostile surveillance motives.
We're just going to have to lock down OSes. Giving any application carte blanche network access is like when apps had open access to all of system RAM and the whole filesystem back in the MS-DOS / Windows 3.x days.
Maybe unauthorized (by the user) telemetry should be like invalid memory access and trigger something akin to SIGSEGV. That'll send a message.
The discussion would be a wholly different one if the data could be used for anything to e.g. sell, target ads or even deduce how/whether an individual has even used the application.
There's a big difference from functionality not existing to do this, and functionality existing that can change at any time. rsc's attitude about this comes with an air of "I want to alter the deal. Pray I don't alter it further"
That would be asshat-y but as long as no PII is transmitted, it would stop short of being a legal issue at least.
This is the thing though I think: people don't like to give up control (in this case, about what their software does). It's not about privacy, it's about control.
THIS !!!
I suspect like many people, I feel that if they had said "We're going to introduce this ... but don't worry, it will be opt-in", then they would have avoided 99.99999999999999% of the negativity.
Instead, by making it opt-out, they are basically doing a classic Google shit move.
And that's before we start getting into questions about GDPR and opt-out.