Add opt-in transparent telemetry to Go toolchain
github.com
github.com
They get users accustomed to the idea that now telemetry is in the toolchain. Next step will be to "accidentally" turn it on for everyone.
Then it will be "oopsie daisy" everyone is okay? See. Nothing happened, so we'll leave it opt out.
In contrast with how Homebrew implemented their opt-out telemetry, opinions obviously did matter to Go. Yes, I'm still salty about that project being so tone deaf.
Basically, there's a very obvious in group that has to pay lip service to the proletariat but otherwise can do whatever they please.
Telemetry is valuable for making decisions and identifying issues, so it's no surprise the group making decisions about the future of the project would welcome the proposal.
You can always fork an open-source implementation, but that's a lot less useful when it means your code is no longer compiler-compatible with other people's code. So languages tend to be run by some committee, and a lot of the popular ones have a Benevolent Dictator for Life or a small committee.
Guido was Benevolent Dictator for Life of Python until 2018. C++ is standardized by JTC1/SC22/WG21 (and Microsoft still has huge influence on it based on simply whether they decide or not to incorporate a feature into MSVC). Ruby is an ISO standard. Common LISP is an ANSI standard. Modifying a widely-adopted language in a way that will be seen by most users is at least as much "Can you work with those with the political influence to decide 'yes'" as "Is your recommended modification technically good?"
Coming from a language that is slowly but surely trending to an "everything and the kitchen sink" PL (aka swift), believe me, i appreciate A LOT the care with which go team makes sure any new feature composes perfectly with the existing state of the language and doesn't add too much complexity.
Continuing to hold this opinion does not imply that their opinions have somehow been manipulated or changed.
As for the Overton Window - telemetry has been a part of software since the internet was a thing. And it will continue to be a thing long into the future. I just fight to keep it as a choice for the user.
Not true. As I mentioned in a separate comment, in the 90s (and even somewhat into the early part of 00s) adding any kind of phone-home functionality to code was pretty much universally seen as outrageous privacy violation.
Go back to, say, year 2000 and post on a technical newsgroup about having noticed some code phoning home and watch the outrage fly. This was seen as a line that must absolutely never be crossed.
So yes, the overton window has done a massive shift.
Yep.
And I remember a few very heated fights between dev teams and marketing teams over it. Devs universally fought adding telemetry because it's an abuse of users and user trust. Marketing universally fought for adding it because it allowed them to sell more efficiently.
The 90s-early 2000s were the era where the internet exploded and the dev community decided en-masse to start shipping software as web apps. Web apps are notable for giving massive streams of 'telemetry' to the web server operators as a natural consequence of how they work. Nobody cared and the change was embraced because the value of those server logs was so high that web-based SaaS companies could easily outcompete people steeped in client-side "your mouse clicks are private" culture.
Eventually, web culture had to be forced on the desktop people. At Google they were distributing desktop apps and the command came from the very top (Larry Page) that Google's software updates should work just like the web. Silent, background, no end user control, no confirmation popups. Stuff just updates. Of course, huge uproar. End users should be in control of their computer etc. Nothing worked like that at the time. Page insisted: do it or you're fired, so it got done. Users loved it, that model turned out to be extremely successful and became a key competitive advantage for Chrome when it launched, one that others have since copied.
The Overton Window shifted because the window only existed at all in programming culture. It took confident and in-control CEOs like Page and Jobs to kick developers out of their ideological cubby-hole and into a place that better suited the huge majority of end users, who really couldn't care less if some guy they never met knows how many seconds they spent looking at the welcome screen. But they do care a lot about whether their tech actually works reliably.
Ehhh...
And the vast majority upvoted this proposal.
It doesn't really, not necessarily anyway.
Think of all the awful terrible laws that get introduced "for the children". While in politics there are certainly a lot more malicious actors than in development, there are certainly many people who truly believe they are doing good "for the children" and they simply don't stop and think what are the drawbacks of their proposals.
I think it's just human nature to look for solutions to improve things which they are involved in (which is great) without (often) giving enough weight to how it harms other people (which is not good). That's why it's important to push back against these things, whether in software or in law.
Go already has a bad reputation for political reasons.
For instance, it would make sense to create a set of selinux rules (or something like it) to make sure that a compiler cannot do anything other than reading its input files (and system headers/libraries/etc) and writing to its output directory, even if for instance a buffer overflow triggered by a malicious source code file led to running shell code within the compiler. Having to allow access to the network for the telemetry would require weakening these rules.
It reminds me of the classic "confused deputy" article (https://css.csail.mit.edu/6.858/2015/readings/confused-deput...), which coincidentally also involved a compiler tracking statistics about its usage.
The telemetry is opt-in,and failing to send them won't fail the compile (it won't even run on every compile). It's not really preventing you from applying your SELinux policy if you want, even if it would have been opt-out.
Define "input files." Tools have to do a combination of reading, parsing, and sometimes even version unification/downloads just to get the complete set of inputs to feed to the compiler.
Of course you can define the compiler as the tool that parses text and writes machine code, but then you're just shoveling dirty water around.
I think that's the wrong way to go about things; instead it's more useful to ask "will this be useful?"
There's a long list of real-world use cases in part 3 of that blog series.
I miss telemetry in my app sometimes too; there's some features where I wonder if anyone actually uses this, and I also don't really know what kind of things people run in to. Simply "ask people" is tricky, as I don't really have a way to contact everyone who cloned my git repo, and in general most people tend to be conservative in reporting feedback. I have found this a problem even in a company setting with internal software: people would tell me issues they've been frustrated at for months over beers in the pub, when this was sometimes just a simple 5 minute tweak that I would be happy to make.
Can I make my app without telemetry? Obviously, yes. And I have no plans to ever add it. But that doesn't mean it's not useful.
Well, that's also the wrong way to look at it. Because everything, no matter how broadly bad it might be, is useful to someone somewhere.
Of course telemetry can be useful to the developer of the application (if they look at the data and act on it). But at the same time it violates the privacy of all its users, who vastly outnumber (at least for most projects) its developers.
For any argument we need to look at pros & cons, not just the pros.
Android, iOS, Windows, macOS, Chrome, games consoles, Docker, VS Code, IntelliJ etc. They all collect telemetry and stats on how they are used. That didn't stop them becoming monster success stories even amongst the developer population because nobody cares.
The Go team are right to do this. Really, we should all be following their lead. Software telemetry is essentially pure win with no downsides for end users, which is why everyone has adopted it. It can also be done in better ways than we do now, like by recording stats in human readable form in files that sit around for a while before they get uploaded, so uploads can be turned to manual mode for inspection by the 0.1% of people who do seriously care.
I think that's the wrong question as well. The right question is "does this provide benefits in excess of the costs"?
You should only be adding telemetry if you can verifiably prove it will give you information you can't get otherwise, that you will actually use that info (what features hit the most bugs doesn't freaking matter if you are spending 95% of ever sprint doing completely different things) and only if you can find a way to legally guarantee that info is used for NOTHING else.
Microsoft did not have broad telemetry in Windows in the 90s, and yet Raymend Chen had no trouble getting popular software and running it to find out what problems it ran in to. When vista had basic telemetry, all they found out was that nVidia makes crash prone drivers (50% of all Vista BSODs) but that didn't help them at all. People still blamed Vista for all the problems not caused by Vista, and nVidia was not getting that telemetry.
Telemetry is a weird crutch that people keep latching onto before they even break their legs.
In Rust we've had large migrations like that (the "new" borrow checker comes to mind) and we had a really long periods of time where they were tested against the latest crate versions on crates.io until crater came back clean, and even then we had bug reports about regressions in the wild only after it was released on stable.
For me personally there's one big blind spot when testing only against published code, no matter how big the corpus is: humans are excellent fuzzers and the malformed code they write and try to compile is hard to replicate. Having visibility into uncommon cases that are only visible on users machines would be incredibly useful. An example of this could be the "botched" 1.52.0 Rust release[1], where invalid incremental compilations were changed from silent to visible Internal Compiler Errors in nightly for several releases to the point where the team felt all of the outstanding incr comp bugs related to them had been addressed, but when turned on in stable immediately hit users in the real world, making it necessary to do an emergency dot-release reverting that change. If we had telemetry on stable compilers, instead of turning on the silent to ICE change we could have added a metric for the silent error and know ahead of time that the feature wasn't ready for prime time. With more telemetry we could have known what was causing this. Snooping over a user's shoulder while they hit the case could make identifying the bug trivial, but of course then that user would then turn around and ask me in unfriendly terms "who are you and how did you get in?"
Arguments can be made against even implementing telemetry, and they can be compelling enough to elect against it, but shouting down the conversation from even happening is not helpful.
That's probably because it's happened in real life quite a lot.
In fact, it seems to be the opposite in some cases. Firefox and Windows, for example, have generally become significantly worse for me over time, despite the telemetry that they're collecting.
In the "best" case, software like Visual Studio Code and Homebrew have merely remained mediocre.
I've seen much better results from developers who base decisions on feedback and bug reports that have been manually submitted by users, rather than trying to make assumptions based on automatically-collected telemetry data.
I have used telemetry for mobile apps in the past and now at work to monitor performance of certain software we deploy in client data centers, and there’s a lot of things I notice in the telemetry that users wouldn’t find and report. For example, I remember I was able to fix an issue in a mobile app because I noticed startup times increasing each time users opened the app, and I could fix the problematic cache quickly. I bet most users didn’t really notice. Same at my current job, we’re able to detect slowdowns and processing bottlenecks that for the users just show as subtly erroneous data. Do they notice those fixes? Nope, but that’s the point, I want to be able to fix things before they notice and report them.
Many likely be doing a lot of switching away from Go intentionally, due to this.
When your code is running on application servers and your applications are composed of components, all the tools were already there, in the OS and as add ons, like dtrace, and in whatever monitoring tools came with your application server. Today, instead of components, we compose systems out of (lightweight) processes, and processes can be created on any device, and the replacement for the application server is the whole gamut of k8, terraform, elastic, ..., etc.
Nothing has changed in the abstract structure of our systems, its just that the current approach has the beast dismembered and spread out and loosely connected via protocols, instead of a linker or a dispatch mechanism of a platform.
You can make many arguments for and against telemetry in developer tools. Not acknowledging that telemetry helps with visibility into how those tools actually work in the wild, which in turn helps lower the incidence of bugs and speed up development, is disingenuous. You can arrive to the conclusion that even inert, opt-in telemetry is not worth it, but don't disregard out of hand the utility of it in helping their development as if it were some crazy idea.
NASA couldn't obtain that same data locally, they had to know how the vehicle and software behaved in the real situation. NASA's telemetry also ran on their own hardware, not on arbitrary users'.
Contrast that to Go, which is used in real situations internally at Google. They don't need to spy on their users, they have first hand experience using the tool.
And that’s a big disadvantage if I were to scale the usage of Go then I would say startups and middle orgs occupies like 80% if the Go team work according to logic at google scale, Go wouldn’t be this successful
Back in the 90s in companies I worked for, attempting to add any kind of phone-home code was a fireable offense. Or at least, would get you a very stiff talking to from a few very high up people in the organization. You'd never even think of doing that again. Customer trust and privacy was paramount.
As we all know, spyware slowly started creeping into end user apps and later became a flood. Now it's difficult to find any consumer app that doesn't continuously leak everything the user does.
It's become so normalized that now even developers tools seem to think it's somehow ok to leak user data.
Well done, golang team. Other companies with supposedly open languages (looking at you, Microsoft) can learn a thing or two from you.
There's always the risk that they'll roll out the telemetry setup now as an opt-in feature and then switch it to opt-out down the line, but I don't think this is the current team's intention.
Like pinky promise, trust me bro.
This whole thing looks delusional. I hope someone is going to create a fork as Golang team has lost their marbles.
The point is, this is just introducing another security threat to worry about. You now have to be aware that the toolchain can call home and ensure it doesn't happen. Someone might misconfigure it or you don't know if a release comes that has a "bug" and it sends telemetry regardless of settings.
This is just completely wrong and should be nipped in the bud.
The mere fact that they are pressing ahead with this is sinister. These things are never about "oh just don't opt in".
If someone really feels the need to send telemetry to Google and be spied on, they should use a completely separate tool that is not included in the toolchain.
If you don't trust developers' promises, why is one of those worse?
Their motivating example was something like “golang stopped working on clean macos, and no one noticed for months”, which kind of proves my point.
The whole point of the proposal is that things could be improved with some telemetry.
Now of course you could weigh the tradeoffs and say the potential for improvement doesn't outweigh the risk of misuse, but it seems clear to me that reasonable people can disagree.
> I don't think they need telemetry. Go is clearly a successful language and community without it.
Right! So why "pee in the soup"?
> things could be improved with some telemetry
Or they could not do those things? (I read "Why Telemetry?" and remain unmoved.)
> you could weigh the tradeoffs
I have.
> the potential for improvement doesn't outweigh the risk of misuse
That's not my argument. It's not even misuse that I care about here (I do care about that, but that's a separate concern, and one that doesn't arise if you don't collect data in the first place, eh?), I care about use. I don't want my compiler to make network connections to Google or anybody, for any reason.
> reasonable people can disagree
That's what I said.
I assume what the GP means is to infer a strong type for a struct that has all the data a context has in it. This is much harder, for many reasons. A context that comes in from one path may have a RequestSource in it, but another may not; the resulting type of the function is rather complicated to infer. We prefer not to have two different functions as a result. There's also the problem that such inferences end up strongly tying types across many functions together, such that up in some middleware for your web site you add a new value into the context, and if there was a strong type for that context, that strong type would cascade throughout everything that could someday possibly touch that context. The result is much like checked exceptions in Java.
This is perhaps the hardest practical type problem I know. It seems to me to be very related to a similar problem, which is that of trying to strongly type errors. The sort of cross-scope type inference we envision in our heads is extremely unwieldy in practice, if you take the time to try to scope out what that would convert to in a real program with many nested scopes and arbitrarily complicated paths in to those scopes, plus arbitrarily complicated closures being passed around that further impact the types.
Contexts are at least a value you can see and manipulate rather than having something attached to your thread, and in particular, can pass across thread boundaries if you need to.
One way or another, you end up needing some sort of scope that carries values that can't be strongly typed because the way you're composing those values doesn't work with any known strong type system in practical use. (I've seen some super theoretical ones that can in theory do it but I've never seen them brushed up into something practical and successful.)
C# is pretty neat in having an AsyncLocal[0] that goes beyond what thread locals enable
[0]: https://vainolo.com/2022/02/23/storing-context-data-in-c-usi...
I don't know I never actually remember the commands. :o)
https://en.wikipedia.org/wiki/Telemetry
Look at that list under Applications. What makes software so different that it shouldn't be included?
Though at this point it's been widespread long enough that I suppose it's just a normal term that's "always been used", to some developers, not a new, alien-feeling part of the software lexicon.
Microsoft OTOH calls their data collection the "Customer Experience Improvement Program" - now THAT is whitewashing.
It's only been a decade or so now that you could rely on most computers always having an internet connection. That's probably why the term feels new - it only started being used when the practice became technically feasible. Maybe others remember things differently. /shrug
Like neural networks?
https://en.wikipedia.org/wiki/Dynamic_programming#History
Sounds super-fancy and advanced, but when you dig in it's like, "oh, that's all?"
Edit: have been doing software since the mid-nineties. You can downvote all you want, but this is a more recent usage.
The opt-in thing is fine, but some of us are still stuck, I guess, on older standards for software ethics, and find it entirely unacceptable and alarming that opt-out was ever proposed in the first place, about as bad as if they'd proposed adding an opt-out bitcoin miner to it to help fund the project—it's disturbing they'd consider that OK to even propose.
Telemetry is generally for the benefit of the marketing team, law enforcement and development team.
(some time in the back half of the '00s, I reckon)
I for one plan to enable it, and I hope others will do the same.