VS Code – What's the deal with the telemetry?
roboleary.net
roboleary.net
[1] https://imgur.com/a/RbrsuyA
You will want these three settings in your default settings JSON (the first two just to be safe):
{
"telemetry.enableTelemetry": false,
"telemetry.enableCrashReporter": false,
"telemetry.telemetryLevel": "off"
}> ...
> return TelemetryLevel.NONE;
You gotta love how the documented function return value (TelemetryConfiguration.OFF) doesn't match the actual return value (TelemetryConfiguration.NONE).
One of the comments mentioned that there's already a setting to disable it, ain't that enough..?!
I'm once again recommending OpenSnitch: because Microsoft won't take "No" for an answer.
It uses https://open-vsx.org, an open source host for VSX plugins.
They try to remove all telemetry, read more about that here.
...
However, VSCodium can’t shut out all the data collection as it is the same codebase. And since extensions act independently with regard to data collection, you still need to be mindful of what extensions you install.
Vscode as an IDE is quite useless without it's extensions. And that includes official(?) extensions from Microsoft that also might have other policies regarding tracking.
I think it's great that open license terms enable these sorts of forks and experiments so that the community can work together by sending pull requests.
Is there a good way to install code-server (hosted VSCode/vscodium in a browser tab) plugins from openvsx?
(E.g. the ml-workspace and ml-hub containers include code-server and SSH, which should be remotely-usable from vscodium?)
source: https://github.com/VSCodium/vscodium/blob/master/DOCS.md
The case that we allow users to control privacy needs to be scrutinized more.
Telemetry is a privacy nightmare for users, since it sends data to the developers outside of an average user's control - data which is probably easily associated with an individual and kept for eons (disk space is cheap after all).
I think telemetry collection is starting to become a professional ethics issue for developers, and should be talked about more.
This is is actually a terrible trap, as product usage tells you next to nothing about which features are useful. I suspect it' a big part of why a lot of software has turned to shit last decade or so.
Look at it this way: Your average fire extinguisher sits mounted on a wall its entire life and is never actually used. You still rather have it and not need it than need it and not have it.
Agreed.
> You still rather have it and not need it than need it and not have it.
And there's the trap. Data is a liability, but until a company gets burned for losing it, they won't do anything about it. Well, being real, even if they do get burned it's just the "cost of doing business".
> You still rather have it and not need it than need it and not have it.
Given the context of the fire extinguisher, "it" in this case is a useful (possibly critical), not-commonly-needed feature, rather than a collection of data.
Do we think the people making calls like that would suddenly have great ideas in the absence of data?
Meanwhile, there are plenty of folks out there who can reason about data in less blindly idiotic ways.
Option A generates 10 times more searches than option B. This may mean the search box is 10 times more useful than in the B case, or the users require 10 times as many searches to find what they want. Looking at the data, you really can't tell. Even without the data, you can dogfood your application and fairly quickly tell whether the feature is good or bad.
I'm not saying anyone is stupid. Naive, perhaps, but not stupid.
The problem is that interpreting data is difficult. Incredibly difficult. Scientists, who construct experiments and interpret experimental data for a living, with a decade of education behind their back, get this wrong all the time.
Personally, I would be so bold to suggest that yes the availability of data for people ill equipped to interpret it can lead to WORSE decisions than the absence of data.
That is because statistics and the general skill to correctly interpret data is actively counterintuitive. Untrained people will be more likely to generate wrong conclusions than right conclusions.
So, yes data in the wrong hands can and will be harmful.
Certainly, there is a lot of software out there with poor polish. My pet peeve is that devices like multimeters that used to have segmented VFDs, so had a certain design aesthetic by default. Clean, straightforward, readable. Now those have moved to low resolution LCDs, so you notice that the font they picked is terrible, and the font rendering engine they picked is even worse. It just looks bad, in exchange for more flexibility and features.
But, none of this is indicative that the software itself is getting worse. In the 80s, trivial software bugs were literally killing people: https://en.wikipedia.org/wiki/Therac-25 In terms of radiation therapy machines, the underlying business hasn't changed much; people with tumors want them to be killed, and the difference between now and 40 years ago is that the machine doesn't kill them because of integer overflows.
(One interesting note about the Therac-25 is that it reused software for previous models that had hardware interlocks that masked the software defects. They didn't have any telemetry on when the hardware interlocks triggered, so they didn't know that the hardware interlocks were masking bugs in the software when it came time to remove them. If they had the data, they might have kept the hardware interlocks, or fixed the bugs before 6 people died. So maybe lack of telemetry is a greater ethics issue than including it!)
The main effect of the telemetry in self-driving cars is that we can all Monday morning quarterback the exact defect that caused a particular crash. With human drivers, we don't have as much data to get to the root cause. If someone falls asleep at the wheel and drives into a truck, was it their medication? Their sleep schedule? An emotional conversation they just had? The music that the radio station chose to play? A combination of all of those factors? What actions do we take to prevent it from happening again?
And people are terrible at putting in the wrong coordinates or dosages in medicinal machines. It's as interesting a data point as the Therac.
> The main effect of the telemetry in self-driving cars is that we can all Monday morning quarterback the exact defect that caused a particular crash.
Unnecessary, and has been unnecessary for almost as long as we've had cars. Investigators (the NTSB) are fantastic at looking at wrecks and determining the underlying cause, even in the absence of the black and red boxes. There's no need for always-on-always-phoning-home telemetry.
And this is where I get annoyed - there's no value in companies having that data when their most public use of it is to get popular opinion on their side in a fatal wreck.
Seriously, fuck Elon Musk for that horseshit.
This seems exactly backwards to me. No one who has ever managed a deployed product with good failure telemetry would say that it tells you "next to nothing". Telemetry detects failures that mere testing simply never can and never will.
And no one who lived through earlier ages of consumer software would ever say that it's "turned to shit". Software is getting better. It's getting better steadily and inexorably, and it's been doing that for half a century. It may not be doing what you want, but it's doing it with higher quality than people in the 90's or 00's would have ever dreamed. When was the last time[1] you saw a major consumer application (the TikTok or Instagram client, or your web browser, or VSCode) crash hard and fail on you? This was a daily (or worse) event for most of software history!
And certainly robust telemetry practice is one of the big drivers of this quality revolution we've seen.
[1] Heh, cue the peanut gallery of everyone wanting to post crash reports. That's on me for poor framing, I guess. But seriously, everyone: Both anecdotally and via actual studies, consumer-visible software is getting much better.
Huh. I've used software since the mid 1990s, and I can't recall using software that crashed daily, let alone frequently. Which is honestly a bit of a miracle, since a lot of software back then was shipped without the ability to patch it at all.
Product usage and failure reports are different. It wasn't called telemetry when apps just asked to send failure reports.
> When was the last time you saw a major consumer application (the TikTok or Instagram client, or your web browser, or VSCode) crash hard and fail on you?
An hour ago or so.
Google search on Chrome on my home computer has a habit of just returning a white page after about a day of computer uptime. I have to switch to a different profile to get it to show anything.
Magento CMS will silently hang if you try to perform any query on a tab that's logged out. Default login cookie expiration time is one hour.
iOS keyboard has started suggesting german words in autocorrect. A third of the words I get on swipe input are now worthless. German is not in the languages list on that device. I read a suggestion to uncheck "German" in the dictionaries list. It was indeed checked (Why?? Because I frequently type in my German last name?) but unchecking it did nothing.
Discord app on iOS has a hard crash bug on the emote picker if you type a query that has no results, backspace it, then try to pick one of the new results. I actually tried reporting that one and they told me to fuck off. (Which is the default response of every customer service team to any bug report, of course, since 99% of bug reports will be from civilians who have no idea how to report a bug)
The apps I use are now riddled with hang bugs rather than crash bugs. Huzza.
Applications not often, but video games hard crash all the time.
Yesterday. The kindle app. It has a habit of soft locking on a book's cover. TikTok's failure mode tends to be more of a "refuses to play a video" kind. Especially if you're scrolling through videos and one does not exist. It doesn't like that. Instagram soft reboots pretty often (you can tell because you're back to the home feed, as opposed to where you left off), especially after being suspended.
I have to reboot my Mac weekly, both for updates and to keep it from misbehaving oddly.
It's one thing our development culture actually fosters these days - let it fail and restart in a known state; as opposed to doing what you can to keep the software alive through non-fatal errors.
Before that, Fastmail on Android, this morning.
Slack on iOS gets stuck a few times a week.
Frequent use of a feature is just as much a signal that the feature is an obstacle as it is a signal that the feature is beneficial.
* Feature X is good and I like to use it.
* Or feature X is broken so I have to try multiple times.
* Or feature Y is broken so I have to try using another feature instead.
Like, most of the times I open a settings window it's not because I like settings but because the program is not good enough as it is. Or if I follow a lot of links on your web page, it might not be because I like your web page but because I'm unable to find what I'm looking for.
Personally I refuse to work with or consult for companies that don't get this. It's a matter of professional pride. If I wouldn't use it myself, I'm not going to foist it on my users.
So they include an OFF switch. Does it also switch off checking for new versions of VSCode? Does it turn off consulting online sources for dictionary updates? Does it turn off viewing the extension library? Those requests are technically also telemetry but if you disable them they look like broken features. The license agreement is written with open language to say "There are some things you can't turn off" because MS doesn't know a priori whether a court would consider "updating the extension library" telemetry until someone drags them into court and presses the issue legally, and they want their asses covered either way.
And the OFF switch doesn't switch off any extensions because the API is open enough that extensions can do their own network access independent of VSCode.
In practice, it turns out to be very hard to build one guaranteed-to-work OFF switch. Not that it isn't a laudable goal. But in general: if it has online features, assume telemetry == true.
Corporations demand more certainty than "might carry the day" from their legal departments when drafting terms of service.
As a user and engineer, I am more than happy to share the usage and crash data with them, but at the same time, I would appreciate if they can be fair to us the user.
The exchange was rarely made explicit in any kind of formalized consent sense.
The veil is thin between the internet and offline computing these days.
Extremely true. Data gleaned from actual usage instrumentation is radically higher quality than self-reported.
> Telemetry is a privacy nightmare for users, since it sends data to the developers outside of an average user's control - data which is probably easily associated with an individual and kept for eons
None of that is usually true. In the common case, users don't care (after all, MS has the telemetry to estimate how many users disable telemetry). And product improvement data is pseudonymized and usually worthless after a version iteration.
I think the folks that have an anxiety about product use data collection don't realize how worthless individual data points are. The data is pseudonymized out of necessity if for no other reason; it comes in as such a firehose that it has to be immediately analytically bucketed or it overwhelms storage at the volume these tools get used. However, the telemetry options get defaulted to "on" because volumes of data are invaluable for concrete analysis of how the product is used, counts on feature access frequency, understanding pain points in usage flow, etc.
I've done some work in this space (not for MS) and am happy to answer any questions people might have that aren't "Who did you work for?"
My 20TB drives say otherwise
I have come to think that telemetry used to track actual product usage is a bad thing. It encourages shallow thinking about how users really use the product, and often leads to terrible product decisions.
Examples: replacing/removing external IDs, ensuring data can't be linked to users, aggregating data, not retaining data, only processing data within data centers (no data extracted to laptops), deleting old data, etc.
This is why it's slow and painful to work at big companies compared to startups. It doesn't mean that accidents don't happen, but it's pretty well thought out. The same applies to how they do qualitative research. It's got reasonably high standards which are fairly consistently applied. It's not perfect, but nothing is. There is a lot of job security and motivation for people to make it better.
Smaller companies have almost zero understanding of how to deal with this due to lack of resources and expertise. The only way they get compliant is by using good tools which force the user into a safer, more compliant posture.
When the whole world of internet services began consolidating around the web and then later when the whole web began consolidating around a few powerful walled gardens, I saw how these trends began chipping away at our anonymity, privacy, web performance and software performance. Computers are ubiquitous but privacy is under attack. Hardware is faster but software is slower. Internet pipes are faster but websites are sluggish.
In all this turbulence I thought at least my software development tools are not affected by these terrible user-hostile trends. Emacs has become faster with faster hardware. While Vim was always fast, I am sure the Vim folks too would agree that they can now run Vim with a lot of plugins thanks to faster hardware. These editors do not do hidden telemetry. They don't add user hostile features. The core editor experience remains more or less the same year after year.
Even after all the disruption (I mean literally, disruption, not in some positive metaphorical way) in the rest of tech world, I have found consolation in my modest code editor, Emacs for me, Vim for some, other editors for other people. My code editor has always served my best interests. It helps me write code and documentation without any distractions. But when I read articles like this about VSCode and its telemetry, it really makes me anxious. Perhaps Emacs or Vim will never be afflicted by issues like this. But still ... If developer tools meant for ordinary software developers like me are going to start sneaking on my data and begin bundling user-hostile features, what hope is there for all other kind of software tools!
I still don't understand why developers use Microsoft tools. At all.
About the only valid use case I can think of is creating Windows native apps and _maybe_ testing on IE/Edge.
Otherwise, why? Just why would you wade in the shit when there's a lovely clean pool right over there?
this pool isn't filled with shit, and the water over there isn't very clean, thats why.
Two things helped get me into Emacs full-time (and this is after > 15 years of using vim):
1. I went step-by-step through Susam's Emfy Emacs config [0]. That helped me understand some of the basics at a foundational level. I extended that base configuration a little bit and became comfortable with the environment.
2. I then went step-by-step through the entire "Emacs from Scratch" playlist that System Crafters put out [1]. I pushed my personal configuration pretty far with that over the course of 2-3 months.
I eventually moved to Doom Emacs and married in pieces of my own configuration. That's been my daily driver for months now.
[0]: https://github.com/susam/emfy
[1]: https://www.youtube.com/playlist?list=PLEoMzSkcN8oPH1au7H6B7...
> Why have licenses like these if they are just concerned with product improvement?
That nails it for me. I have no reason whatsoever to use vscode.
There certainly is a valid argument for Internet access, for example for documentation lookup, schema validation, database editing, etc.
There would be a strict, and yet fallible, vetting process on the extension marketplace to make sure every extension complies with whatever rules are defined.
Expose the setting to extensions (if it isn't already), and set a policy in their extension catalog that extension telemetry must honor the global setting. It isn't perfect, but it's better than nothing.
> extensions may separately send telemetry, that do not have to adhere to the main VSCode telemetry configuration settings.
From what I see things that should never happen and should yield a block from the marketplace.
Like you mention, in the end of the day that is not enforceable, but more transparent data privacy policy would show some good faith from their side.
Here is one from a random person on the internet which has the license "All rights reserved": https://marketplace.visualstudio.com/items?itemName=LiamNevi...
Here is another by Adobe that doesn't have a license linked to it (presumed proprietary): https://marketplace.visualstudio.com/items?itemName=com-adob...
> Extension authors who wish not to use Application Insights can utilize their own custom solution to send telemetry. In this case, it is still required that extension authors respect the user's choice by utilizing the isTelemetryEnabled and onDidChangeTelemetryEnabled API.
I suppose the quote in the article is technically correct, because there's no guarantee that _will_ follow this. I'm curious if I could report an extension for abuse and have it removed if it doesn't honor the global setting.
The article also says
> Microsoft’s C# extension (ms-vscode.csharp) sends data to Microsoft. There does not appear to be any setting offered by the extension to turn telemetry off.
I unzipped the extension and looked at the package.json, and it appears to use Microsoft's recommended extension-telemetry library, so I presume it is following the global setting.
I wish Microsoft required extensions to publish detailed telemetry info (or, really, info on any and all external connections an extension might make) on their Marketplace page.
[1]: https://code.visualstudio.com/api/extension-guides/telemetry
That's a legal minefield.
It's what they do. It's been 20+ years. I'm not even old. I just don't understand how people haven't figured this out yet.
They were losing hearts and minds in the Web dev space by waging war against open source. It makes far more sense to co-opt it, as per the Halloween-document strategy.
TBH I saw this coming 20 or so years ago, about the time when IBM started putting on its open-source cheerleader outfit. Open source is formless, shapeless, like water per Bruce Lee. You cannot oppose it with force, it will just get everywhere; but by aligning yourself with it you can direct it. That's what I saw IBM doing and I knew Microsoft would follow suit out of lack of choice. And now here we are.
I think we're still dealing with the same Microsoft that we've dealt with through the 90s. They are not a champion of open source, and they are still up to their old tricks.[2]
[0] https://news.ycombinator.com/item?id=31966414
[1] https://keivan.io/the-day-appget-died/
[2] https://social.platypush.tech/@blacklight/108719097530863121
These solutions have privacy nuance they are not black and white/evil v good
Give user option to enable to improve product -> Good
Our goal should be to adapt and ensure privacy via regulations and standards. Wishing, or enforcing, that every application not have telemetry by default is both unrealistic and unwanted.
But the point isn't being good or evil; businesses have no sense of morals, they don't obey to common sense but rather take the path that leads to the higher profits, and these days profiling users unfortunately brings high profits. It's not a matter of which company is good or bad, or if telemetry can be done right or not, but rather how long until a software with built in telemetry will use it the wrong way, because soon or later everyone of them will be presented the scenario in which doing a bad thing brings them more profits; and this principle applies way beyond telemetry.
My point is - something being useful is a very low bar. It's not the point of the discussion here. Maybe some privacy zealots won't acknowledge this, but telemetry can be very useful, and at the same time a breach of privacy.
And regarding Microsoft, it's also about their attitude towards it. It's their way or the highway. It has always been like this - the software might be the way it is, but regarding their business dealings they are and always were ruthless. That's what made them great. So now what we see is a continuation of their previous attitude. And a slice of the current zeitgeist which is always-on internet with everyone slurping up user data.
The solutions should have a privacy nuance... but don't. For example from TFA, collecting git branch and repo names. That seems like a terrible idea, since it's not like Microsoft's devs will have access to reproduce issues from those repos. It's also a potential leak of sensitive corporate data.
I wish they could find a way around some missing plugins, but it does what I need it to.
https://github.com/VSCodium/vscodium/blob/master/DOCS.md#ext...
Thank you! Was always annoying when coworkers would tell me about a plugin that I could not install (or worse, had an old broken one)
You can go a bit different way: run code-server or OpenVSCode Server in WSL and open them in a browser.
I'm too afraid to run it.
Does this exist?
What would be even better was a solution that had profiles for common usages and apps, and the profiles could be created by all and shared via a community or repo and could be extended and tweaked.
Does this exist, is it feasible, or does anyone have a better simpler idea?
Sure, Sublime Text.
alternatively, go climb the ladder at any company that collects this terrible scourge of invaluable personal explosive data and get them to stop.
I'm ok with this. and just like you didn’t like me belittling your position, as i did above, i’ll kindly ask that you understand my reaction here to being told (without asking, by the way) that telemetry is an unparalleled bad thing.
i wish people who are convinced that some illuminati-like cadre is behind all of these things were employed at large companies which could collect this info. there is nowhere nearly enough cohesion between employees to make anything useful out of all the data collected, much less to act upon it, outside of the developers trying to chase the bug that has haunted them for 36 months, and the UX team trying desperately to reduce the amount of choices in menus. i mean, manage a few projects with people who are all on board with the effort - even those groups are hard to keep focused, most of the time.
these companies use this data for the stated purposes; to help make their products better. if i could explain some of the technical considerations behind the apparently odd choice of data to collect, it would make a lot more sense, but i’m pretty awful at communicating, and even if i weren’t, and i were able to succinctly convey that info, i feel like most of you would just argue with me anyway.
If you use VSCode or its extensions you can't keep commercial secrets from those working at Microsoft unless you block network traffic from VSCode. This seems to be true even in the open source builds like VSCodium.
Now the program does have a variety of spots where it is designed to contact Microsoft services, like say the extension store, or update checks, or other similar functionality. Those services will (obviously) keep some level of logging data, which you cannnot opt out of if VSCode Talks to those services. This is what the terms are talking about when they indicate that not all collected data can be opted out of.
There is no "Don't talk to any Microsoft cloud service" feature toggle, sure, but would not really be a sane toggle people actually want. For example that toggle would need to forcibly disable the git support to ensure those features don't try to talk to GitHub if you open code in some GitHub checked out repo. It would disable the extensions store. It would prevent the update checking, etc. VSCodium would end up wanting to rip chunks of that toggle's functionality back out, since for example, VSCodium uses a different non-Microsoft extension store by default, so would want want the extension store to remain disabled.
Which is ironical given that you should be able to completely strip any telemetry code if the complete source is available ...
I don't believe my coding habits can be used for advertising.
I don't work at Microsoft but I work somewhere similar on a product that collects similar types of telemetry and those words are not empty, they are very, very true. We absolutely add, remove, and change things about the UI based on what our telemetry says people do.
And it's not just about moving things around (which I know HN hates), it's about making the page load faster (which I know HN loves!). We are going through a concerted effort to lower page TTVC at the moment and a big part is going to have to be turning off some features that show by default that take too long to load and guess how we'll be making the call about which features to turn off?
Please consider stopping with that UI changing thing :)
* Profiling your work hours
* Tracking the physical locations of software professionals
* Profiling the nature of the projects you are working on
* Perhaps: training GitHub Copilot?
On the other hand, sites such as StackOverflow and Google (possibly even Bing) can get a good guess of the others as well.
After reading, this seems impossible, short of going down to DNS or network blocks of some kind: the "disable telemetry" setting doesn't; extensions have their own telemetry; MS consider the data sufficiently anonymous to be exempt from GDPR, even the vscodium maintainers say they found it impossible to block vs code's eager telemetrizing.
I wonder what is the most effective way of actually completely stopping it phoning home, apart from yanking its ethernet cable?
The Rust and web related extensions work pretty well. I've been using that for a while.
There are a few more around.
I would like to see some analysis of the telemetry Jetbrains sends and whether you can turn it off. It's a paid pro IDE so it wouldn't surprise me if you can really disable it because some commercial customers will not want telemetry. In high security environments it's often considered bad due to the possibility of accidentally exposing secret information.
Use your operating system's firewall app to block all outbound connections from it.
I wasn't aware I could configure firewall rules for specific applications. Any hints where I could start looking at that?
Source for this? Does VSCodium with no extensions send telemetry, or is it the extension issue?
EDIT: Oh the quote from VSCodium is in the article:
Even though we do not pass the telemetry build flags (and go out of our way to cripple the baked-in telemetry), Microsoft will still track usage by default.
The funny thing is that I was about to install VSCodium for the first time TODAY, because I want live markdown preview. (I want a doc editor, not necessarily a code editor)
The rationale for VSCodium says it's about telemetry, although strictly speaking it's doesn't promise they removed all the telemetry!
but the product available for download (Visual Studio Code) is licensed under this not-FLOSS license and contains telemetry/tracking
-----
I should also say that on the dev side, crash reports are EXTREMELY useful and really do improve the quality of the software. But you used to have to confirm this data would be sent.
All this talk about how this data can be used to target people seems far fetched. Microsoft doesn’t have a serious ads business. What they do have are various products and services to sell to developers - Github and Azure among them. That’s the whole point of VSCode - give it away for free so they can understand developer practices and also upsell paid products. And if anyone doesn’t like that, they can simply … not pay Microsoft any money.
What if you added domains to host list
Yes, it’s always possible to turn it off but what I meant to say was that the VSCode team makes it hard.
Metrics and crash statistics are fine. And if they are collecting some dystopian level of data, then I bet that google and facebook has them beat already.
Telemetry is a major focus of windows since the current ceo too over. It’s kind of his thing
Note there is an open source version called vscodium that iirc has telemetry disabled but I found my extensions didn’t work with it last year when I tried it
I know it's not just me because I did find the github issue and there was no fix so I just stopped using VSCode
I had an issue around a year ago with cpptools on Linux, where a single 80k LOC header would cause obscene memory usage (10+ G) and the extension would basically stop working.