I hope this project motivates MS to open source the runtime and making tracking a user- selected option.
I really like Code and think MS has really helped the dev community by making it so great and free. But I would like to see them embrace f/l/oss for all their non-core products and stay on track for customer/dev-friendliness.
Very few users change the defaults. To the point where you don't get enough data for useful insights.
This is literally the only reason that telemetry is almost always on by default. There is no illuminati secretly buying developers to learn your secrets via application telemetry, and it's laughable to think that's the case, in my mind.
Tl;dr; people usually accept the default so your prompts yield very different results.
* How many times you clicked Edit -> Paste vs Ctrl/Cmd + V. This is not personal information. It cannot be used in any way to identify you.
* Your location/medical information. Guard this with your life if you must. Share it with literally no one. If it's uploaded somewhere, delete it asap.
All our information falls on a spectrum between these two extremes. Let's please acknowledge that not all information is super sensitive, that it's ok to share information like the paste example.
Second, this information can be useful, and can change the direction of the products we're building. For example, the Office team wanted to re-design the old menu bar interface to surface the features people were using the most. You know what people used the menu bar for the most? EDIT -> PASTE. You'd think everyone knows about Ctrl + V, but apparently not. The vast majority of the world used to click Edit, then click Paste. Knowing that users were doing this allowed the interface to be moulded to suit the needs of the silent, non shortcut using majority instead of HN power users - https://imgur.com/a/WLm4UJd.
Third, if you acknowledge that this information is useful, then it follows that it's only useful when it's opt-out. If it's opt-in and only 0.01% of your users opt-in, you can't make any reasonable conclusion from the data because it wouldn't be representative. If you tried to convince folks that no one uses keyboard shortcuts to paste based on this tiny data set, they'd laugh you out of the room. Collecting opt-in data in this case is almost useless.
Fourth, products that collect this info become better. If you're ok with webapps collecting this but not native apps (like VS Code), then the consequence might be that web apps end up being much better than native apps. Whether you prefer that or not is moot, but personally I like having options. Let's not force native app makers to stop collecting non-identifying information when the downside is minimal/non-existent.
Search engines are recording really detailed information as you type, as you browse other websites using their invasive trojan-like analytics. With this data you can probably identify anonymous users from their behaviour (even if the ipv4 address is "encrypted" hah). All it takes is for data from one company to end up at another company and I believe this happens very often.
I think the problem is that you need less and less information to personally identify someone as you have more data. We are not intelligent enough to know what data is identifying or not so therefore the only option we have is to stop giving out our data.
Tell me, who exactly is running ML on this data? They're going to great lengths to anonymize this data. If they wanted identifying information, they'd just upload your unixname, they wouldn't anonymize and then de-anonymize with ML.
I don't even know how you conflated the metrics on menu clicks with uploading paste data. That's ridiculous. Did you even notice you made that leap?
I don't know what data they collect now or what they will collect later so I am just speculating. But I am sure the ToS has a section about how they can change the terms however they wish.
Please tell me about the great lengths they go to anonymize this data because I believe it to be very difficult to do, and absolutely not in their interest.
Edit: With pasting habits I meant if you paste with shortcuts or go through the menu, not the actual paste content. Sorry for the confusion.
So what? VScode can also be used as web app already, it's not about what others do or don't with their online services. The question is about the expectation and respect to the end user, having telemetry on by default is disingenuous to say the least.
One single link to your identity, and you now have user tracking. If course, there's probably easier ways to achieve it.
If you don't like that switch to a non-open source editor.
It’s not like it’s human sifting through it by hand.
As grandparent OP said, in recent years this has moved waaaaay beyond theory territory and been shown time and time again that <Corporate Sector> + NSA + FBI + CIA + intelligence agencies around the world all employ different ways of collecting analytical data broadcast over the internet.
NSA just taps the servers without telling anyone.
FBI sends National Security Letters containing gag orders preventing companies from telling you that the federal government is now a data sharing partner.
CIA just pays companies for it.
FISA court issues secret rulings justifying the legality of it all.
Whether that bothers you or not is up to you. Most people don't care. I usually don't. I wouldn't say "logic dictates this won't happen" when it is pretty much only the multi-national multi-billion dollar companies subject to this kind of tampering, and most incentivized to monetize the analytics by allowing these and unknown third parties in.
To avoid the security exhaustion, some people would simply prefer their text editor not be "smart", which is a euphemism for internet connected.
I recall a recent thread here on HN where I learned that the Windows Calculator was sending data ("telemetry") back to the mothership.
Nowadays I think I'd be more surprised if a Microsoft product were released that didn't phone home.
I mean, sure, it might be the case that there were indeed tens of thousands of ALSA users with tracking turned off, but... From my perspective, it seems more likely that it was just a handful, and really there's no way to tell the difference. If you turn off telemetry, and be aware of and accept the downsides.
I trust Mozilla, so my telemetry is on, but for many other applications, I often opt to turn off tracking - with the understanding that it's harder for them to tell what my needs are.
No, that still falls into fallacious usage.
The user doesn't justify why the steps of their assertion follow after the other. Just that they... do.
Over the course of the last 20 years, we've seen that once a data collection and digital surveillance framework is put in place, the surveillance tends to expand.
Slippery slope arguments, sans good reasoning, tend to be fallacies. However, don't fall into the trap of thinking that an argument backed by historical record is a slippery slope just because it's predicting an outcome. We might call that the "history is all slippery slopes" fallacy. Stating "this has happened before multiple times before, and each time has lead to x" is a very different argument to stating "this has happened, so the logical extrapolation is x".
For example, I slept in 'til 8:30 today. OMG, a slippery slope. Next thing you know, I'll be sleeping until 3PM. Til midnight! I will never wake up again. But as it happens, sleeping in isn't a slippery slope. I don't think there's any solid evidence that telemetry is either.
“... and that was me before we had children.”
By itself no, but it is often used fallaciously. Such as in this case, when someone is opposing some good thing on the basis that that good thing might, some day, lead to bad things.
Our entire legal system is predicated on common law precedent. So it is very valid in many cases to argue that allowing something good now, might set us up for something very bad later.
You think people elect for bad things? Bad governments, bad software or privacy violations? They chose things that are full of rhetoric and promises of good things then those bad things get snuck in off the back off relaxed regulation, or existing software adoption, etc.
Take Facebook as an example, people didn’t sign up to it thinking “I wanted to be tracked around the internet so I can have personalised adverts” nor dis Zuckerburg think “Wouldn’t it be good to create a platform that could latter be used for rigging elections”. No, instead we got there because of a serious of good ideas that slowly got abused.
There is a saying that goes “The path to hell is paved by good intentions.” I’m not a religious man but I think that beautifully illustrates how slippery slopes are not a logical fallacy.
The "page" in settings consists of just two options, and the complete descriptions of the types of information they collect are "crash reports" and "usage data and errors". That seems the opposite of transparent and granular. Am I missing something?
There is an application in the Microsoft store you can install that lets you view all telemetry collected by Microsoft, and the actual data is encrypted. The metadata (which tells you WHAT is being collected) is not. Knowing the people I have worked with in the past, this is to securely prevent modification by users before the data is actually sent. I've worked with lots of people who would modify that data to attempt to get a feature added that they wanted or just to screw with MS.
That tool also gives you the option to delete all telemetry sent from that machine in Microsoft's possession.
Setting that to "Trace" will show the telemetry being sent in the Output pane of the interface, amongst a bunch of other stuff, I am sure.
"Logic and critical thinking textbooks typically discuss slippery slope arguments as a form of fallacy but usually acknowledge that "slippery slope arguments can be good ones if the slope is real—that is, if there is good evidence that the consequences of the initial action are highly likely to occur. The strength of the argument depends on two factors. The first is the strength of each link in the causal chain; the argument cannot be stronger than its weakest link. The second is the number of links; the more links there are, the more likely it is that other factors could alter the consequences.""
https://en.wikipedia.org/wiki/Slippery_slope#Non-fallacious_...
I can also see them collecting code searches done within the app as a way to check if their search system is working well for real use-cases.
Neither is outside the realm of possibility - you just have to put yourself in the mindset of a dev who is assigned to track down a rare crash or to “improve the search experience” who might want a little more data to work with.
Not saying I agree with any of this collection - it’s terrible and definitely falls under “the road to hell is paved with good intentions”. Companies should be extremely clear about what they will and won’t collect - and never cross the line even if it would be useful.
> Lol, probably just an oversight because they made the search a lot better.
> So I think it only does this on the settings file and not on other files ;)
> https://code.visualstudio.com/blogs/2018/04/25/bing-settings...
"workbench.settings.enableNaturalLanguageSearch": false
That's unacceptable.
>> it only does this on the (VSCode) settings file and not on other files ;)
The data could be accidentally broadcast or left in a vulnerable place. Even the payroll data for the national security establishment was once reported to be compromised. Everything is vulnerable. Computer science is in such an abysmal state.
Even with properly configured servers, OSes, and databases that are up to date, they are still vulnerable to zero-day attacks because they are not formally verified and have enormous and largely unnecessary complexity. Then throw in the crazy complexity of processor instruction sets, creative side-channel attacks, and stuff which exploits the physical properties of the hardware (rowhammer).
It is reasons like these why we should never really trust transmission of sensitive data over the internet. The concept of secure voting systems, for instance, is literally a joke. insert obligatory xkcd here
There really isn't such a thing, in the way you've used that phrase.
Capturing telemetry on how I use a tool from within that tool is perfectly fine, to me. Collecting telemetry on my search history in the browser by that same tool isn't. THERE ARE NO INTERIM STEPS that makes the second of those ok. There is no slope. If there is, it isn't slippery. There is a series of discreet decisions and at some point (which is different for everyone) a line is crossed. There was no slope or slip that brought you there, only a series of mostly unrelated decisions.
To think that Microsoft's long-term goal is to install a keystroke logger via a multi-decade and multi-phase plan that begins with application usage telemetry in a free developer tool thanks to "a slippery slope" is just simply not realistic.
Telemetry is used either with naivety or malice. There is always some risk to the user.
That being said, it’s somewhat amazing we are now in a time when a Microsoft product could have aspects users don’t approve of and rebuild it without it: everyone wins.
I can think of two reasons for the data to be made public :
- for trust and transparency purposes : It is for the same reason than an election system should be observable and reproducible, from data collection to final decision, including counting methods, etc. Otherwise it would be like "trust us, the data shows it" and you don't show the data to anyone to prove it.
- for coordination and sharing knowledge : some people might interpret the data differently, chose to focus on a niche market by doing different bets than MS, and create a complementary editor to the one from MS. MS has no obligation to support minorities, but someone else might be interested, and those minorities are detectable in the data
Not only, that's the main problem. And I can't simply trust it for company without dark M$ reputation and EEE experience.
I hope this is sarcasm. If not, what did I miss? This is the same empty phrase that facebook, google, etc. use. Why is Microsoft more trustworthy in that regard? I am 100% sure they use the data to make money in short terms or in a long run. They for sure use it to make VSCode better, but only to get more people use VSCode and make them dependent on it. VSCode is a prime example of the Embrace, Extend and Extinguish strategy. I already see them grasping for the Python community.
A need for telemetry indicates that contributing is too difficult.