Telemetry in Front-End Tools
telemetry.timseverien.com
telemetry.timseverien.com
It is true that, in some cases, telemetry can make it faster and easier to improve software, but this is ultimately placing the convenience of the developers ahead of the privacy and autonomy of the users. I believe that software should serve the interests of users, and informed, affirmative consent for phoning home is an absolute baseline for respecting their wishes.
Is it though? Putting aside the issue of whether you trust the person collecting telemetry for a second. You can collect telemetry completely anonymized. You don't have store IPs or any sort of personal identifier. I know a certain crowd still won't like it but in practice how does that harm the privacy of the user?
You can't.
You can decide to discard identifying information (although only if you aggregate and don't insert some sort of "anonymous" identifier), sure. But you're still going to collect it.
So it all boils down to trusting not only the developer, but any company or interest that might acquire the developer or product in the future. And I think it has been repeatedly demonstrated that this trust is not well-placed.
Of course you can! Just route telemetry through Tor.
This is why I don't use web apps or web-based services unless I have no other option.
I'm surprised that anyone claiming to be security- or privacy-minded would prefer a native desktop app to a website running in a sandboxed web browser. Even if you build the code yourself, you're probably safer with the web app. At least in the browser, you can monitor and block any connections from the web app. Good luck doing that for a native binary without inspecting the source and build process for underhanded code exfiltrating data from your machine through some unknown number of obfuscation techniques.
Amusingly enough, the OP article is about telemetry of "frontend tooling," which does not refer to "web apps," but to "native binaries for building web apps."
That shouldn't be so surprising, really. It's a question of what threats you are the most concerned about. That decision is pretty individual.
> Good luck doing that for a native binary without inspecting the source [...]
I don't have to do all of that. I firewall off all outgoing traffic by default. If a binary is trying to exfiltrate data, it won't get past the firewall. That's hard to do in a web context.
> which does not refer to "web apps," but to "native binaries for building web apps."
Indeed so! I'm not saying that just using native binaries all by themselves is sufficient. I'm saying that I have more tools available to mitigate the problem when it's a native binary.
Just look at the impossibility of effectively stopping browser fingerprinting for an example of the difference between the two things.
At what level? Unless you're running every native binary on its own hardware, or maybe within a VM on an isolated VLAN, how can you be so confident that your firewalling method is less leaky than the battle-tested sandboxing of Chromium or WebKit?
Also, why not both? The most secure option might be running a web app in an isolated Chromium process, with a Chromium extension allowlisting outbound connections, and then also firewalling the Chromium process itself at the operating system level.
> how can you be so confident that your firewalling method is less leaky than the battle-tested sandboxing of Chromium or WebKit?
I can't be 100% certain, of course -- although that's equally true of Chromium and webkit. But I have a monitoring system that checks my firewall logs to catch anything suspicious.
> The most secure option might be running a web app in an isolated Chromium process
Because fingerprinting. I can't think of any way to stop fingerprinting. The best I can do (and I do) is to disallow JS from running, but that still leaves many possible signals.
This is even worse for native apps. webapps do fingerprinting because they don't have access to lowlevel native information. Any native App does not need to do Fingerprinting, they can just read device IDs directly, much more reliable that any method of fingerprinting.
Not if those apps can't phone home.
I in no way claim that my approach is airtight. It's purely a "best effort" sort of thing. But it's far better than nothing. I'm better off for it even if it doesn't stop everything.
this is a biased framing of the problem. first of all, the telemetry is often more useful for PMs (not developers) to make decisions on what efforts to prioritize vs not, and to assess relative success of launches. less of a convenience thing, but an industry standard way to get some meager signal from all your users, not just the ones that fill out surveys (and i'm sure theres plenty of complaints about surveys being unrepresentative too).
second of all, the work that developers and PMs do with this anonymized data is usually intended to help users.
pitting developer vs user is a false binary.
Which speaks volumes.
They do not understand, they just shrug “I guess devs need this”. They expect us to know, to care, because it is our job. We don’t and it’s sad.
That’s why things like public opinion surveys have to be random and why jury duty selection isn’t volunteer.
OK, then it places the developers and PMs convenience ahead of the users privacy and autonomy.
Sure, one can make the argument that if booking.com upsells rental cars more effectively they earn more money and provide a better service and experience for users. Haven't seen this happen once though.
I disagree with the practice, but it's definitely the way it is.
- which operations were running slowly on real user's computers (and how commonly they ran slowly), so that we could fix these;
- whether users were faced with out of memory errors, or timeouts, etc.;
- whether users had abnormally large database, preferences, etc. files, which suggested that we should optimize for such cases;
- in some cases, whether some features were used at all because if they weren't, it wasn't worth optimizing them.
I'd qualify these as definitely in the interest of the user.
I know that not everybody enjoys having Telemetry collected on their usage, but without this, Firefox would never have been able to catch up with Chrome on user-visible performance.
Terrible conclusion, by that reasoning a drunkard should only seek for his keys under the streetlight
If you do the former, there is a chance you may be missing out. You do the latter, you are certain that you are missing out.
I don’t want browsers or tools to undermine the trust of their users by quietly tracking unknown things. But I agree collecting usage data matters in ways that aren’t always appreciated.
I don’t know how to do it transparently and preserve/establish trust. But I think the instinct to distrust any metric collection probably isn’t the right balance to strike.
Not in the modern software ecosystem it's not. How often is telemetry used to make a product better vs. helping to add more dark patterns to increase "engagement"? No amount of telemetry seems to stop Microsoft from shoving more user hostile garbage into Windows 11.
If that is so, just ask and explain and be more specific than “this helps improve things”. How hard is it to say how exactly this vague ball of data is helping The Users?
The Next.js CLI phones home to Vercel with statistics about how many times e.g. `next dev` is being called, along with details of the developer's OS and CPU, as well as what Next.js plugins they're using in their project.
(And you can opt-out and they display a warning, tough of course that is indeed not an opt-in.)
If the next.js project is collecting ip addresses together with this info they are processing personal data under GDPR. They need to do so under one of the 6 bases for processing, which in their case is either consent or legitimate interest. If consent, opt-in is required and opt-out is a violation. If legitimate interest then opt-out is alright and in fact not even required, but they have a high bar for clearing that standard, especially since opt-out is offered (which somewhat disproves the claim of legitimate interest).
I assume the project is in non-compliance and one complaint to a regulatory authority away from a proceeding that may lead to a fine if they don’t switch to an opt-in model.
That said, I'm not sure if consent or legitimate interest are the only potentially applicable bases. Knowing when the software breaks so you can fix it seems like it might be in the data subject's interest. And if it's not PII (which I'm not sure it's not, given that an IP address can be exposed, even if not logged), those bases aren't even necessary.
This is a slippery slope for “understanding more about our users” and I challenge anyone to show a meaningful improvement that derived from telemetry of this kind. There are better ways to figure out how a project is used without harming privacy (ex: analyzers like BuiltWith). Projects doing this should get a red mark to discourage it.
As far as I'm aware, the contract of OSS is "you can fork this to change it, if you don't like it" (license permitting).
A big part of the reason people advocate for OSS is so they can read the source code to understand what's going on, and change the parts they don't like. Telemtry sits squarely under that usecase.
I agree that it's good manners for projects made users aware of telemetry, but there's no "contract" that implies they have to.
We do this (even then, strictly for non personaly identifiable crash/error reporting only), and I would encourage everyone who isn't a friend/family of the team, close community member, investor, etc. not to opt in. It isn't necessary. (Though having a small number of tens of people from those groups who do opt in is reasonably useful.)
Which features of our software are people using, and where do we need to concentrate the most with our limited resources and bandwidth to provide our users the best possible experience.
What are the biggest build- and run-time errors, and are they a result of the developer experience, and can we fix this through either better documentation or better error messaging.
What are the paths to successful adoption of our software, and can we nudge potential users (through documentation or developer outreach) to get on one of those paths.
We polled our users and they said that they want feature A, but our usage data shows that a vast majority of our users need feature B to be more successful faster.
Our software is used most often with other-tool-A. Can we strike up an agreement with the team that makes A, in order to give our users a better experience when integrating both tools through documentation, tighter API integrations, etc.
The best possible experience starts with not being spied on, especially against my consent. No feature is worth that.
> What are the biggest build- and run-time errors, and are they a result of the developer experience, and can we fix this through either better documentation or better error messaging.
Make it easy to report bugs. This is a solved problem (c.f. any popular open source application).
> What are the paths to successful adoption of our software, and can we nudge potential users (through documentation or developer outreach) to get on one of those paths.
As a user, I never want to be "nudged". Ever. Do not "outreach" to me especially if you've made the decision to contact me by spying on my data.
> We polled our users and they said that they want feature A, but our usage data shows that a vast majority of our users need feature B to be more successful faster.
So your users told you what they wanted and you didn't give it to them and justified the decision based on analytics. This is where people say that analytics are good for the product manager's job in helping justify its existence but at the expense of the users.
> Our software is used most often with other-tool-A. Can we strike up an agreement with the team that makes A, in order to give our users a better experience when integrating both tools through documentation, tighter API integrations, etc.
Create great solutions. If that means doing biz dev or integrations to make that happen, go for it. You do not need to spy on your users to build meaningful partnerships.
This kind of knee-jerk reaction says that there is no comment in the universe that would make you entertain a different perspective. That's fine, that's not what this comment is going to be.
Instead, I'll say this: the default assumption of ie "collecting crash data equates to invading my personal privacy" is just wildly off-base. I will freely admit that there are plenty of companies who have abused the concept of the always-on-network connection and are doing hardcore data-mining almost at the keylogger level. That's bad, 100%.
But there is a large majority of data-driven companies who are actively trying to make your experience with their software better, with the small team that they have. To say "well f them, they can try to divine how I, Power Linux User, would even use their software" is reductive and puts an engineering team at a distinct disadvantage and slows their overall output. This is not what you, as a user of their software, want, and it is ultimately a self-defeating argument.
"Create great solutions" does not happen in the year of our lord 2023 without data. That data needs to come from somewhere, and "make it easy to report bugs" isn't it.
You're correct in that there is no comment in the universe that would make me ok with being spied on. That's not a "knee-jerk reaction" though, it's privacy and consent 101.
> "Create great solutions" does not happen in the year of our lord 2023 without data.
It's true that there's a trend toward dark patterns and spying, which is exactly why we as users should resist and not allow it to be framed as the status quo. Great software has been created for decades, a lot of it still in use, without a hint of spyware or analytics.
When it comes to simple anonymized telemetry, one user or subset of users, spamming the same error over and over will skew your data, and unless you track them even more you won't be able to tell the difference. As I said earlier, it's a slippery slope to stand on.
Comprehensive telemetry/tracking is fine for say, an e-commerce website or any kind of cloud application where you're already receiving all the user input anyway. It's also fine, with consent, for desktop apps, especially complex ones like IDEs etc. It is not OK for third-party software packages that you'll ship to your own users - regardless of any promise to not add tracking to the resulting build, since at this point they already broke that wall and it only takes one "well-intentioned" PM to make it happen.
[1] a system that evolves automatically by running stochastic A/B tests is something many have dreamed up including myself. I'm really curious what that would end up looking like. That's being truly data-driven!
How can you make it easier to report a bug for the average user than providing a single button to the user that lets them upload a crash report after a crash?
But I offer two interesting and related observations.
The first is that it's not that uncommon for truly terrible feature decisions to be made on the basis of telemetry. Things like removing features that turn out to be really important even though they are rarely used, etc.
The other is that I see some people who are using software that they know is collecting telemetry alter their use of the software in an attempt to influence decisions based on that telemetry. Things like using critical features more frequently than they otherwise would, in the hopes that the feature won't be cut.
> Which features of our software are people using, and where do we need to concentrate the most with our limited resources and bandwidth to provide our users the best possible experience.
That a team of product managers and UXers can answer this question even half correctly with a bunch of metrics and A/B tests is one of the biggest fictions of our time (not to mention pure arrogance). Talk to your users, or better yet be serious users of your own product.
> What are the biggest build- and run-time errors, and are they a result of the developer experience, and can we fix this through either better documentation or better error messaging.
Invest enough in quality so that users see errors so rarely that they get in touch directly when they do.
> What are the paths to successful adoption of our software, and can we nudge potential users (through documentation or developer outreach) to get on one of those paths.
This is an anti-pattern. Your adoption is irrelevant, only the utility of your software and user happiness matter. If those things aren't compatible with you making money, you shouldn't make money. Using telemetry for this purpose decimates the already weak value of telemetry for improving your software.
If you are building applications with telemetry for this purpose, please stop.
> We polled our users and they said that they want feature A, but our usage data shows that a vast majority of our users need feature B to be more successful faster.
This is the height of arrogance. Quite incredible that anyone would think like this. I would immediately end the relationship with any provider who believes that their usage data, of all things, is enough to reliably decide that the users don't know what they want and need and the provider knows better.
> Our software is used most often with other-tool-A. Can we strike up an agreement with the team that makes A, in order to give our users a better experience when integrating both tools through documentation, tighter API integrations, etc.
The level of telemetry required to know not only how I use your software but what other tools I'm also using is quite terrifying. Pretty much total surveillance. Try talking to some users instead.
"If I had asked people what they wanted, they would have said faster horses." Henry Ford
In a world where we have companies like Google/Facebook and massive levels of surveillance it's bizarre to me that this is the hill some of you want to die on.
The tools themselves provide a way to disable it, so I don't see what any of the fuss is about.
This happened recently with a telemetry proposal in the Go community and everyone threw their toys out of the pram. The Go team was forced to back peddle and now the telemetry collected by the toolchain is worse off because of this.
Any connection that is not neccesary for the purpose of the tool is exposing the IP of the user without his consent.
Pinky promises that IP of the connection will not be (ab)used do not count.
See https://rewis.io/urteile/urteil/lhm-20-01-2022-3-o-1749320/
> It is sufficient that the defendant has the abstract possibility of identifying the persons behind the IP address. Whether the defendant or X. has the specific opportunity to link the IP address to the plaintiff is irrelevant.
What about server logs? If I send an http request to a server, do I have any right to say they can't log that request? What about the database metrics?
Yes, absolutely.
That website's code is being executing on my machine, consuming my power, then using my internet service to phone home with a data package of arbitrary size.
It's very difficult to tell what data is really being reported by an application. The majority of the time, all we have to go by is what the developer says is being collected.
But the developer claims cannot be considered trustworthy by default. Many developers have simply lied in the past about what they collect. And even if the developer is truthful, they can only speak for what the software is doing at the moment. There's no guarantee that the developer won't change their mind in the future, or the product won't be sold to another company that isn't so considerate of their users.
Also though, interestingly Walmart actually really does do the equivalent of this: https://bernardmarr.com/walmart-big-data-analytics-at-the-wo.... They have real-time detailed metrics for individual customer transaction behaviour in stores, and I would not be surprised at all if they tracked lots more, e.g. the total number of cars in parking lots, how long different cohorts of customers spend in stores, etc.
And the problem is even bigger, how do i know what data is being collected? even if it declared publicly somewhere, do i need to check on each update if the collected data has changed?
> otherwise they'd be begging for a GDPR problem
They have in fact a GDPR problem but like in the case above where it was not Google to answer for the GDPR Problem but the site that used Google CDN, in this case too, most probably it will not be these toolmakers who are taken to court but the companies that asking their employees to use these tools without disclosing this personal IP leak. If this becomes the norm how much burdon do we have to go through to verify each tool whether we need to ammend the employee GDPR consents or not? I see this already having an effect... Just standardise only on Microsoft tools so at least we need to gather GDPR consent only for Exposing data to Microsoft...
Data about me, my machines, or my use of my machines is private data. Nobody has any right to it besides me. If it's being collected without my informed consent, that's just spying.
Personally, I like to implement a “Report a bug” button which sends detailed telemetry when it really matters. To keep an eye on more general usage data, I use Cloudflare Web Analytics. It’s easy to build a toggle for analytics. Personally, I find that more than sufficient, especially when combined with traditional user studies.
Having the backend respect the toggle too wouldn’t be that hard.
To do anything less than blocking every connection until you've allowed it is putting a lot of faith in software (vendors) that they have already demonstrated they will abuse. It's nice that some software grants you the privilege of asking it politely to not do a telemetry[1], I guess. But it (and the vendor) already demonstrated they are happy to breach your trust (and, arguably, your security) for their own arbitrary reasons.
So many of these tools present themselves as free software. But if you're not in control of it, it's not free.
[0] Unfortunately, it's really nice when you control the operating system, productivity software suite, game development studios & publishers, email, world's best and most helpful good boy with your best interests at heart online chat assistant, scalable web services provider, programming languages, software forge, proprietary 3D graphics stack, laptops, game consoles, ...
[1] What does `telemetry = no` even mean anyway? Does the tool report that telemetry was disabled? Does the tool now do nothing on a network that isn't explicitly declared, or what is required to perform the actual work you ran it to do?
I know that this sort of data is very useful to a business. But I also know that it's abusive and wrong to extract this data without my customers agreeing to it.
While this data is very valuable, not having it isn't an existential issue for any company (or, if it is, that's because there's something very wrong with how the company operates.)
> The truth is to be found in the middle
Hard disagree. This is an issue of consent, and I don't see how there's a middle ground for consent.
The fact that you may have incentives to do bad things is just a reason society should make laws against you doing them, not a reason I should feel sorry for you and let you spy on me.
> But then no one opts in!
OK that is your business' problem to solve.
I never liked Postman and that's how I mostly stopped using it. Completely stops working on analytics request denial and then starts sending reports to Sentry (but even if you allow those, they don't learn, I've tried)
It's also pretty interesting to see the comments here, in a privacy focused thread, and contrast them with comments elsewhere about software performance. If you like faster software, developers need to measure it.
Those categories feel a little too broad to me, especially without clearly stated definitions. "Environment" in particular made me second guess if dev tools were leaking secrets. But, thankfully, the projects I looked at were very clear that they don't collect anything from the shell environments.
Maybe all it really needs is some detail for what's in the category. "Device (OS, core count, available memory)" or "Usage (command, plugins, timing)"?
docker ce - collects telemetry in the installer before you have even installed anything
brew - telemetry on by default
balena etcher - telemetry
anything apple
Holy shit! A glorified dd if=/of= that uses Electron and has telemetry – now that's something
If you are using a free - as in free beer - and not free as in free software - you are the product!
True. And increasingly, paying for the software doesn't make this any better.
Or another example, if architects decided to start putting cameras in the house everywhere. The camera, somehow, blocks out your face but monitors everything else. "But, we need this data to figure out which part of the house you're using and how often you're in these rooms. PINKY promise it's anonymous and no one will know who you are."
Opt-in consent is the only right way.
I'll start with the caveat that I'm not a lawyer, but I wouldn't be so sure. For example, knowing when your tool breaks for a user so you can fix it sounds like something that you could argue is necessary to make sure the tool works for the user.
Does submitting an order form for some goods in any way depend on the telemetry requests you sent before/after? Would the order still go through if those requests were missing or fed fake data?
If the answer to the above is true then "strictly necessary" would be difficult to apply. Of course the point is moot because so far GDPR enforcement is not only lacking to begin with but the cases that are investigated aren't given the right technical expertise to adequately determine those answers.
This example requires a finer-grained analysis than the broadly chosen "opt-in" and "opt-out" categories to describe telemetry collection, and I wish for this information initiative to improve in clarity and thoroughness (e.g. does VSCode also scour the local network?).
Addendum: Thanks for asking for clarification. The grandparent comment could've been ambiguous in intent.
The vscode source is not everything
What command line switch I prefer seems less problematic than say my browsing habits. One hell of a slippery slope I know, but it does seem to me like a spectrum
This is not true in Europe Anymore. Any Optout telemetry is illegal because they have to establish a connection to the telemetry server and thus leaking the user IP (which is PII for GDPR) without the user's consent.
See the latest case that even linking a font from CDNs is illegal. https://rewis.io/urteile/urteil/lhm-20-01-2022-3-o-1749320/
relevant part: > It is sufficient that the defendant has the abstract possibility of identifying the persons behind the IP address. Whether the defendant or X. has the specific opportunity to link the IP address to the plaintiff is irrelevant.
Any connection that is not neccessary for the primary purpose of the tool/app needs to have the consent of the user thus all opt-out telemetry is illegal because the tool/app can execute its purpose just fine without the telemetry
Note that this is not a problem only of the toolmaker, this is a problem also for the employers that are exposing the developers to these tools creating unneccesary burdons for them to gather the consent of their employees.