Fedora considers “privacy-preserving” telemetry
lwn.net
lwn.net
Looking at the current developments in AI, I am concerned that AI models can easily de-anonymize and guess end point users when being fed with "telemetry data" of hundreds of thousands clients.
I hear a lot and read a lot of software and hardware vendors saying that "telemetry" is supposed to somehow magically improve the user experience in the long run, but in actuality software tends to get worse, unstable and less useful.
So, I would like to know how exactly any telemetry data from Fedora Linux clients is going to help them, or how is it going to improve anything.
As seen in Firefox..
It also doesn't really make strategic sense to focus on the lowest common denominator. Chrome already has that group. The one place they could eke out a loyal userbase is specifically the users that Chrome fails to capture because they have unusual needs or requirements.
You have to be just mindbogglingly oblivious to not see how this has been one of their biggest problems the last 20 years.
Would you rather attrite a fraction of 95% of your users or a fraction of 5% of your users.
Without data you can invest your time into the wrong features.
Meanwhile, Brave's vertical tabs were done by a single developer.
> We also want to know how frequently panels in gnome-control-center are visited to determine which panels could be consolidated or removed, because there are other settings we want to add, but our usability research indicates that the current high quantity of settings panels already makes it difficult for users to find commonly-used settings.
Personally I'd like to see more transparency in their usability research, because GNOME is best know for removing features, which is what they'd like to do this time around as well.
[1]: https://lwn.net/ml/fedora-devel/CAJqbrbeOZrHvYjvMCc=qGZD_VXB...
Apple also does a lot of (non-anonymous) user testing, which can give very detailed feedback.
That is, if you suspect they'd change their minds and start trying to deanonymise previously collected data anyway — remember that open source distributions (I don't know fedora-the-organisation specifically) are generally made up of volunteers like you and me. Notable exceptions obviously exist, like for-profit Canonical; that's not the org type I mean or trust.
If we can't trust people ever, what's the point in doing anything?
You need police and a judicial system and you fix them whenever they break. But you don't need telemetry, it's entirely optional and shoddy implementations translate into unnecessary risk.
Also:
https://lwn.net/ml/fedora-devel/H5JEXR.LLU011IQ4I6K@redhat.c...
I'm no fan of privacy invasion, always bother clicking through the banner to find the reject cookies option, and sent out plenty of GDPR requests whereas almost nobody else I know ever sent one. I'm not in favor of tracking, but collecting anonymous statistics, especially when they open with "privacy-preserving" and the business is not Facebook or Google or so where we know there's shit about to hit the fan, the cynicism and mistrust in this thread baffles me. Nobody minds when they browse the web and every site keeps access logs invisibly, but oh boy if someone announces keeping a visitor counter for a configuration screen to see if people can find their way to it
I am not in Europe. The GDPR doesn't help me.
> collecting anonymous statistics, especially when they open with "privacy-preserving"
I am far from convinced that such statistics are gathered in a "privacy preserving" way, but that's neither here nor there.
> the business is not Facebook or Google or so where we know there's shit about to hit the fan
The problem is that you can't just trust the current devs. You also have to trust all future devs and companies that may buy the thing. It's not Facebook or Google now, but it could be in the future. And this is Fedora, which is connected to Red Hat, which is connected to IBM.
And it's also not just about privacy. It's also about impact on product development. It's not exactly rare that software has been made much worse as a result of decision-making based on telemetry data.
> Nobody minds when they browse the web and every site keeps access logs invisibly
No? I think quite a lot of people mind this. But there's nothing that can be done about that. It's still worth trying to keep everything from getting even worse, though.
- It does not allow a reasonable decision as it does not show the data before it sends it
Which is a very large "if". There is a tendency for people to think that the only personal data consists of what legally counts as PII, when in fact there is much more personally identifying information than is covered in those definitions.
It says nothing about whether or not you can join the output of multiple ML models with other telemetry to build a deanonymization model.
I can almost guarantee you that the US government has a tool where you can input a few posts from a person on an anonymous network and get back all of their public profiles elsewhere. Fingerprinting tools beat all forms of VPNs and the like. Our privacy and anonymity died like maybe two years ago, there is no stopping it.
And telemetry is important. We have limited resources. How do we determine the number of users impacted by a bug or security vulnerability? Do we have a bug in our updater or localization? Are we maintaining code paths that aren't actually used? Telemetry doesn't magically improve user experience, but I'd rather make decisions based on real data rather than based on the squeakiest wheel in the bug tracker.
We can certainly make flawed decisions based on data, but I'd argue that we're more likely to make flawed decisions with no data.
What I've seen in practice so far is that the use of telemetry has harmed software quality more than helped. It often leads developers to optimize for the wrong things and make poor design decisions. This happens because they tend to think that "the data never lies", ignoring the fact that telemetry always gives a skewed and incomplete picture.
I’ve been a product manager for products that had no telemetry, and that can be a rather undesirable place to operate, especially if you’re in the enterprise space where product changes can impact the operations of businesses.
I think it’s certainly possible to focus on the wrong things, but I don’t see that as an outcome of telemetry itself as much as an outcome of a product team that doesn’t understand the problem space or customer base.
The attributes to capture are presumably based on what teams understand to be key indicators about their app/service. I think confident incorrectness armed with bad data is just a slightly different version of a complete lack of data. Such a team was operating on whatever they imagined to be important before, and they continue to do so after, albeit with greater conviction.
But good telemetry in the hands of a good product team can be immensely beneficial for decision making and can protect customers from bad decisions. Anecdotally, my ability to pull numbers about certain attributes has been key to my ability to shut down executive pressure to make changes that would have drastically impacted customers if not for the direct evidence that it would.
I’m also not claiming that downsides don’t exist, and privacy is always my primary concern, but there are a range of outcomes based on the maturity of a team/company, and as long as the PM understands that data is not an alternative to having a relationship with customers, I think data is pretty important.
Honestly, I can't actually remember specific examples. It's not something I dwell on. But I know that's it's happened several times that software has been made much less useful to me because features have been removed on the basis of being rarely used, ignoring the fact that even though they're rarely needed, when they are needed, they're indispensible.
More often, though, the bad telemetry-based decisions I've seen are around UI changes. Things like a laser-focus on reducing the number of clicks it takes to perform things, even though sometimes reducing the number of clicks for a thing adversely impacts the usability of it.
For bad telemetry-based UI decisions, my standout example if Firefox, although that's hardly the only one.
> But good telemetry in the hands of a good product team can be immensely beneficial for decision making and can protect customers from bad decisions.
This was actually my point of view a few years back, when telemetry started to become popular. And, as a dev who sells software commercially, I totally understand the value on that side. My experience with products that have used it, though, has shifted my view.
All that said, I do agree that it's possible to use telemetry in a way that is good for users. But I don't think it's common, and I think the reason for that is economics and human nature.
Once you start measuring a thing, that tends to become a goal rather than just a data point. And since the industry is all about maximizing velocity, that effect is even stronger. Doing proper usability studies is a slow and expensive process. Telemetry can be a useful thing as part of that process, but the tendency is to make it pretty much the entire process. That does a disservice to everybody.
> the PM understands that data is not an alternative to having a relationship with customers, I think data is pretty important.
Not just the PM. The entire company. But a relationship should be consensual, not forced. I have zero issues with opt-in telemetry. When it's not opt-in, though, it's an invasion and adversarial. I presume that's not the sort of relationship a good PM wants.
> But a relationship should be consensual, not forced. I have zero issues with opt-in telemetry. When it's not opt-in, though, it's an invasion and adversarial.
That's it. No matter how many ways you dice it - collecting data without consent or forcing opt-out is an invasion. As more and more of our lives shift to being online, our privacy and our sense of autonomy in a digital world is ever increasingly paramount.
Ideally, yes. In practice, and especially in larger shops, it’s the PM’s job to own this relationship and to make sure the important players have this understanding.
> But a relationship should be consensual, not forced.
Absolutely agree here.
A large software company in Redmond, perhaps?
Why are there code paths nobody uses, that have to be maintained, in the first place?
But how do we know the code path is no longer in use? Are people still using this CSS property (e.g., a vendor prefix)? Are people still using gopher or this one configuration variable? The more configuration options you have, the more combinations you need to test and maintain.
It can be done with error reporters like the "System program problem detected. Do you want to report the problem now?" popup in Ubuntu. In my experience, many users are willing to send error reports, and they're extremely useful, although 90% of reports are garbage.
Do you work for Boeing or something?
When I've worked on mission critical (so, safety critical, in practice), we made sure the probability of catching a failure in testing was 100x the chance of catching it in production.
Modern software development techniques like fault injection and fuzzing make this pretty easy to achieve.
We use de-identified voluntary safety reports filed by pilots, air traffic controllers, and others, along with flight telemetry data from the aircraft and other data to identify and study potential safety issues in the national airspace. Privacy-preserving techniques ensure that we can collaborate on safety and trust that the data stays non-attributional (and thus, non-punitive since participation is voluntary) despite competing interests.
We can't really do fault injection or fuzzing for real-world systems to understand, say, the impact of false low altitude alerts on risk of undesired aircraft states (e.g., controlled flight into terrain) at a certain airport.
Data doesn't always help! It can lead to assumptions, often really bad ones. And IBM isn't to be trusted with it.
You don't need AI for this. This is done by real humans right now, using data points correlated from multiple sources.
That's time consuming and expensive without ai, so you can't do it at scale to a comprehensive degree. That hasn't been practical until now. It still isn't quite cost effective to do this for every human, everywhere, but soon it will be. Give it 5-10 years
Thanks to ai
User ANON-123 with default font x and locale y and screen resolution z installed package x1
Is clearly a big hazard, but statistics on what fonts, locales, and resolutions have is not really. Even combinations to answer questions like "what screen resolutions and fonts are most used in $locale?" should be safe as long as the entropy is kept low. It is less useful, since you have to decide on your queries a priori rather than being able to do arbitrary queries on historical data, but ethics and safety > convenienceIt doesn't take very many bits of information to deanonymize someone once you start combining databases.
It won't improve anything for users. It might improve something for IBM.
Some metrics like startup time and crash counts lead to clear improvement, while others like pointer heatmaps and even more invasive focus tracking are highly dubious in my opinion.
On a related note, I’m coming to the opinion that A/B testing is harder to pull off than many think. And serving a single user both A and B at any point can confuse them and get in the way of their trusting the consistency of the software. Much like how when you search for something twice and get different results in Apple Maps. OK, now I’m just ranting…
I absolutely detest that Catch-22 argument, which some distro (not Fedora) actually tried to use on me in the past.
It's a black box with no incentive to participate unless you're one of those specific types of users that is dedicated enough to put up with all of that, or users that have never done it before and are trying hard to contribute back.
I figure it must be due to an abdication of responsibility-- absent information, the product must at least appeal to someone working on it who is making decisions about what is good and what isn't, and so it will also appeal to people who share their preferences. But with the power of DATA we can design products for the 'average user' which can be a product that appeals to no single person at all!
Imagine that you were making shirts. To try to appeal to the most number of people, you make a shirt sized for the average person. But if the distribution of sizes is multimodal or skewed the mean may be a size that fits few or even absolutely no one. You would have done better picking a random person from the factory and making shirts that fit them.
When your problem has many dimensions like system functionality, the number of ways you can target an average but then fit no one as a result increases exponentially.
Pre-corporatized open source usually worked like fitting the random factory worker: developers made software that worked for them. It might not be great for everyone, but it was great for people with similar preferences. If it didn't fit you well you could use a different piece of software.
In corportized open source huge amounts of funding goes into particular solutions, they end up tightly integrated. Support for alternatives are defunded (or just eclipsed by better funded but soulless rivals). You might not want to use gnome, but if you use KDE, you may find fedora's display subsystem crashes out any time you let your monitor go to sleep or may find yourself unable to configure your network interfaces (to cite some real examples of problems my friends of experienced)-- you end up stuck spending your life essentially creating your own distribution, rather than saving the time that you hoped to save by running one made by someone else.
Of course, people doing product design aren't idiots and some places make an effort to capture multimodality though things like targeting "personas"-- which are inevitably stereotyped, patronizing, and overly simplified (like assuming a teacher can't learn to use a command prompt or a bug tracker). Or through efforts like user studies but these are almost always done with very unrepresentative users, people with nothing better to do then get paid $50 to try someting out, and you learn only about the experience of people with no experience and no real commitment or purpose to their usage (driving you to make an obscenely dumbed down product). ... or by things like telemetry, which even at their best will fail to capture things like "I might not use the feature often, but it's a huge deal in the rare events I need it." or get distorted by the preferences of day-0 users, some large percentage of which will decide the whole thing isn't for them no matter what you do.
So why would non-idiots do things that don't have good results? As sibling posts note, people are responding to the incentives in their organizations which favor a lot of wheel spinning on stuff that produces interesting reports. People wisely apply their efforts towards their incentives-- their definition of a good result doesn't need to have much relation to any external definition of good.
Now, with telemetry, they can say quantifiable things like "we've driven catastrophic root filesystem loss and permanent loss of network connectivity to 0% of installs!", and prioritize any contrary bug reports away in a data-driven, quantifiable way.
(Because, of course, weak telemetry signals are more valuable than actual humans taking the time to give you feedback on your product.)
As coined here (copy and paste the link into your browser if you don't want a 'surprise' from jwz) ttps://www.jwz.org/doc/cadt.html
Now how long until someone reinvents peer to peer networking or document databases again…
I created a bug report [1] for tigervnc-server in Fedora because the Fedora documentation [2] for setting up a VNC server didn't match any more what was coming from dnf.
In the bug report I provided the info that would need to be fixed in the documentation. Now after two months, seemingly nothing has been done to fix the situation.
[1] https://bugzilla.redhat.com/show_bug.cgi?id=2193384
[2] https://docs.fedoraproject.org/en-US/fedora/latest/system-ad...
> Optional telemetry, of course, but again - this creates a selective and unpredictable reality. Ordinary people don't care either way, and nerds will always make deliberate choices that often have nothing to do with product or profit or anything else entirely.
So, by that logic, users who opt out of telemetry are aware of why they are doing it. People who don't care share their chaotic usage patterns. This creates a false picture of usage reality, and makes software worse. In conclusion, there are only two remaining choices: 1. Make telemetry non-optional 2. Ditch telemetry and rely on QA studies
But then users who care about this issue will just block the application's ability to phone home, or use a different one that doesn't spy (my definition of spying is any time data about me or my machines is collected without my informed consent).
Ambiguous: they're aware of why who is doing it? I opt out of telemetry because I don't know why they're doing it - the data collectors. I mean, I know why they say they're doing it, but I don't know if it's true.
I also don't want my computing resources and network bandwidth used to further a goal that I might not support. Even if the only reason for collecting data really is to "improve the product", perhaps that'll result in them making the product dependent on systemd, which from my POV would be an adverse outcome.
And if you're worried that your user case (e.g. not-systemd) will be deprecated then that's a good reason to keep it on and be represented. I bet people who used Firefox's RSS feature were more likely than usual to turn off telemetry, yet that didn't save their feature. If anything, it may have made Mozilla more confident in removing it
That phrase can be viewed in many different ways, depending upon who the decision makers setting the goals are.
> If anything, it may have made Mozilla more confident in removing it ...
Of course. People know this, that's a good part of why people are against metrics.
It leads to dumbed down software, as there's always a 5% of features that can be cut (on every budget or dev iteration) to "focus more on the majority uses!".
Even if that's true, organizations like to keep this stuff for longer than they should and are not immune to hacks or changing business priorities (like AI). If there was serious liability attached to leaking or "off-label" use of this data then I'd feel better about it.
Also, using telemetry to justify removing features is not about making the product better but about cutting costs. How is Firefox better without RSS support? The fact that only a small number of users use it tells you nothing about how passionate or influential those users are, nor how such a change will affect wider perception of your products. See also: how pissed off many on HN still are about the death of Google Reader even though it was a "niche" product.
Maintaining stuff has a cost in time and/or money. Added new features or fixing bugs has a cost in time and/or money. Time and/or money is limited, so decisions must be made on what is the best use of it. Removing features can (sometimes) result in a lot of savings in the form of simplified codebase (which means improved developer productivity), reduces surface area for bugs, etc. I (really) hate when products remove features, and I think it's often done assuming a lot more savings than actually happens, but there is some logic to it.
I don't know what Firefox used the RSS support savintgs to do, but if they used it to build containers (which I use constantly) then for me Firefox is much better. If they squandered that time, then of course it's not better off. But without having access to roadmaps and planning documents, it's dificult to know which tradeoffs were made.
This is the key actually.
they could be a single engineer tracking down a pernicious segfault.
or
they could be marketing "they gave us the data, we can do anything we want"
It seems to me that even with excessive levels of telemetry, software remains buggy and sluggish most of the time.
Restaurant management wanted to compare different soup offerings by counting orders from each soup to determine which ones were more popular. They have selected the most popular two offerings, the rest were scrapped in order to safe money on ingredients. Soon after, not only did the order numbers of those two soup offerings drop, the total number of soup orders dropped. How come? Well, maybe nobody has asked the customers if the offered soup was tasty at all. A quick survey revealed that customers make the popular choice, find out it's crap, and then do not ever order soup again in that restaurant, or in rarer cases give the other one a shot. It turned out, the most popular offer was basically cheap crap nobody wanted to eat, and when there's nothing else, they keep ordering the same, or never visit that restaurant again.
Telemetry does not tell anything about user preferences. Who ever is selling you that idea, does not either.
Demonstrable, perhaps; but to actually demonstrate it, they'd have to share their telemetry results.
Also, demonstrably better than what? Demonstrably better than a Google-free kernel with no telemetry? How can you demonstrate that?
> I can speak as a GNOME developer—though not on behalf of the GNOME project as a community—and say: GNOME has not been “fine” without telemetry. It’s really, really hard to get actionable feedback out of users, especially in the free and open source software community, because the typical feedback is either “don’t change anything ever” or it comes with strings attached. Figuring out how people use the system, and integrate that information in the design, development, and testing loop is extremely hard without metrics of some form. Even understanding whether or not a class of optimisations can be enabled without breaking the machines of a certain amount of users is basically impossible: you can’t do a user survey for that.
Perhaps instead of adding telemetry, they need to... actually start listening to their users (finally)?
Ha! Like that will ever happen.
For instance, I've never worked with a competent release manager who said "we need more field telemetry!"
Instead, the good ones invariably want improved data mining of the bugtracker, and want to increase the percentage of regression bugs that are caught in automated testing. They also generally want to increase the percentage of automated test failures that are root-caused.
I don't really need or want this pressure on Fedora -- 'the premiere desktop... blah, barf'. This is how we got Canonical and their brazen licensing
I'm happy enough with it basically being RHEL-next. Package the misc things (GNOME, KDE, Sway, etc) with the appropriate SELinux policies and bam, a decent desktop OS.
That way, you can just type "zyp ..." for stuff which is pretty easy, whereas "zypper" I always found to be more of a pain (for typing). "zypper" seems to attract typos, for me anyway. ;)
You could do the same thing with an alias too I guess... :)
I have a System76 laptop I use mostly for travel / edge cases and Pop_OS is decent enough that I'd consider it on the desktop.
The recent closing of build files of RHEL, now telemetry in Fedora. At least it doesn't have ads like Ubuntu amirite...
Spyware everywhere.
Not that I'd ever use Fedora as my main Desktop OS. Arch has won that battle. And if I want a simple installer, where everything just works, Manjaro.
https://lwn.net/ml/fedora-devel/CAJqbrbeOZrHvYjvMCc=qGZD_VXB...
Really, given that "One of the main goals of metrics collection is to analyze whether Red Hat is achieving its goal to make Fedora Workstation the premier developer platform for cloud software development." is right there in the text it would seem pretty clear to me that the 'Fedora Community' (in sofar as one exists) is just Red Hat in disguise serving its own interests here, and at arms length to see what the response is. Think of it as a trial balloon which will certainly be followed by more clicks of the ratchet.
Are you the CEO or within the decision making structure? Otherwise really it just sounds like you just aren't in on the backchannel where these instructions are issued. Oh and inb4 "There is no backchannel!", IBM is subject to Sarbanes–Oxley there's a backchannel.
They bought Red Hat because Red Hat was profitable, the stock was going up, sales were in the early stages of a massive swing upward, and IBM was dropping stock price and waning in market power. It's the same playbook they've been using for decades.
It seems more far fetched to me to believe that IBM bought Red Hat so they could destroy it.
I seriously doubt it.
It seems so much more likely to me that the Fedora team just wants telemetry data (which is very standard in the industry) for decision making, exactly as they said.
Yes, I think that's well within the realm of possibility. IBM bought Red Hat for some reason, so they obvious care quite a lot.
I'm not saying they're actually doing this, of course. I don't know. But we're talking about large corporations here, so suspicion seems warranted.
> It seems so much more likely to me that the Fedora team just wants telemetry data (which is very standard in the industry) for decision making, exactly as they said.
I think this is true! Both things can be true.
Standard in the industry, though? No. It's standard in the commercial software world (which is one of the reasons why I avoid commercial software), but it's not standard in the FOSS world.
The reason why I object so strongly to these sorts of moves from Fedora (and other such outfits) is because there's clearly a strong effort being made to normalize this in the FOSS world, and I think that would be tragic.
Is there any 'relative autonomy' entity for Red Hat, like a board of directors or a council of any kind. Or is Red Hat simply a department in a larger structure?
This question is not meant to be rhetorical or challenging.
Most of the decisions/changes which people attribute to IBM fall into one of a few categories:
* What tends to happen to any company when it doubles / triples in size to >20,000 employees, as Red Hat did from 2017 to 2023.
* Leadership shuffles, such as when Jim was replaced by Paul when Jim became President of IBM. (This happened at the same time as the acquisition, but just because one CEO made choices differently than another might have doesn't make IBM directly responsible for those decisions).
* The rise of Amazon, Google, Microsoft / PaaS and SaaS / containers sucking a lot of air out of the traditional Red Hat market segment. RHEL is still an important cornerstone of the company, but remaining a successful company 10 years from now will require finding additional niches and creating value in ways that are not as vulnerable to entities 100x our size.
So no separate governance/board of directors and a CEO who reports to IBM.
Some services not merged (yet).
Redundancies in a landscape that is changing.
Best of luck.
I'm struggling to think of any company, anywhere, that has that kind of arrangement after an acquisition. Can you point one out?
>Some services not merged (yet).
No services have been merged as far as I know, other than the ones that are literally impossible to not merge, such as the employee stock purchase plan. Please do not twist what I said.
Health insurance is separate, 401k is separate, IT systems are separate. As far as I know the expense system is separate (I haven't needed to use it in years, and I'm on vacation so I can't check).
>Redundancies in a landscape that is changing.
What is even the point of saying this? You asked if there was separation and then dismiss evidence of separation as "redundancies" [that won't last]. I can't predict the future, but the way HN talks about Red Hat you would think the entity known as Red Hat no longer exists (or only barely so). The separation has remained constant for 4 years so far. Not sure what else to say.
It's kind of the culmination of dozens of threads on Hacker News over the last few years and increasing frustration on my end. Mainly the frustration is because I have criticisms of Red Hat, but every conversation seems to jump straight to some variation of IBM and it drowns out the (IMHO) reasonable discussion, or it's (rarely but still happens) a Red Hat person who doesn't think a single decision they've made is bad.
My reference to redundancies arose from my own experience (at massively smaller scale) when organisations merge. Duplicated functions are removed over time.
I think I've said enough for this topic and I hope you enjoy the rest of your holiday.
For instance, in this particular case, the proposal by the Red Hat Display Systems Team may well be considered by the Fedora Engineering Steering Committee and the Fedora Council which sounds like what I mean by 'relative autonomy'.
However it appears that Red Hat itself does not have any 'corporate body' that is distinct from IBM.
I actually feel like the archinstall[0] tool included in the official Arch ISOs really nails easy installation. It's an official way to install Arch that is incredibly user friendly and fast, in my opinion.
I ran Gentoo back in college, did LFS once -- all good learning experiences -- but Fedora or Ubuntu can get me a usable system in 20 minutes and I don't have to think too hard about anything except basic partitioning and a password.
This looks like it'll get me something working, but not commit me to a full-on Manjaro install.
Absolutely! Archinstall is so nice to just get a system up and running quickly. And not many installers (i.e. on other distros) give you so many options for desktop environments. archinstall is a great tool!
https://discussion.fedoraproject.org/t/f40-change-request-pr...
Luckily, OpenSuse Tumbleweed looks to be a pretty good alternative to Fedora. There’s even an immutable version of it, like Silverblue!
For me the big question is why? Proprietary software needs telemetry because the user is not in control and features are only added by the owner of the software, thus the centralized owner needs to know what features to add.
Open source is different. It is decentralized. Anybody can tweak the system to make it better for themselves. In addition, as opposed to proprietary software which sells licenses, the most common open source monetary model seems to be selling support in which case the people buying support can ask for the feature without telemetry.
To put it another way, you need telemetry for a cathedral since decisions are made centrally. A bazaar doesn’t need telemetry, since decisions are decentralized.
And then they proceed to list a bunch of non-ethical and extremely specific data.
99% of users just use whatever the default packages their ON gives them. Practically none of them are digging into the code and making changes.
>as opposed to proprietary software which sells licenses, the most common open source monetary model seems to be selling support in which case the people buying support can ask for the feature without telemetry.
Proprietary software which are sold are in the business of creating products with enough value that people are willing to pay actual money for them. It is smart to look at what these businesses are doing because they are incentivized to make good software so they get to be efficicient about doing so. Meanwhile most open source programs are given out for free and do not care about user experience or providing features for the mass audience.
Asking for features publically is just one, biased signal. There is more to telemetry like seeing what features are actually used. What investments should be double downed on. What features should be made more clear how to use. What crashes are the worst. What parts of your program need to be optimized. etc.
Open source projects can use telemetry to prioritize bug fixes and feature enhancements. You can't fix every bug, but with telemetry you can focus development on issues and features that affect the most users. Here are some use cases that come to mind.
AVX-512 is complicated; there are 21 extensions. What percent of users have some AVX-512 support and which extensions are popular?
Which buggy devices, clients, and applications are worth adding workarounds for?
Which protocol settings can we safely turn on by default?
What are the most popular GPUs? Where would we get the largest benefit from performance improvements, testing, and bug fixes?
What percentage of users are affected by a new Intel CPU bug? How should we prioritize developing a workaround? Does it need to be deployed this week or this month?
Which GNOME applications and configuration settings are frequently used?
What are the most common display resolutions?
What are the most popular installed packages? Should we add them to our base system?
Edit: This was a response to the earlier version of your post that only said “This makes no sense”.
My post was not an IBM-owned-project but it evoked the idea of one inside the reader, much like Magritte
Patents.
I think the Red Hat eco-system is turning IBM Blue.
I'll take this as the warning to move off Fedora, to more forward looking distributions.
Red Hat already lost my work laptop.... now it'll lose my personal one :(
https://discussion.fedoraproject.org/t/f40-change-request-pr...
Opt-in telemetry is garbage. I’m going to stop responding to comments that are
requesting opt-in because I’ve made my position clear: users who opt-in are not a
representative sample, and that opt-in data will not be accurate or useful.
Accurately summarised to "Fuck off dickheads, your privacy is getting in the way of us doing development!".https://lwn.net/ml/fedora-devel/CAJqbrbeOZrHvYjvMCc=qGZD_VXB...
=== What data might we collect? ===
We are not proposing to collect any [...] particular metrics
just yet, because a process for Fedora community approval of
metrics to be collected does not yet exist. That said, in the
interests of maximum transparency, we wish to give you an idea
of what sorts of metrics we might propose to collect in the
future.
One of the main goals of metrics collection is to analyze
whether Red Hat is achieving its goal to make Fedora Workstation
the premier developer platform for cloud software development.
Accordingly, we want to know things like which IDEs are most
popular among our users, and which runtimes are used to create
containers using Toolbx.
Metrics can also be used to inform user interface design
decisions. For example, we want to collect the clickthrough
rate of the recommended software banners in GNOME Software to
assess which banners are actually useful to users. We also want
to know how frequently panels in gnome-control-center are
visited to determine which panels could be consolidated or
removed, because there are other settings we want to add, but
our usability research indicates that the current high quantity
of settings panels already makes it difficult for users
to find commonly-used settings.
Metrics can help us understand the hardware we should be
optimizing Fedora for. For example, our boot performance on hard
drives dropped drastically when systemd-readahead was removed.
Ubuntu has maintained its own readahead implementation, but
Fedora does not because we assume that not many users use Fedora
on hard drives. It would be nice to collect a metric that
indicates whether primary storage is a solid state drive or a
hard disk, so we can see actual hard drive usage instead of
guessing. We would also want to collect hardware information
that would be useful for collaboration with hardware vendors
(such as Lenovo), such as laptop model ID.
Other Fedora teams may have other metrics they wish to collect.
For example, Fedora localization wishes to count users of
particular locales to evaluate which locales are in poorer shape
relative to their usage.
This is only a small sample of what we might want to know; no
doubt other community members can think of many more interesting
data points to collect.
That last piece "no doubt other community members can think of many more interesting data points
to collect" sounds pretty bad for telemetry that's enabled by default, with people having to opt out of it. :(> A new metrics collection setting will be added to the privacy page in gnome-initial-setup and also to the privacy page in gnome-control-center. This setting will be a toggle that will enable or disable metrics collection for the entire system. We want to ensure that metrics are never submitted to Fedora without the user's knowledge and consent, so the underlying setting will be off by default in order to ensure metrics upload is not unexpectedly turned on when upgrading from an older version of Fedora. However, we also want to ensure that the data we collect is meaningful, so gnome-initial-setup will default to displaying the toggle as enabled, even though the underlying setting will initially be disabled. (The underlying setting will not actually be enabled until the user finishes the privacy page, to ensure users have the opportunity to disable the setting before any data is uploaded.) This is to ensure the system is opt-out, not opt-in. This is essential because we know that opt-in metrics are not very useful. Few users would opt in, and these users would not be representative of Fedora users as a whole. We are not interested in opt-in metrics.
Those last three lines are essentially adding insult to injury. I have no problem with telemetry, I do have a problem with telemetry being on "by default" as it puts the burden of safeguarding privacy squarely with the user. As an individual you have to make sure the toggle is set to "off" yourself, and assume that telemetry really is switched off.
Back in february, the Go project proposed to add opt-out telemetry in the Go toolchain and they were almost torn in half over a weekend.
Rightly so.
It totally fails to ensure that. The underlying setting is off, but if you don't take deliberate action on installation, it gets switched on. It might as well be on unless you act to switch it off.
The meaning of "opt-in" was settled a long time ago; it means that if you don't take deliberate (and informed) action to opt-in, then you haven't opted in.
[Edit] Sorry, mis-read parent comment. They're trying to say it's opt-in, but the proposal is to make an opt-out UI, with an underlying off setting that gets set to $SOMETHING at installation, defaulting to on.
Sigh. Yet another Linux distribution is becoming adware. Ubuntu already shows ads (and spreads FUD) in apt, but they haven't resorted to tracking yet.
IBM / RH Salesperson: "But think of the opportunity! GNOME Software could be a REAL app store[0] just like the Apple App Store or Google Play and bring the Linux desktop into the 21st century. We could even call them 'donations' so that users feel good about paying. We will need to have Flatpak integration to allow less technical users to get the software we are moving to community support like LibreOffice. With payment account setup and processing, we could easily add commercial software and offer vendors more customers with minimal effort on their part."
"Once we get telemetry enabled and users setup their online accounts in GNOME Online Accounts[1][2], we can make deals with the online account providers for targeted advertising metrics. We won't show any ads in GNOME for now. But after users are used to seeing some promotions for Red Hat products and services, it is just a slight change to show other promotions."
[0] https://wiki.gnome.org/Design/Apps/Software
[1] https://wiki.gnome.org/Projects/GnomeOnlineAccounts/Provider...
OK, so they need telemetry to measure the effectiveness of their ads. I think the rest is padding.
I fear what that means where the preference is saved, and how spins (or users who simply choose to not have GNOME), may feasibly opt out
Where's the demarcation? Is this some dconf thing that a timer will read, a service, or what?
I lack trust in their handling in certain matters. For example, every Fedora device 'phones home' for AP checks:
$ cat /usr/lib/NetworkManager/conf.d/20-connectivity-fedora.conf
[connectivity]
enabled=true
uri=http://fedoraproject.org/static/hotspot.txt
response=OK
interval=300
Including those that are wired... and that's rather unnecessary. I generally get a sense of haste these decisions, lately.To disable it, mask the /usr file with one in /etc:
touch /etc/NetworkManager/conf.d/20-connectivity-fedora.conf
Another example: systemd-oomd on anything with > 64GB installed; entire user scopes randomly killed with oodles free.I say this in the softest way possible, I don't really mind it... but it raises an eyebrow towards eagerness/attention.
<snide>They already get to know how well they're doing being "the premier OS" by seeing how often they get hit.</snide>
* The proposer has clearly not done any research on how to actually collect anonymous data (they'd never heard of differential privacy for example).
* They want a plug and play solution (they specifically say they don't want to do more work than that)
* They are not open to discussing privacy regulations such as GDPR
* They are not willing to bend on the most contentious points of their proposal
* The system they want to use collects invasive metrics that can be de-anonymized and has only been used by a niche distribution
Because the de-anonymization bit might not be clear, let me summarize some of the things that the Endless OS metrics collect:
* Country
* Location based on IP address to within 1 degree lat/long
* Your specific hardware profile
* Daily report that includes your hardware profile, along with the number of times the check ins have occurred in the past
* Detailed program usage (every start / stop)
* An unspecified series of additional metrics that can be sent from anywhere else on the system via a dbus interface
Additional this proposal wants to explicitly collect:
* What packages and versions of such are installed
* Specific application usage metrics (the example they give is the gnome settings panel)
They discard the IP address, but how hard do you think it is to differentiate users based on the combination of hardware profile, +/- 1 degree of location accuracy, their specific set of packages (and knowing the history of package installs/uninstalls already through their package manager). The proposal doesn't meet its stated intentions of being anonymous, and the proposer actively understands that users don't want this but believe their desire for the metrics overrides the end users desire of not being tracked.
Maybe we need a particapatory privacy stack that produces valuable anonymous data and also contributes to it. You might be able to do it with homomorphic arithmetic that increments defined counters (like the hash of a package or version), and we already have distributed ledgers for collecting and distributing the data. We can do queries with differential privacy, and zksnarks.
It's not a viable product because people who actually use data want the real data, the discretion is power to them, but as a tool for coordinating a cooperative effort, we need to build something new to say that this is how we do things now.
A few years ago Ars Technica had a site redesign. When it first came out, it didn't have a dark mode, and the comments on the article announcing the change (after the design had been rolled out) were full of people who were upset with the lack of dark mode. Turns out, many of the subscribers and power users that used dark mode had adblock and the like that blocked metrics collection, leading the Ars web designers with the impression that basically nobody used it. Since then I've opted in to telemetry in programs I trust to be good stewards of that data so that I can at least do a little to ensure that whatever weird setup I might be using, like dark mode, is seen as relevant to the project.
But that does bring up the point that just relying on telemetry doesn't always present an accurate picture of what's going on with all of your users. Probably the best way to see what your users are doing would be a combination of telemetry (to see what the average user is doing), surveys (to see what the enthusiastic user who self-selects into completing the survey is doing), and user studies (for specific design decisions). I'd like to see what Fedora's policy for data collection in the last two categories is and if they'll integrate all of it together for more comprehensive decisions.
Also, I'm not surprised by the instinctive reactions that people have to the word "telemetry" in the headline, but the proposal is really well done. It addresses many of the complaints I see in these comments, like if it should be opt-in or opt-out, if it's legal under GDPR ("Fedora Legal has determined that if we collect any personally-identifiable data, the entire metrics system must be opt-in. Since we are only interested in opt-out metrics due to the low value of opt-in metrics, we must accordingly never collect any personally-identifiable data."), and the like. I think that this proposal was done by a person or team that genuinely sat down and thought through what a real privacy preserving telemetry implementation would look like, rather than the typical corporate claim "it's totally privacy preserving (also our product is not available in the EU)!!!"
Stop trying to treat my computers like they are a part of your test lab.
Ideally, the only statistics you have are from me and my Sybils, and all of society's energy is dedicated to improving life for me.
Telemetry isn't enough if they really want to imitate Windows 11. They need to add advertisements as well. And more preinstalled crapware.
And consider requiring (or at least strongly recommending) paid cloud-based login for all accounts.
not my os, not my problem
is it: fedoraproject.org ?
> all of the components of the server (discussed below) are open source, and we will provide instructions for how to run a simple server yourself and view its metrics database. You can redirect metrics from Fedora’s server to your own by changing a URL in a configuration file.
I'm fine with these changes as long as they are transparent.
So, to me, as long as they use a specific and dedicated (and known) fqdn, I may just block the whole telemetry adding that entry in /etc/hosts file), can't I?
https://retrace.fedoraproject.org/faf/reports/
This is privacy respecting and useful.
The privacy respecting part here seems to be handled carefully, I'm not sure if it is opt-in or not (in my opinion it should be) and the payloads appear to be benign, but, that could still go off the rails if there is a bug in ABRT.
In principle I want my machines (desktops, servers) only to initiate network calls and to respond to network calls that I allow. Outbound firewall rules are there for a reason.
it's useful to me as a user.
Just do it.