If you want usage statistics for packages just track how often individual packages are downloaded on the server side. A maintainer has no need to know who's installing what.
If you want usage statistics for packages just track how often individual packages are downloaded on the server side. A maintainer has no need to know who's installing what.
To be clear: Homebrew has no idea which users are installing what. We only store counters for package install, failure, etc. events, and everything that's stored is visible on the Homebrew website[1].
Homebrew's architecture doesn't really have a "server side" in the way your suggestion requires: the formulae and bottle components rely heavily on public services like GitHub Packages and GitHub Pages, which don't offer those kinds of analytics.
FD: Member of Homebrew.
Yes: Homebrew deprecates and/or disables packages if we see evidence that they're unmaintained and not actually supported on the platforms we support, or only used by a tiny fraction of users while also requiring disproportionate maintainer time (e.g. due to complex or flaky builds).
The goal is to balance conflicting user interests: 99% of users want maintainer effort focused on the top 100 (or 500, or 1000) packages, and many of those packages also require significant maintainer effort (e.g. making sure that they don't cause transitive breakages).
According to https://docs.brew.sh/Analytics they use it to measure how often formulas fail to install, to get overall metrics on which OS versions are used, and to correlate those (i.e. to tell on which OS versions specific packages fail to install correctly).
> A maintainer has no need to know who's installing what
Aside from the IP, they don't know who's installing what, and in the new model announced in this post they now don't store IPs or any other user token at all, so it should be purely anonymous aggregate metrics.
Edit-Also, this is for Mac OS. Chose a few standard OSes to support and test them. If a system update will fix the issue then it shouldn't be fixed at the package manager level.
Based on what I read on the site, that looks like exactly what they are doing, and they are explicitly NOT storing information that would identify "who".
Do any of you actually work in this industry shipping software products to end users? Without telemetry the problem there is literally one of trying to read the mind of your end users to figure out what they're doing, hoping that your internal CI manages to reflect the configuration in their environment.
I don't see why something that's little more than a file server needs telemetry.
You're doing an awful disservice to Homebrew.
I don't like telemetry at all and I believe we have to find other ways to do QA. Hence my strong reaction.
The most valuable one I’d guess is package install error rates. Seems pretty useful to me.
If so, that's... not ideal for sure.
However, we can't catch everything: Homebrew has millions of users, and those users have all kinds of different setups. We can't predict every possible host and software interaction; basic failure analytics help bridge the gap there.
Where can I learn more? Can you point me at the right place in the source?
I'll not be banning Homebrew telemetry.
I've linked Homebrew's analytics data and the source code that collects it elsewhere in this thread.
And, to be absolutely clear: it is perfectly fine for you to disable Homebrew's analytics. There are an infinite number of legitimate reasons for doing so, including the most basic one of "I just don't want to." My sole goal is to dispel the small number of inaccurate beliefs about what Homebrew collects, why we collect it, etc.
Which the software that I used to be employed maintaining has actually broken homebrew compiles when they've been installed at the same time (which I think I made better but I never got the PM who actually owned the product to spend the resources to properly fix).
A good example of how the configuration in the end user environment can affect package installation.
In addition to being INCREDIABLY slow, now I have to worry about what it might spy on. If I have a problem I'm more than happy to go to GitHub (or which ever site it's hosted on), and report it.
It’s very hard to write and maintain good software without knowing how it’s used. No package manager needs to know how you specifically use it, but aggregate data and the ability to identify scenarios it does not handle well are both very important for SW lifecycle.
https://opencollective.com/homebrew#category-BUDGET
They seem to be receiving about US$2k/month via Patreon too:
https://www.patreon.com/homebrew
I think their Patreon was around the same when I looked ~12 months ago.
It's also pretty likely they get a lot of their resources for free / sponsored too.
So, US$2k/mo might be fine (no idea).
Do you really not see any advantage to maintainers having visibility into what packages people actually use?
Are you paying for the compute?
I imagine they can get much richer metrics through this as opposed to only tracking downloads on the server side.
I'm not saying I like it. In fact, I plan to keep it disabled. I'm just saying it's a bit naïve to think client-side analytics are the same as server-side download tracking.
If anything they need money, of course, and to know their software works for their users. Prior to release have a test system install the full base, test those packages work, and you know anything less will work too.
- Time to install packages - Versions of things - Has the compilation (when required) failed? What dependency versions are installed? - CPU architecture - OS version ...
There's a lot more that can be sent from the client that's not available on the server side.
> I fundamentally don't 'get' what richness they actually need.
That's fine. Perhaps you could ask them instead of ranting about what you don't know or don't 'get' in a public forum?
> to know their software works for their users
Sounds like you're not very far from understanding why they want better telemetry.
> Prior to release (...) and you know anything less will work too.
Things break in unexpected ways. OSs are complex systems and there's a lot of interactions between components. Homebrew's user base is enormous and very diverse. There's 2 different architectures, many OS versions, lots of environment variables that might be set differently in each user's systems, different versions of libraries, ... I could go on but I think you get the picture.
Edit: s/collected/collecting/
So even if I have an issue with telemetry, good on the Homebrew maintainers for ignoring this MR.
I think you know developers who reject informed consent will never adopt an informed consent model. The proposal was the best users could hope for realistically. Did you never compromise?