HomeBrew Analytics – top 1000 packages installed over last year
brew.sh
brew.sh
I've been using it instead of grep the last few months and I could never go back. Check it out if you haven't! Here is the repo and a technical breakdown by the author:
It's the one I personally use simply because I discovered it before ripgrep came out. Plus the 'ag' commandline is super easy.
I'm willing to bet (and happy to lose) that most people, even in the subset who use grep "a lot" (defining "a lot"...), wouldn't see a significant improvement. They are people (I'm betting fewer) who need speed above all other concerns, and those people already make it to the top 1000.
> I'm willing to bet that most people, even in the subset who use grep "a lot" (defining "a lot"...), wouldn't see a significant improvement.
Daily, but not for the reasons that I think you're thinking. I work in Python a lot, which means there is typically a virtualenv in the tree somewhere, sometimes more than one. Typically, I want to search the code base itself — not a virtualenv, not the .git directory, etc. ripgrep, by default, will ignore entries in the .gitignore (and the virtualenvs are listed there, as they're not source, and cannot be committed), and repository directories like .git, and will thus not even consider those files. For my use case (searching my own code base), this is exactly what I want, and culling out those entire subtrees makes rg considerably faster than grep.
Yes — I could exclude those directories with grep by passing the appropriate flag. But it's time consuming to do so: ripgrep wins out by doing — by default — exactly what I need.
I also greatly prefer ripgrep's output format; the layout and colors make it much easier to scan than grep's.
Most of the people I've recommended ripgrep to are using grep, and passing flags to it to get it to do what essentially rg does quicker and/or by default. Ripgrep is an excellent tool.
(I used `git grep`, which is also considerably faster for similar reasons, prior to rg. But `git grep` requires a repository — for obvious reasons — and thus fails in cases where you're not in one. I often need to search several codebases when doing cross repository refactors, and ripgrep has been quite useful there.)
But yes, the feature sets are very similar.
You'll easily notice the difference on any recursive search. grep is really slow.
Before I wrote ripgrep, I was a grep user. I hadn't migrated to ack or the silver searcher because I didn't see the need. (Sound familiar? :P) In my $HOME/bin, I had grepf:
#!/bin/sh
find ./ -type f -wholename "$1" -print0 | xargs -0 grep -nHI "$2"
and grepfr #!/bin/bash
first=$1
shift
grep -nrHIF "$first" $@
And that was pretty much all I ever needed. If ack had never come along, I'm not sure I ever would have changed. The tools I had were good enough.ripgrep didn't begin life as something that I intended to release as its own project. It began life as a way to test the performance of Rust's regex engine under similar work-loads as the regex engine in GNU grep. In other words, it was a benchmark that I used. (In fact, I used it quite a bit to reduce per-match overhead inside the regex engine. The second commit in ripgrep says, "beating 'grep -E' on some things.") I didn't really start to convert it to a tool that other people could use until I realized that it was actually as fast---or faster---than GNU grep. That, plus I was bored in an airport. :-)
A lot of people are happy with their tools that are good enough. I know I was. Has my life been dramatically changed by using ripgrep? No, not really. But I do like using it over my previous tools. It's a minor quality of life thing. It turns out, a lot of people care about minor quality of life things!
But yeah, I hear roughly the same sentiments that you say from a lot of people. All it really comes down to is different strokes and different common workloads that magnify the improvements in the tool.
Second, I include myself in the users of ripgrep (and the silver searcher before), I also dislike the slight waiting time when I have a better alternative.
Third, I'd like one those to be a default package in Debian.
> Third, I'd like one those to be a default package in Debian.
Yeah, I'd love that too! I know there have been people pushing on this, but AFAIK, it's stalled on "how do we package Rust applications in Debian."
(I don't use Debian and I'm not terribly familiar with their policies, so I'm not really familiar with the details.)
https://crates.io/crates/debcargo is also a big help.
But typically, I am not really looking for string "foo". I guess most users who are grepping log files are also looking for strings/text that appear slightly before and slightly after the match. This mostly happens when I am searching for an error/exception that triggered "foo". I find usability of grep frustrating when I need to search around something. It usually means, I have to restart the search with `grep -C` or something like that and even then line numbers I specified may not be enough.
rg foo -C 20 -p | less -R
Also, yuck, that command line. Glad I have those hidden behind shell scripts.
#!/bin/sh
exec rg -p "$@" | less -RFXIn particular, in order to better show results, I kind of feel like the search tool needs to know something about what it's showing. Right? How else do you intelligently pick the context window for each search result? For code, maybe it's the surrounding function, for example.
The grep -C crutch (which is also available in ripgrep) is kind of the best I've got for the moment for a strictly line oriented searcher. `git grep` has some interesting bits in it that will actually try to look for the enclosing the function and emit that as context. I think it's the `-p/--show-function` flag. ... But that doesn't really help with your log files.
In any case, I am very interested in this path and even have an issue on the ripgrep tracker for it: https://github.com/BurntSushi/ripgrep/issues/95 --- I'm not sure if it really belongs in ripgrep proper, but I would really love to collect user stories. If you have any, that would be great. Examples of what you'd like the search results to look like would be great!
It's better than ag in a lot of ways, but there are little pain points like that which make me shift back and forth between tools.
$ rg -g l foo
No files were searched, which means ripgrep probably
applied a filter you didn't expect. Try running again
with --debug.
Empirically not! :)EDIT: I see, you have to do "rg -g '*.l' foo". Well, that's a bit silly. Why force people to put asterisks inside of single quotes on the command line? Asterisks have a specific meaning in a shell setting. It's five times longer than -G l, the ag equivalent.
EDIT: Thanks for all the explanations.
Actually it's just:
rg -g '*.js' query
And if it's a known file type, like js, you can do (-t type): rg -g -t js
>Why force people to put asterisks inside of single quotes on the command line?*Because else the shell will auto-expand the asterisk before it even gets to rg, and rg will instead get the expanded list of files that match the pattern. E.g. if you have
a.js b.js /foo
in a folder, then: rg -g *.js query
will be expanded by your shell to: rg -g a.js b.js query
and THEN run. rg will never see the asterisk in that case (and it also wont search inside /foo).The asterisk is part of standard globbing. You can also write it as `rg -g \{STAR}.l foo`, if you find that nicer.
If you want to match a list of extensions, then you can fall back to standard glob syntax: `rg -g '{STAR}.{foo,bar,baz}' pattern`. Or, as others have mentioned, if you're searching for standard file types, you can use the `-t/--type` flag. e.g., To search HTML, CSS and Javascript: `rg -tjs -thtml -tcss foo`.
Basically, ripgrep's `-g` flag is supposed to match grep's `--include` flag, which also uses globs and requires the same type of ceremony. I'd like to add --include/--exclude to match grep's behavior more precisely (which is based on user complaints wanting those flags).
N.B. Replace {STAR} in text above with a literal asterisk symbol.
rg -g*.l fooThanks for the kind words!
If you're on Windows, ripgrep will also (seamlessly) search UTF-16.
it's just yet another tool which i'd have to learn, and yet another tool that isn't going to be installed on the remote machine.
is the productivity win worth installing it, let alone learning it compared to going from `grep` to `git grep`? for me probably not, diminishing returns. doesn't mean it isn't great software!
(also, for the sake of stats, homebrew is macOS only, although i appreciate the info)
It’s also only very recently that the crown has passed from ag (the silver searcher) to rg.
I've been using ack-grep for years, seems fine.
My blog post linked elsewhere goes into more detail, although I left out ack because it was too slow.
> Notably absent from this list is ack. We don’t benchmark it here because it is outrageously slow. Even on the simplest benchmark (a literal in the Linux kernel repository), ack is around two orders of magnitude slower than ripgrep. It’s just not worth it.
[Disclosure: I'm the original author of popcon so I'm biased :)]
> Homebrew's analytics are accessible to Homebrew's current maintainers.
https://github.com/Homebrew/brew/blob/master/docs/Analytics....
Yeah, I know, my tinfoil hat is quite elaborate.
EDIT: Like I said, my tinfoil hat is quite elaborate; I already have analytics and automated updating turned off. I like to retain some control over what information is reported back to Google or other projects.
That doesn't somehow boost my confidence in Google and volunteer-run open source projects to properly respect the privacy of their users.
brew analytics off
You can check the current status with: brew analytics # disable homebrew from tracking metrics
# see https://github.com/Homebrew/brew/blob/master/docs/Analytics.md
HOMEBREW_NO_ANALYTICS=1I wasn't part of the creation of that particular page, but one thing we (the maintainers) commonly find ourselves doing is publicly referencing install statistics as justification for removing an unused formula or taking extra care during a version bump. Having public statistics for the top 1000 most popular packages makes those considerations a little bit more transparent.
Edit: I've created a PR to fix the language: https://github.com/Homebrew/brew/pull/3120
This does not, in my mind, help the case at all. This culture of deletion makes no real sense to me; something not being used for a month or a year doesn't mean it's not being used at all. What's the real cost of just leaving those formula around? A slight problem with it not working immediately when it is used? At least then it can be fixed, instead of seeing the "No formula found" text.
The culture of deletion surrounding the long tail of digital artifacts just doesn't make any sense. Yeah, yeah. Get off my lawn, too.
(It's pretty easy to contribute your own formulas, too! The checklist is pretty small.)
That said, I've occasionally run into a broken `brew cask install X`, so I get that maintenance is a thing. But in that situation it seems best to let someone else notice that and contribute a patch rather than remove it entirely. I understand that might eventually lead to a large percentage of broken casks though.
In the best case, the community does pick up the work of patching and maintaining formulae (and casks). However, that's the best case, and low-volume formulae rarely fall into it.
In my experience, formulae tend to get abandoned by the upstream that originally submitted them, causing them to eventually break when the system or surrounding dependencies change. One potential solution is to have a chain of custody for formula maintenance, but that hasn't worked so well for MacPorts.
IMO that's self-inflicted by homebrew's policies. from https://docs.brew.sh/Acceptable-Formulae.html :
> We frown on authors submitting their own work unless it is very popular.
I'd expect "very popular" software to have relatively large populations of users capable of maintaining the formulae (and authors less likely to bother with packaging), while in lesser-known software, the respective authors are most likely to maintain long-time interest.
Maintenance.
When unused and unmaintained projects and formulae pile up in Homebrew, we end up spending a tremendous amount of time and effort patching mostly unused software for a very small part of the userbase. That dis-proportionality hurts the 95% of users who expect timely and well-tested updates to major packages.
We used to provide a "boneyard" tap for unused/unmaintained formulae, but even that led to a lot of requests for support that we simply can't provide. If something is being removed from the core tap, our current recommendation is to put it a personal tap[1].
> The culture of deletion surrounding the long tail of digital artifacts just doesn't make any sense.
Keep in mind that the "artifacts" in question are still available, since Homebrew and all Homebrew taps are just Git underneath. You might not be able to build an old formula for compatibility reasons, but all prior work is available for reference.
[1]: https://github.com/Homebrew/brew/blob/master/docs/How-to-Cre...
IOW, HomeBrew seems to be big enough that it needs to start engaging in some PR.
You don't see that text but instead text indicating the formula has been deleted.
> The culture of deletion surrounding the long tail of digital artifacts just doesn't make any sense.
They will be preserved in the Git history for as long as we continue to use Git.
Both the above comments make it seem like you've not really done any research into your own objections, here.
Anyways. Having done my research:
> You don't see that text but instead text indicating the formula has been deleted.
If, and only if, you know the exact name of the package. If you do a search, you get the "no formula found". If you attempt to use the braumeister.org web site, it also will not show up unless you explicitly craft the recipe URL.
> They will be preserved in the Git history for as long as we continue to use Git.
How does one go back into git history and revive deleted formula, given how automated the entire process is around ensuring the most recent version of brew is always in use? By your "deleted" help docs, I infer that it comes down to creating your own cask?
The one example I came across quickly is "abi_compliance_checker'. It was deleted because it requires GCC 4.7. Yet Homebrew is still quite capable of installing and using gcc@4.7. Not that it was broken in a recent build. Not that it was a major maintenance burden - the updates for years consisted of version updates.
This isn't something I need, but it's a great (and quickly found) example of seemingly arbitrary deletion of an otherwise active project.
It may well be that some relatively-little-used tool nontheless has a very significant use. Popularity is one factor within the mix, but only one.
We generally only refer to popularity after a formula comes up on our radar due to failing other sniff tests. To give you an idea for some of them:
1. Multiple subsequent releases without anybody bothering to update the formula
2. Historical problems with the package (flaky builds, complex build systems)
3. Historical problems with the upstream (patches ignored, unwillingness to cooperate with packagers, unreliable download servers)
Since homebrew is aimed mostly at technical types and devs, build tools themselves are probably fairly highly featured, but those are easy things to lose in large general-public releases. (Android's abysmally poor shell tools come to mind.)
Sounds like a pretty good basic set. The flaky builds criterion in particular seems like a strong signal of a low-quality upstream.
Why is telemetry collected through Google? I suppose that’s because they have some particularly convenient APIs, however please consider that many user are concerned by the amount of information Google has about their lifes.
(I know, it’s possible to opt out, but good choices should be the default!)
No-one else provides anywhere near the same scale for free. We have over a million monthly active users and every other solution we tried (including FOSS ones) either fell over at that load or we couldn't find anyone willing to provide hosting for us.
We don't send any personally identifiable information to Google. You are tracked by a randomly generated UUID (which you can regenerate at any time) and we tell Google to not store IP addresses.
> No-one else provides anywhere near the same scale for free.
That's a bit naive. Google is not a charity and provides that service by making a profit out of user data.
> We don't send any personally identifiable information to Google. You are tracked by a randomly generated UUID (which you can regenerate at any time) and we tell Google to not store IP addresses.
Even with no IP, Google can easily cross-references searches and the random UUID, since a typical use case is that a user installs something through Homebrew after Googling it.
Please rethink this bad default choice. And sorry if my reply sounds harsh, but I think that Google tracking by default is an extremely bad choice for your users.
So what if they make profit -- they provide a fantastic service that actually works well, and it allows homebrew (and many other systems) to continue to provide it's services for free.
Opt-in telemetry clearly stating that Google will record the first three bytes of your IP? I’m OK with that!
Yep and that's an acceptable trade-off for the maintainers of Homebrew given that we need analytics to do our (volunteer, free-time) job adequately and we do not have financial resources for other alternatives. If you're willing to provide those financial resources indefinitely: get in touch.
> Even with no IP, Google can easily cross-references searches and the random UUID, since a typical use case is that a user installs something through Homebrew after Googling it.
That may be technically possible but I see no evidence that it is the case.
> Please rethink this bad default choice. And sorry if my reply sounds harsh, but I think that Google tracking by default is an extremely bad choice for your users.
We disagree. Please consider another macOS package manager. MacPorts is a good alternative.
Yes, you do. While you use[1] the "Anonymize IP" option, a packet is still sent from the user's IP. Google's business model includes gathering as much data as possible so it's foolish to think that they are throwing data away in this situation. You may disagree and trust Google to honor the "Anonymize IP" option, but trust is not transitive so you shouldn't ever assume users agree (use opt-in in every situation).
However, claiming you don't send pii to Google makes me wonder if you have actually read the documentation for GA? The "Anonymize IP" (aid=1) option is blatant doublespeak. From their own documentation[2]:
> The IP anonymization feature in Analytics sets the last octet of IPv4 user IP addresses ...
They are only masking out the last 8 bits of the address, which are the least interesting bits. You can still discover the ASN from the remaining data. At worst all that option did is add a 1-in-256 guess when correlating your analytics data to the rest of Google's tracking data. That is trivial to overcome Google's massive databases of tracking data.
You even provide a unique additional per-install tracking number that lets Google track users when they move to a different IP address. Once a correlation exists between your analytics data and everything else at Google, your analytics events provide a reliable report about that can allows other tracking data to be correlated to the new IP address.
Why does that option exist? It's possible that it was designed to mislead developers into sending Google tracking information, but their own documentation[2] suggests a different hypothesis:
> This feature is designed to help site owners comply with their own privacy policies or, in some countries, recommendations from local data protection authorities
This is a feature designed to check boxes on compliance requirements, not to provide any provide actual anonymity to users.
[1] https://github.com/Homebrew/brew/blob/fd4fe3b80cab9902437016...
[2] https://support.google.com/analytics/answer/2763052?hl=en
> Yes, you do
PII is a term of the art which the GP is using in its standard sense and you are not. https://en.wikipedia.org/wiki/Personally_identifiable_inform...
(This is independent of the deontolic status of your comment.)
Similar data could probably be compiled from Google Trends, but this clean and authoritative view is something that can be trusted, and I think the result makes the open-source ecosystem stronger.
I'll happily tell you my toolset, but I don't use home brew and I wouldn't allow it to send my usage to google even if I did
Just because someone on the internet had a particular setup doesn't mean I want to follow it. Or that I have time to track down several people's opinions.
Getting install stats directly from the homebrew project, which I know because I use it, is infinitely more useful to me and much more easily discoverable. that's just my opinion though and you're entitled to your own.
That also gives you more context, because it's answering your actual question, rather than trying to answer your own question with a bunch of vaguely related data.
It also handles the dependency issue. Someone asked why imagemagick is so popular, but its probably actually just a dependency for language-level bindings (e.g. php-imagick), not that people are using `convert` or `identify` directly on the CLI.
Heck, consider the case of front-end developers who have a toolset that depends on nodejs. They may never write any server side code, but if they follow recent trends they probably need nodejs for their css/js "toolchain" - the stats don't tell you that though. They just tell you that nodejs is installed a lot.
This list was added specifically because a bunch of people expressed concerns about this data not being public. You can please some of the people all of the time, etc.
Interesting; I don't see apt, dpkg, nix, or yum maintainers releasing top lists of packages. Even the package providers rarely provide a "top" list, especially one with compile-time flags.
Those flags being provided in this list are very explicit, which could be easily correlated with data provided (or harvested) by other parties.
> a good way to measure the health of the Homebrew ecosystem
What value does this really provide to the general public, other than a, ahem, "member" measuring contest? You already see this occurring in this very thread - the MySQL vs. PostgreSQL comments. I fully expect a "Node vs. X" one as well.
I can't remember whether the default is opt-out or in with the current installer.
It also doesn't sent data to a privacy abusing company like google.
Actually the data is sent to Google.
If you don't trust us to respect your privacy why are you trusting us to download and run code from the internet on your machine?
We need analytics to effectively run the Homebrew project (as you've noted: we are volunteers). If you no longer trust our project: go use something else.
I can't on yours (or Google's).
EDIT: I'm sorry, was there something factually wrong with this statement to warrant downvoting it into oblivion?
I would estimate I download tmux via homebrew once a year on average. It is possible 112,000 represents a good estimate of all mac carrying tmux users in the world [*]. For comparison this is roughly the same number as employees at Apple.
Assuming these stats aren't opt in.
And what are the pros and cons?
document.querySelectorAll("td > code").forEach(code => code.innerHTML = `<a target="_blank" href="http://brewformulas.org/${encodeURIComponent(code.innerText)}">${code.innerText}</a>`)
Cut and paste that in your Javascript Console (`View/Developer/Javascript Console` on Chrome) when you're on https://brew.sh/analytics/install-on-request/.I found myself tabbing out to google constantly like 'fdk-aac, wow, that's a thing? cool~'
Great software by the way.
The automatic LetsEncrypt stuff is great and has removed crufty cron jobs to make sure the certs are up to date, and support for HTTP/2 out the box (with 0 configuration) has seen a marked improvement in the performance of my site.
Also, Perlbrew is not necessarily easier, simply because having one package manager manage your packages is better than having many package managers (cargo, npm, gem, etc). It suggests that maybe there should be a package manager for all these package managers... I'd personally want to use the simplest way possible to install something unless I had a reason.
Random pedantry: the list on that page is the stuff where people have typed `brew install perl` (or `brew upgrade perl`) and it hasn't been pulled in as a dependency. As to why: I have no idea :D
I have no data to back this up other than my anecdotal personal experience but I'd wager Perl hackers who use Homebrew never use it to install Perl.
Most of the classic Perl documentation walks you through installing from source, which is quite easy, and modern documentation/tutorials usually refer to perlbrew.
The Perl community often has established solutions within their own ecosystem, and this makes them seem like they have a smaller presence then the otherwise would have.
#575 algol68g 3,485 0.02%
Now that's awesome."Ah, wget. I guess I do need that."
Could it be some node package that has a dependency on imagemagick?
What I'm having trouble wrapping my head around is so many people being aware of imagemagick and reading through it's documentation to learn how to use it.
many, many use-cases. i use it all the time, it's just more convenient than pixelmator/gimp/photoshop for those tasks.
There are also instructions on how to opt-out.
Python3 #5, Python #6, Pypy #439 ...
Go #22, Scala #78, Elixir #83, Rust #141, Typescript #584 ...
Groovy #150, Kotlin #195 ...
Go #22, Ruby #33 ...
#85 fish
I am glad to see these 2 in top 100.
Interesting to see MariaDB in the middle of that list though...