It sparks joy in my heart whenever I see shade cast against pandas.
It sparks joy in my heart whenever I see shade cast against pandas.
Close second is the plotly library.
The user guide material absolutely needs work, and the examples in the reference docs tend to be a little contrived. But I absolutely have seen worse-documented libraries, such as Gunicorn and Pydantic.
(I also wrote a Pandas book or two... So there's that)
But for a lot of people who use it infrequently its documentation is a frustrating mess. Simple problems turn into significant time sinks of trying to find which page of the documentation to look at.
A lot of issues are made worse by shit-awful interop between libraries that claim to fully support dayaframes, but often fail in non-obvious ways... meaning back to the documentation mines.
I'd argue that because there's a market for a single author to write two books about it is indicative of documentation problems.
However, I always though the 10 minutes to Pandas page was decent for getting started. I picked up Polars recently and thought it was more difficult than Pandas because there wasn't any quick intro docs. What projects have great introductory docs for you?
Also, I am curious to learn more about the specifics of interop libraries you are referring to.
Learning a new tool is generally a challenge. I think another challenge with a lot of data tools is that non-programmers tend to be the major audience. I make my living teaching "non-programmers" how to use these tools.
That said, I always teach "go to the docstrings and stay in your environment (to not break flow) if you can." The pydata docstrings are better than most, including Python (the language).
However, maybe it makes more sense that it's just a mess that's hard to document.
No, I want you to force me to provide my data in the right way and raise a noisy exception if I don't.
I agree that the magic type auto-detection is a bit too magical and sloppy, but you have to realize that data analysts and scientists have historically been incredibly sloppy programmers who wanted as much magic as possible. It's only in recent years that researchers have begun to value some amount of discipline in their research code.
One look of dplyr code over pandas would of course disabuse anyone of the notion that R is trash and the tragedy is Python will in the current state never have anything like that. That's the advantage of the language being influenced by Lisp vs not.
I agree that it is a trash language and that, outside that many frontier academic ideas are available and some plotting preferences are solidly prescriptive, it should be thrown into the trash bin.
Python, Julia when it gets its druthers for TTFP, Octave, Fortran, C, and eventually Rust. These are the tools I've found in use over and over and over again across business, government, and non-profits.
Everywhere R is used by the org I have seen major gaps in capacity to deliver specifically because R doesn't scale well.
I agree that the standard library is what you might call "a chaotic disorganized mess".
"Trash", despite its connotations of lacking value, is really just a chaotic disorganized mess of something made by artifice with dubious reclaim/reuse/recycle value. Being a subjective assessment, it is natural that one person's trash is a treasure to another.
It's fine that people like it. What's good about it isn't unique, and what's unique about it isn't that great. And there are certainly switching costs for some orgs to consider.
x <- 3