Dplython: Dplyr for Python
github.com
github.com
But I feel like it would be better to use method chaining for the piping of transformation rather than overloading dunder method operators. It would preserve one of the nice things about dplyr -- composing complicated transformations from a simple vocabulary, but more pythonic. This is a relative weakness I see in the design of pandas and would love to see ported over.
But also, dplyr is a thing that really goes beyond pandas. It's really an elegant, SQL-like DSL for transforming (mostly) arbitrary data. In this way it's more like LINQ than a specific implementation/API of a data structure.
(Of course, I may be biased in that I work on a commercial product which also has this characteristic.)
edit: After some stalking, I see you work on PowerQuery. Great product! Improving the SQL editing capabilities would make it much better!
Also see this:
"R package to bring forward-piping features ala F#'s |> operator."
That said, the hoops you have to jump through when interfacing with stock R code and dplyr make me think that having an operator that is less generic than `%>%` makes sense -- `X %>% foo()` is equivalent to `foo(X)`; except if it isn't -- `X %>% plot(x~y, data=.)` is not equivalent to `plot(X, x~y, data=.)`. And sometimes you need tricks like `X %>% { plot(.$x, .$y) }` to work around the nonobvious behaviors.
I find most impressive the work that this project did to emulate the lazy evaluation that is a hallmark of dplyr. I wasn't aware that python had anything like this ability.