(caveat emptor: I have a phd in cognitive psychology but am unfamiliar with aphantasia research).
1,274 karma · joined November 19, 2013
Working on siuba, a data analysis tool for python:
https://github.com/machow/siuba
(caveat emptor: I have a phd in cognitive psychology but am unfamiliar with aphantasia research).
(Both has2k1 and I work for Posit, which supports plotnine work, but authoring its guide was mostly an act of passion for me :)
Take two examples of dataframe apis, dplyr and ibis. Both can run on a range of SQL backends because dataframe apis are very similar to SQL DML apis.
Moreover, the SQL translation for tools for pivot_longer in R are a good illustration of complex dynamics dataframe apis can support, that you'd use something like dbt to implement in your SQL models. duckdb allows dynamic column selection in unpivot. But in some SQL dialects this is impossible. dataframe apis -> SQL tools (or dbt) enable them in these dialects.
Here's a comparison:
* Method chaining: `df.pipe(f1, a=1, b=2).pipe(f2, c=1)`
* Pipe syntax: `df |> f1(a=1, b=2) |> f2(c=1)`
Piping generally chains functions, by passing the result of one call into the next (eg result is first argument to the next).
Method chaining, like in Python, can't do this via syntax. Methods live on an object. Pipes work on any function, not just an object's methods (which can only chain to other object methods, not any function whose eg first argument can take that object).
For example, if you access Polars.DataFrame.style it returns a great_tables.GT object. But in a piping world, we wouldn't have had to add a style property that just calls GT() on the data. With a pipe, people would just be able to pipe their DataFrame to GT().
"The space between us: stereotype threat and distance in interracial contexts"
It mentions being run at Stanford, and was pretty popular (Claude Steele discussed in his book Whistling Vivaldi).
https://psycnet.apa.org/doiLanding?doi=10.1037%2F0022-3514.9...
There's something really inspiring from realizing how far back tables go.
[1]: https://posit-dev.github.io/great-tables/blog/design-philoso...
But if the shape drawing process isn't random, I think the author's experience of feeling unable to articulate the rules AND gravitating to a set of behaviors is a good example of procedural memory (implicit vs explicit).
Explicit rules would probably help speed things up, though!
1. Many of the early pioneers in statistics were psychologists.
2. The econ x psych connection is strong (eg econometrics and psychometrics share a lot in common and know of each other)
3. Many of the people I see with math chops trying to do psychology are bad at the philosophy side (eg what is a construct; how do constructs like intelligence get established)
I come from a similar cog psych background as the Bjork Lab, so am a big fan of their research, but books like 10 steps come from instructional design, which is a bit more focused on the big picture (designing a whole course vs individual mechanisms).
I just wanted to say that Rich is the only software developer I know, who when asked to lay out the philosophy of his package, would give you 5,000 years of history on the display of tables. :)
Can't speak highly enough of the inkstitch maintainers, though--really welcoming community!
The underlying study is not about just rate of learning, but rate of learning under favorable conditions.
The article describes where the data comes from:
> In particular, we model learning using 27 datasets with over 1.3 million student performance observations from 6,946 learners in 12 different courses ranging across math, science, and language learning, across educational levels from late elementary to college, and across educational technologies including intelligent tutoring systems, educational games, and online courses
And the authors argue this is favorable learning conditions because of providing things like immediate feedback on errors, etc..
Lots of room for nuance, but "favorable learning conditions" is key here.
You often get a ton of methods / attributes on the child class, which gets hard to keep tabs on. You have to be careful not to use the same names for things. (The worst is that someone uses this broken encapsulation and relies on things across parent classes via the child).
I get that inheritance can go okay for simple cases, but have seen enough chaotic uses of it that I'd just encourage people to write the boilerplate for forwarding methods.
* duckdb example: https://siuba.org/guide/workflows-backends.html#duckdb
* supported methods: https://siuba.org/guide/ops-support-table.html
(And some types are defined in the typeshed so only exist to be imported during type checking; eg the type checker lib itself is a dependency in this case)
I wonder if a challenge for doctests in R is they often have to test larger, more realistic outputs?
For example, in dplyr's mutate doc, one example is this:
starwars %>%
select(name, mass) %>%
mutate(
mass2 = mass * 2,
mass2_squared = mass2 * mass2
)
This example's output is a dataframe with 4 columns and will display first 5 rows.On the other hand in siuba (a port of dplyr to python), I often have to truncate the example output, because it's hard coded in the docstring:
(cars
>> mutate(
cyl2 = _.cyl * 2,
cyl4 = _.cyl2 * 2
)
>> head(2)
)
cyl mpg hp cyl2 cyl4
0 6 21.0 110 12 24
1 6 21.0 110 12 24
It's nice you can see the full example in the docstring in python, but also very handy seeing complex examples on R doc pages:https://dplyr.tidyverse.org/reference/mutate.html#ref-exampl...
https://siuba.readthedocs.io/en/latest/intro.html#Working-wi...
You can use the verbs collect() and show_query() like in dbplyr.
If you DM me on twitter, would love to set up time to hear about your work / walk through siuba!
Right now it supports a decent number of verbs + SQL generation. I tried to break down why R users find pandas difficult in an RStudioConf talk last year[2].
Between siuba and tools like polars and duckdb, I'm hopeful that someone hits the data analysis sweet spot for python in the next couple years.
The trickiest thing to explain to people is that Cantonese is a diaglossika. It's spoken, but to write you would use mandarin. So there aren't a ton of giant written language corpuses to train on.
AFAICT that's why tools like this often use simple heuristics to do word segmentation. I remember going down a deep rabbit hole, before finally just packaging a small tool to do cantonese word segmentation using the A* algorithm:
https://pyscript.net/examples/
(This was presented today as a pycon keynote talk!)
It's definitely related to ecological fallacy in the sense that both underestimate relative error and inflate effect sizes.
There are a lot of problems you encounter when using arrays for data analysis, like some of their funky behavior with strings [0], but it seems like extending arrays, or building new types of numpy arrays would have been better than new data structures like the pandas Series.
(pandas folks thought a lot about these problems so I could be very wrong).
(I haven't really used it, but it looks promising)
Siuba has come a long way since I wrote this, and now can optimize for fast grouped operations!:
* https://github.com/machow/siuba
* https://siuba.readthedocs.io/en/latest/developer/pandas-grou...
* If f() converts grouped data to something ungrouped, then you can't use a similar function f2(f(group))
* If f() returns a grouped object, then you can't do basic operations like f(grouped) + 1, because DataFrameGroupBy, SeriesGroupBy do not define basic operators like addition. Let alone operations against other grouped data.
A lot of this is worked out in siuba now, and this doc explains a bit more:
https://siuba.readthedocs.io/en/latest/developer/pandas-grou...