HNHacker News
TopNewBestAskShowJobs

closed

1,274 karma · joined November 19, 2013

Cognitive psychologist / data scientist!

Working on siuba, a data analysis tool for python:

https://github.com/machow/siuba

submissionscomments
closed··on GitHub Projects – Customizable, flexible tool for planning and tracking work
I've been working on using github projects beta for the past couple months. I'm super excited for it, but right now it's hard to use because because it lacks basic features.

Things projects have that beta projects don't:

* No easy way to automatically add issues from a repo to your project. You have to add them one by one, or use the graphql API, which lacks batch operations.

* No issue preview on the board.

Basic feature missing from sheet view:

* You can't sort issues along the dimensions available in a repo's issue tab (e.g. newest, oldest, recently updated).

I thought I must have been missing something, until I realized that github's public roadmap repo itself uses a bot to add issues [0]. There's an issue there listing out the parity items [1], which will be super handy!

[0]: https://github.com/orgs/github/projects/4247/views/1

[1]: https://github.com/github/roadmap/issues/287

closed··on The Problems with Deliberate Practice (2020)
> What I’m willing to bet, however, is that you’ve heard of deliberate practice in the context of Malcolm Gladwell’s ‘10,000 hour rule’ — the mistaken notion that 10,000 hours of practice would turn anyone, at any age, for any skill, into a master practitioner.

Ericsson spent a lot of time and page space trying to distance himself from this "rule", but it's clear from his writing that he introduced it.

See his 2007 HBR article (which lists it as a necessary condition):

> By now it will be clear that it takes time to become an expert. Our research shows that even the most gifted performers need a minimum of ten years (or 10,000 hours) of intense training before they win international competitions.

It's a confusing enough situation that Macnamara et al review the many places he discusses 10k hours (or 10 years) in their book chapter The Deliberate Practice View.

closed··on The Rule of Three
I think the quote flouting the rule of 3, and the discussion in the article is analogous to a similar phenomenon with face attractiveness.

If you average across many faces, you get an attractive face. It's a safe bet for an attractive face.

Yet, many super models are not an average across features, but have extreme features. The challenge is that unlike an average, extremes can generate unattractive faces.

Rule of 3 is a recipe for a pursuasive message, but you can be pursuasive by flouting it (and extremely pursuasive writing often flouts rules).

https://en.m.wikipedia.org/wiki/Averageness

closed··on How to Test S3 in Python (2020)
On my laptop each run of boto3.client("s3") takes about 4 milliseconds. I'm guessing this is okay for most crud apps..!

Edit: especially compared to any interactions with s3

closed··on Human Memory
Weirdly enough, Momento is a big hit with a good chunk of memory researchers, and had been used as a stimulus in memory experiments!

https://www.aalto.fi/en/news/film-memento-helped-uncover-how...

closed··on Making better decisions with the Brier score
The article mentions Brier score is just mean squared error, so it's connected to binomial through that (e.g. where correct prediction is 1, incorrect is 0, it is the mean of the binomial).
closed··on Programming and Writing
> Bad code doesn't work.

> It's very hard, and maybe impossible, to determine if a novel "works".

I wonder if there is a an assumption here about what it means to "work" vs to be "bad". Psychologically, maybe it's helpful to view each person as their own interpreter, and there is more variation there compared to (e.g.) a specific python interpreter.

But even in python I could write totally not-python code, and have the python interpreter run it (e.g. by writing a codec). And I could write beautiful code that throws an error, and have a person debug it for <some_purpose>, and in meeting that purpose it might be working.

I think the challenge here is that "work" is being defined in a narrow, technical sense for code, but is recognized in a much broader, social/cognitive sense, for novels!

closed··on Practical SQL for Data Analysis
For what it's worth, I maintain a library called siuba that lets you generate SQL code from pandas methods.

It's crazy to me how people use SELECT * -> pandas, but also how people in SQL type a ton of code over and over.

https://github.com/machow/siuba

closed··on Practical SQL for Data Analysis
In case you're interested in what's missing--I maintain a port of dplyr from R to python called siuba, and gave a talk recently on why pandas might be hard to use:

https://www.rstudio.com/resources/rstudioglobal-2021/bringin...

closed··on Data Organization in Spreadsheets (2017)
Is there a specific aspect of pandas research you're interested in? There are a lot of useful guides around table-based workflows that might be helpful :).

I would start w/ different strategies on how to model data in tables. One problem that I often see in pandas data analyses is people treating the data like it's a web app database (many small, normalized tables), rather than joining the data into a few big, denormalized tables. The latter makes it easier for people to answer their own questions / vs relying on a bunch of tiny custom functions someone wrote!

* Hadley's tidy data paper: https://vita.had.co.nz/papers/tidy-data.pdf

* Normalizing data: https://en.wikipedia.org/wiki/Database_normalization

* Denormalized data: https://en.wikipedia.org/wiki/Denormalization

* Emily Riederer, column names as contracts: https://emilyriederer.netlify.app/post/column-name-contracts...

closed··on Some opinionated thoughts on SQL databases
What you described is pretty similar to my experience.

I wonder if part of the author's sentiment about it being more widely deployed can be explained in part by stack overflow trends data.

Basically, MySQL used to make up a much larger % of questions on the site (compared to postgres). But postgres had been growing, and MySQL shrinking, so now they're about the same on the site.

https://insights.stackoverflow.com/trends?tags=mysql%2Cpostg...

closed··on Fuckin' user interface design, I swear
I'm surprised every time someone says:

1: This field requires a lot skill.

2: I am not skilled in this field.

3: It's obvious to me <observation that might require skill in domain>.

And then there's no appeal to an expert in the field, or observed behavior.

It's fair that sometimes negative effects are obvious, and if the writer was observing the deleterious effects of the button placement, I could see where they're coming from.

closed··on Open source projects should run office hours
It seems like the bulk of OSS developers I know do not get paid, but are obsessed with a particular problem domain (or have essentially merged with their tool and become a finely tuned cyborg).
closed··on 10 Years of Open-Source Visualization: Did I learn anything from D3.js?
Have you tried the python port of ggplot, plotnine? I screencast doing live data analyses with jupyter notebooks, virtualenv, and plotnine. I'm definitely much quicker in R, but it's not too bad!

https://youtu.be/z6xNKZZMWgU

closed··on Migrating to SQLAlchemy 2.0
I've been building data analysis tools on top of SQLAlchemy's declarative system over the past couple years. It's got to be the most well documented, carefully designed library I've ever interacted it :).

It looks like most of the changes in 2.0 are aimed at the ORM system, which makes sense. I think a lot of complaints that come up have more to do with the complexity of interacting with a SQL database, so appreciate the effort in the docs not just laying out an API, but essentially educating around the problem domain.

closed··on Why SELECT * is bad for SQL performance (2020)
I'm seeing a lot of comments on why SELECT * is fine or not, but it seems like the bigger issue is that (in general) SQL has two very limited ways of letting your select columns: explicitly naming each column, or getting all columns. This means that when you are concise, you sometimes end up getting columns you don't need.

For example, R's dbplyr library lets you write queries that select all columns that start_with("something_"). This is revolutionary, because now a person can use the semantics of column names in their selection!

Granted, dbplyr ultimately generates a query that explicitly names the columns, so it's not a within-SQL solution, but I've been surprised at how useful the behavior is!

closed··on Implementing the Elo Rating System (2020)
It's worth noting that the general models ELO approximates, called item response theory (IRT) models, could handle larger than 1v1.

Its similar to how you model situations where multiple latent skills contribute to performance (multidimensional IRT). In practice, it's a way easier though to use simple stuff like averaging / constraining groups so everyone has roughly the same skill level.

closed··on Architecture.md
I usually reach for a friend, or someone I've met before, since using the first version of a doc is asking a lot! (And they're often part of the target audience).
closed··on Architecture.md
I love architecture docs, but find they're often written using a funny process:

  1. Spend a long time writing the doc.
  2. Wait for a person to chance upon it.
  3. Hope you anticipated their questions.
It seems like the most important thing a person can do is reverse this:

  1. Say who the doc is for.
  2. Find that person. Ask them to try a lil contribution.
  3. Frantically write / revise the doc.
IMO it's a lot like creating a presentation. The earlier the feedback the better!
closed··on R Markdown: The Definitive Guide
I'm familiar with both Rmarkdown and Jupyter Book. Rmarkdown also uses pandoc. Both are very flexible.
closed··on How to select best UPS and Solar to be off-grid 24 hour for 2kw
AFAIK most prepackeged UPS devices are a battery and inverter. My hot take is that if you need to run this device continuously it will be a substantial (& custom?) build. The biggest factor is how portable it needs to be.

Overview: Most solar outputs 12 or 24v, so what often happens is solar -> controller -> battery (or bus) -> inverter -> your device.

Batteries: Many batteries are rated for 100ah at 12 volts (so 1.2 kw hours). You would need 2+ of them to run the system purely off them for about an hour. You can buy much beefier batteries designed for residential homes (e.g. 3' x 2' x 1' in size, 10kwh, $5,000+ for lithium, etc..).

Solar: The panels you might see on top of a camper often generate 100 watts each, so you would need 20+ of them to continuously sustain your device. Can be 250 watts or a bit more, so 8+ panels.

I've only worked on electrical for a camper, so could be missing other options! AFAIK your requirements aren't crazy for residential solar setups.

closed··on Plotnine: Grammar of Graphics for Python (2019)
One critical piece methods miss is they can't decentralize contribution. For example, the gganimate package in R gives user new ggplot functions. With a `+` users can use functions from any package, so the gganimate approach works.

With method chaining the gganimate author would have to mutate some class, and users would have to load all methods (vs importing what you need).

closed··on Plotnine: Grammar of Graphics for Python (2019)
In case anyone is wondering what the big deal is with ggplot / plotnine, I record myself doing hour long data analyses in python with it!

I've noticed that a lot of bootcamp grads can use matplotlib to do very simple plots, but when it comes to iterating on a data analysis (trying different plots, facetting on variables, etc..), they get tripped up quickly.

I'm trying to use a port of dplyr I'm working on (siuba) and plotnine to show what on-the-fly analyses might look like. Can't speak highly enough of plotnine!

https://youtu.be/z6xNKZZMWgU

closed··on OO in Python is mostly pointless
> Nor do I find charming the belligerent lack of any magical syntactic sugar for `self`. Does Python force you to pass it as an argument to make some kind of clever point?

On the class, you can call the method like a normal function (passing the self arg manually). Seems like a nice connection to just raw functions. Also explains how you can add methods after a class has been defined (set a function on it, whose first param is a class instance).

closed··on Siuba – A Dplyr Port to Python
I'm still debating chaining vs piping, but you can do..

  from siuba import _
  from siuba.data import mtcars
  
  # mtcars is a pandas DataFrame

  mtcars \
    .groupby("cyl") \
    .siu_summarize(avg_hp=_.hp.mean())
closed··on Siuba – A Dplyr Port to Python
Using the experimental fast grouped pandas functions, it should run at the speed of optimized pandas code!

Since siuba functions just run on pandas DataFrames, you can always hand tune for performance, but imo most of the time pandas code runs slow it's because of something like .agg(lambda ...) somewhere.

There's an example with timings here:

https://siuba.readthedocs.io/en/latest/developer/pandas-grou...

closed··on Siuba – A Dplyr Port to Python
Hey, creator of siuba here. I think siuba's big advantage is that it can generate SQL code.

The architecture necessary to pull off executing either pandas or SQL also makes it very extensible (e.g. to spark or dask in the future :).

https://siuba.readthedocs.io/en/latest/key_features.html

closed··on Siuba – A Dplyr Port to Python
siuba does the SQL translation :). pandas is used for local data, since it does a lot of optimization in c++, similar to dplyr's low level code.

Thanks for bringing this up--the docs could be clearer here

closed··on Siuba – A Dplyr Port to Python
Siuba uses type hints to dispatch the appropriate versions of custom functions!

For example, siuba allows users to create custom functions using a thin wrapper around functools.singledispatch.

When deciding how to run...

    df >> filter(my_custom_func(_.x))
It requires that the return type be compatible with the backend being run (e.g. pandas, a SQL dialect).

Would be super interesting to try and lay out what would be needed to do static analysis via mypy. I think it'd require some plugins for singledispatch at least, probably some reworking things in ways myoy expects.

https://nbviewer.jupyter.org/github/machow/siuba/blob/master...

closed··on Siuba – A Dplyr Port to Python
Hey, thanks for pointing out Self--I definitely need to dig into fastcore more!

One motivation for developing siuba is that the grouped agg you show requires users specify only one operation on one column.

E.g.

1. Calculate mean of x

However, common operations like demeaning a column are multiple operations:

1. Calculate mean of x

2. Subtract result of (1) from x

In siuba you can just write mutate(res = _.x -_.x.mean()). This isn't possible from something like gdf.x.agg("mean"), and from what I can tell deeply confusing to analysts :/.

In vanilla pandas I really like to use the chaining method you laid out, and siuba to me is mostly a utility library for making the approach a little more succinct / performant[1].

siuba has experimental autocompletion (thanks to Tim Mastny!), and there's a pretty hefty technical write up on how it uses IPython machinery for that in siuba's architectural desicion record folder[2].

[1]: https://siuba.readthedocs.io/en/latest/developer/pandas-grou...

[2]: https://github.com/machow/siuba/blob/master/examples/archite...

← PreviousPage 2 of 16Next →