Cutting out "lines" isn't always the right decision, but it's been helpful to visualize and discuss.
1,274 karma · joined November 19, 2013
Working on siuba, a data analysis tool for python:
https://github.com/machow/siuba
Cutting out "lines" isn't always the right decision, but it's been helpful to visualize and discuss.
I would say later the article shows examples of standardized effect size, so divides by variance. Whether that means effect size can't be a difference in means (and explains the wiki intro comment) I disagree with--see my Lakens comment for why.
> If you care about units, you can talk about the expected value/difference, but that doesn't make that a meaningful effect size. What you need to do in those cases is to, at least, mention both the expected difference and the variance.
This seems like it is begging the question. If you look at the definition and rationale for effect size, both in the wikipedia and in Lakens, a claim this strong is not there. For example, a CI over difference in means is one way to capture what a standardized effect size that uses a specific variance term might be doing (what Lakens describes as the second ES viewpoint: statistical significance, distinguishing between effect and power).
In any event, thanks for discussing--it's been really helpful to think about and revisit this topic!
pd.period_range(date_from, date_to, freq = "D")
AFAICT, a PeriodIndex and DateTimeIndex function mostly the same, and have many of the same methods, except... * DateTimeIndex can't hold dates far in the future
* PeriodIndex can't easily round to the end of a period (e.g. date + 0*MonthEnd() errors)
* PeriodIndex doesn't handle timezones?First paragraph: "Examples of effect sizes include the correlation between two variables,[2] the regression coefficient in a regression, the mean difference, or the ..."
> Who would consider the same expected difference to mean the same effect size when the population distributions are narrow and when they're wide?
For the case where the unit of measurement has a sensible, relevant interpretation (e.g. if a study measured dollar value of two interventions), and attempts to capture uncertainty via CI or another means, I would consider it one meausure of effect size.
The key to understanding effect size is that its most common focus is around making measures over arbitrary scales become scale invariant. But scales are not always arbitrary.
Daniel Lakens has a great article on ES and puts the motivation for calculating them well..
> First, they allow researchers to present the magnitude of the reported effects in a standardized metric which can be understood regardless of the scale that was used to measure the dependent variable. Such standardized effect sizes allow researchers to communicate the practical significance of their results (what are the practical consequences of the findings for daily life), instead of only reporting the statistical significance (how likely is the pattern of results observed in an experiment, given the assumption that there is no effect in the population).
https://www.frontiersin.org/articles/10.3389/fpsyg.2013.0086...
To the degree that the thing measured has an inherently meaningful scale to the researchers (eg sometimes dollars, time), then it is already an effect size measure.
There is some nuance here, since the level on which you might want to interpret something (eg the spread could be part of the value, especially to an individual who will have 1 and not many codebases).
You might also want to render it comparable to other studies that measure something else, and it's unclear how to convert it to your scale, so you use standardized ES measures (eg many meta analyses).
But an important point is that whether something qualifies is really a question of the scale. (What the best ES is for your specific question is another important issue!)
Difference in means can meet the definition of effect size, and are listed as an example right away in the wikipedia article on it [0]! The key is that it capture the phenomena of interest (and generally, the magnitude shouldn't be a function of number of observations). In psychology, where scales often have arbitrarily defined ranges (eg IQ scales), means usually are not useful effect sizes.
The case you list is a separate issue from effect size (probably the violation of distributional assumptions in some implicit or explicit model). Even using difference in means / variance (say, cohen's d), two observations that extreme would make the interpretation of both mean and variance calculations pretty dubious (a problem not solved by dividing them).
how will you know if your note taking strategy is working?
One thing that has helped me is keeping a high-level study journal. In essence, I have a small calendar that I draw in a notebook each month. Every morning, I write down ~3 key things I did / studied the previous day. Then, I reflect back on the past couple weeks.
If you're in school, you don't want to wait until a final test to realize you've got a lot to study. And in class, you're not in a good position to think about the notes you'd want a week from now. By journaling and reflecting, you can explore note-taking from a review oriented mindset, "what kind of notes do I wish I had taken yesterday / last week?"
It's used a lot in R for this reason. You might want to have a function that operates on different models, but there are 60 modeling packages it might be used on.
They might not be filtering out the bad at the expense of the good. But filtering out some of the good in the name of saving the money it would cost to develop / administer more general assessments.
That's the p(interview_capable) piece whereas the trade-off you mention is the conditional probability (also important!).
p(job_capable | not_interview_capable)
That is, it's crazy that an interview could miss so many people qualified for the job.However, I wonder if oftentimes companies are aiming for..
p(job_capable | interview_capable)
If p(job_capable | interview_capable) is high, and p(interview_capable) is pretty good also, then the company will probably get what it's looking for.This means that the author is right to recognize the test is doing a bad job of measuring their job readiness. A reasonable instrument in this case doesn't have to measure everyone's job fitness (whether there are nasty side effects is another big issue).
But there's been one really nice, overriding factor for me: Twitter pals.
I feel like I've gained a few internet ride-or-dies, even if it's mostly us liking / commenting on each other's stuff.
For example, a grouped filter is very cumbersome in pandas.
Interested to hear if you think it gets at the heart of the problem.
In this case, if black becomes the standard for defining appropriate style then it is superceding pep8 as the standard for style (many styles consistent with pep8 are not black outputs, so if black's style becomes mandatory, it is a ruling against these previously accepted alternatives).
I'm not how this is related to not having any knobs to tune. Aren't there many ways for code to be consistent with pep8? Isn't advocating for only one version of pep8-consistent code in essence an attempt to supercede pep8?
Nowadays I just use maybe 10 notes. If I create a new one for a meeting, etc.., I'll copy in the relevant parts and delete the meeting note after.
It seems like having a proliferation of notes has also been an issue on most orgs I've been in :/.
I think at its core, the groupby issue is a really big problem, and am devoting most of this year to working on it. So if you ever want to pair to work on pandas / pandas wrapping libraries send me an email (link in profile)!
As I've worked on a port of dplyr to python over the past year, though, I've realized the dtype issue (like you said), indexes, and GroupBy being difficult are likely connected. Basically,
* dplyr can chop up a dataframe into 50,000 groups and apply arbitrary functions to it--no problem.
* custom pandas grouped applies are very slow
There are basically three reasons for slow pandas apply methods... 1. creating an index for each subgroup is slow (will not be a RangeIndex)
2. initializing a series for each group is slow (mostly due to type inference being re-run; could be avoided)
3. AFAIK more type inference is run when concatenating results
This leads to a world where grouped calculations can't be run using arbitrary expressions (e.g. lambdas), but have to go through specific SeriesGroupBy methods.I wrote a bit on how I tried to work around that, to enable fast dplyr-like syntax over grouped data in python. Would definitely be interested in your take! There are other libraries, like ibis that do a good job with it, too!
https://siuba.readthedocs.io/en/latest/developer/pandas-grou...
* old attitude: why does pandas have to make things so hard
* new attitude: pandas has a crazy difficult job
I think this is most apparent in the functions that decide what "[d]type" a Block--the most basic thing that stores data in pandas--should be.https://github.com/pandas-dev/pandas/blob/4edcc5541ff3f6470f...
And then, for the ubiquitous Object dtype, often figure out which of the many possible more specific types to cast it to.
If you think that is easy, ask yourself what this outputs:
import numpy as np
np.array([np.nan, 'a'])
Lo and behold--it produces an array where the np.nan has been converted to the string "nan".And yet
import pandas as pd
pd.Series([np.nan, "a"])
Knows this, has your back, and does not stringify it.It also has a pathological fixation on when it tries to convert dtypes, since avoiding all the bad conversion outcomes is a relatively time intensive process (compared to e.g. creating a numpy array).
I realize things could be much easier in pandas user facing interface, but really appreciate the sheer amount of effort that has gone into its dtype wrangling.
This is a very minor question (and I am not concerned about risk to participants)--when you say they signed consent "in accordance with our institutional ethics board", are you talking about Microsoft, one of the two universities, or all?
It's a transliteration of the cantonese word for minibus, 小巴 :).
edit: siu (小) means little!
But grouping data is extremely common in data analysis.
Basically, the strategy with grouped data, is taking the loc approach, and sprinkling in a bunch of additional .transform calls. :/
Would love your feedback :)
(Historical caveats about trusting the results of a single training study apply)
http://www.academia.edu/download/36902540/2015_ActaPsych_Mor...
I wonder if for 90% of users this seems like a critical UI feature!
In general, deliberate practice is about how you practice (e.g. focused on most useful things; w/ quick expert feedback), rather than how often.
https://en.wikipedia.org/wiki/Practice_(learning_method)#Del...
> Part of maturing out of that initial stage comes from when you can autonomously take a well defined task, break it down into steps, and execute it.
It seems like this is the scariest part--since often whether a task is well defined feels more like an assumption. In a lot of ways, I think this is the value of social approaches--you have to answer, "how could I convince so-and-so that this is well defined?"
* Who will I show X feature to?
* When will I show it to them?
* Can I show them a draft?
To me, the biggest risk of solo development is not how I manage a todo list, but that I'll build the wrong thing, because I waited to get feedback.
Some things that have helped me a lot...
* set up a time to show someone your progress on feature X before you feel totally ready to.
* ask someone to try and pick up and tweak some layer of your code (or pair with them)
* if you are developing a library--record hour long screencasts where you use it in a realistic way. Prioritize issues where you say, "oh, I should fix that...". Repeat but with another person driving.
I use github projects and a calendar (the calendar appointments feel more important!)
> As an ex-mormon, we were taught deeply all about Joseph, and that he was not a polygamist that started w/ Brigham
Polygamy is taught in the Doctrine and Covenants, which almost every Mormon will carry to church on Sunday..
https://www.churchofjesuschrist.org/study/scriptures/dc-test...
> Revelation given through Joseph Smith the Prophet ... including the eternity of the marriage covenant and the principle of plural marriage.