HNHacker News
TopNewBestAskShowJobs

pocketsand

600 karma · joined September 27, 2021

submissionscomments
pocketsand··on Is the Reproducibility Crisis Reproducible?
Things are complicated.

To be fair: Everyone I've worked closely with in research has gone above and beyond not to cut corners and produce high quality data and research.

What I have in mind here is a situation where people are actually quite careful but can still end up in a place where they don't know what happened because they don't have good systems for creating datasets and storing code.

For example, graduate students are not always taught to work in a reproducible way. It's definitely gotten better from what I can see, but it was normal for people to get source data and work that data into its final form in a lot of different steps, but not always reproducible steps. E.g., data comes in from secondary source or other provider. It gets cleaned. That file gets saved as something like "clean data 011234.csv".

More work is done, it gets saved again.

Time passes, things are revisited, and a handful of files exist that likely with some care could lead from point A to point B. But the exact process, to say nothing of the dozens, sometimes hundreds, of small decisions data preparation decisions get lost to memory.

Code doesn't go in version control. People get new computers. USBs get lost. Universities migrate to new data systems and so on.

All the while, these students and researchers were very careful while doing the work. They were just never trained to use good version control and pipeline processes. They basically do what they did with papers they write. Save and backup while working through the paper and move on when it's done.

This is made worse when data is proprietary or not legally shareable.

So people aren't necessarily being shoddy or doing bad work, they're just not using good systems.

pocketsand··on Is the Reproducibility Crisis Reproducible?
I think your critiques are fair enough and I think the authors would likely agree they should have been more proactive with the data sharing.

Maybe it sounds like I'm parsing too much, but I nevertheless still think saying "no one knows what happened" is unfair. They know, they shared, and justified what they did when called on it with literally almost no delay. I agree they shouldn't need to be called on it.

Anyone who's advised students or asked even presenting researchers such questions know that often people will literally not know what happened to all their data.

pocketsand··on Is the Reproducibility Crisis Reproducible?
I'll start by saying that if you're going to roast a paper that an econ nobel winner and one of the most famous and respected working statisticians put their names on, you probably want to turn down the volume and double check your claims a little more before hitting "post."

A z score is not at all morally equivalent to a p-value. It's just a standardized measure. Converting measures to z-scores aids in interpretation. They also can aid estimation in some cases: using non-standard parameterization in Bayesian analysis is often crucial to get MCMC to accurately sample from the posterior distribution.

Sure, you can take a z score and look at the area under the curve and come up with a p value. But you don't have to. In the referenced paper, they use z scores to be able to standardize the measures in the papers they draw from, so they're comparable.

The author's other critiques of the paper seem reasonable. It's a problem with all meta analyses: the amount of work it takes to correctly interpret publishes papers and then take those results and aggregate them is herculean. To do it to over 20,000 is inevitably going to lead to some mistakes. That said, those mistakes may not be fatal to the analysis.

Moreover, saying "no one knows what happened to those 11,285 studies" without checking in with the authors is completely unfair. The first author responded with the code showing exactly how they achieved that figure. Nothing mysterious.

Andrew Gelman responded in the comments to him, as did the first author. I find their responses convincing.

pocketsand··on Rage: Fast web framework compatible with Rails
HTML, the language on the server
pocketsand··on Ruby on Rails: The Documentary [video]
Correct.
pocketsand··on Lab leak fight casts chill over virology research
It's a bummer to be a scientist whose work is stymied by these concerns.

Nevertheless, pumping the brakes here is sensible.

I get the sense that researchers feel a certain injustice that geopolitics is affecting scientific decisions by their funding bodies. Funding should be decided on the scientific merit of their proposals. In a vacuum, that makes sense.

Seen more widely, though, the decisions of the NIH and similar funding bodies are shot through with political considerations. For example, it would be hard for a physician scientist to attend a conference in the last few years where considerable resources were not devoted to equity issues.

My point here is that we already contextualize what work we fund by political decisions, for better or worse. In this circumstances, we continue to do so, but we in particular pause funding on research that may, in some form, be responsible for the deaths of millions of people and the disruption of lives of billions. That's a sensible choice.

pocketsand··on Causal inference as a blind spot of data scientists
Social sciences haven't ignored causal inference. Perhaps it’s not everywhere you’d like to see it, but it’s common in quant papers, its the backbone of econometrics, and you’d probably have trouble finding a single top ranked PhD program which doesn’t provide at least cursory coverage of the methods.
pocketsand··on Causal inference as a blind spot of data scientists
One gripe with this article—-regression coefficient doesn’t provide ATE under most circumstances using observational data.
pocketsand··on Show HN: A JavaScript function that looks and behaves like a pipe operator
I suppose it's just what you're used to, of course.

However, in R, you don't need to trailing slashes on new lines. Plus, their pipes are hideous: %>% or now |>.

I do kind of like how the pipe delimiter looks on the left.

pocketsand··on Show HN: A JavaScript function that looks and behaves like a pipe operator
Prepending doesn’t help its case. I like the R convention of pipe at end of line, indent on lines 2+
pocketsand··on Causality for Machine Learning (2020)
This seems well done and well-researched. I appreciate the diagrams and art and references to the likes of Rubin and Pearl.

A few headings down:

> Causal inference provides us with tools that allow us to answer the question of why something happens.

This is not necessarily so.

Randomized controlled trials suffer from black box problems the same as models. This is clear enough when thinking about something like a tutoring program. Suppose I randomly assign a bunch of schools to learn algebra with curriculum X and the rest to continue business as usual.

Program X does better, so we infer the program has a causal impact on algebra learning.

However, we still do not know for sure why program X does better, only that it does better. This is important to inform how to take what works about the program and apply it to other circumstances, adapt it, and so on.

I suppose compared to a big data set, we have a better "why" answer to the variation between the outcome and the treatment. The difference being that we actually know the cause of the observed effect with a trial, whereas with correlational analyses we're not so sure. But that's a very deflationary view of "why." I don't mean to be too cynical here; we can always push "real" causality one more level down. For example, suppose we figure out the secret sauce to better algebra teaching relates to a specifical pedagogical practice. We can then say "but why does that practice work? what does it do in the brain?" So I don't want be too reductive.

But even gold standard RCTs don't always give us a "why?" answer. I remember attending a conference about a decade ago among causal inference-devoted social researchers specifically about "the black box" of causal inference as it pertains to RCTs.

pocketsand··on We Are Not Just Polarized. We Are Traumatized
Counterpoint: We are not traumatized, we are polarized.

You can layer paragraph upon paragraph about how things are bad, but they've often been far worse, and far more tumultuous. Segments of post-civil war, post-WW1, post-WW2 populations were in fact traumatized, from real loss. As have families from the pandemic. Throughout US history, politics has been fractious and at times violent, even in congress itself.

But no, we -- the average person -- are not traumatized. We do not need therapy. Author starts by winking at the self-medicalizing seen on social, trying to separate her serious argument from the frivolous ones on "TraumaTok." But then just dives right in and does the same thing, but with lots of words.

pocketsand··on California moves to silence Stanford researchers who got data to study education
Just for context. I'm a PhD trained in education research who has met Sean Reardon a handful of times, had a meal with him, gone through methods training with him. He sits at the top of the field and has the unconditional respect of nearly everyone for his methodological rigor.

This is not a guy who shoots from the hip.

pocketsand··on UCLA professor refuses to cover for Dan Ariely in issue of data provenance
Check out the datacolada web site. https://datacolada.org/
pocketsand··on UCLA professor refuses to cover for Dan Ariely in issue of data provenance
Pre-registration protects against HARKing and p-hacking, but it unfortunately doesn't protect at all against outright fraud, which we appear to see from the likes of Ariely and Gino.
pocketsand··on UCLA professor refuses to cover for Dan Ariely in issue of data provenance
Seems like behavioral economics really just receives splash damage from what is the massive fraud at the heart of experimental social psychology. At this point I basically assume any catchy finding from the field is at best HARKed or p-hacked and at worst completely fraudulent.
pocketsand··on What is it like to be a bat? (1974) [pdf]
This is an interesting idea, but I don't find it particularly germane to Nagel's question. To someone with a hammer, everything looks like a nail. To someone with an LLM, everything looks like a set of data to be trained on, I suppose.suppose.
pocketsand··on Leader of Online Group Where Secret Documents Leaked Is Air National Guardsman
The dirty not-so-secret about national defense apparatus in this country is that there are thousands of people in these various sub-sub-wings of our military along side mediocre contractors cashing in on defense contracts who go through rudimentary security clearance checks and have access to all this information.

I grew up outside DC and every so often you hear of the security clearance spooks interviewing you or someone you know about some jamoke you went to high school with to determine if they're a security risk. They also ask those people really difficult questions, like, "Would you use drugs at work?" "If you did, would you download illegal documents?" "If you did get high at work and download illegal documents, would you post them on Myspace?" About half of the people actually fail to answer those questions in the expected way. The other half make 200,000 dollars a year from a subcontractor of a subcontractor of Northrop Grumman pushing paper at a desk all year.

It's a national disgrace.

pocketsand··on I created an eBay account and bought an item, today I got indefinitely suspended
And their signals may well be noisy.

This is why it is infuriating. They have some model and some threshold. If an account goes over the threshold, they're banned.

For whatever reason, OP ends up over the threshold, and eBay is completely unwilling to put any customer service effort into reviewing whether they model may have misclassified OP's account. Having the service be completely off limits for a certain % of companies because it's too hard to assess errors is part of the business strategy, and it's an absolutely awful experience for the small number of people it affects.

Luckily, OP can probably live without eBay. Now imagine what happens in scenarios like these, which show up every few weeks on HN and other news media.

1. Sellers who make a living off sites like eBay who have their livelihood thrashed when model goes Beep on their account for some reason.

2. Businesses (and their customers) that rely upon cloud services that have their accounts closed indefinitely for unspecified reasons.

3. People locked out of their Gmail/similar accounts for unspecified reasons, who then lose access to every account tied to those accounts.

If you've ever been in one of these situations, it turns into a Kafkaesque dystopian runaround at best, and complete blackout at worst. I ended upon on the wrong end of an indefinite suspension from Google. I had to appeal a dozen times with no insight from them. I ultimately tracked it down to an off-by-one error on a payment card zip code. I had moved across town and my zipcode went up by one. Google didn't like that and I was a threat until I figured it out. Took months.

pocketsand··on Correlation between the use of swearwords and code quality in open source code? [pdf]
Takes a lot of work to advise theses, and Bachelor's theses are at bottom of list, but this student would have been served well by being encouraged to move all these page-long sections explaining terms and statistical tests to an appendix, or just replaced with a brief description and citation.

Nevertheless, still pretty good thesis for a Bachelors student. I taught statistics to both undergrads and grads in social science and had very students who showed this level of curiosity into methods and then did the work to apply the methods.

pocketsand··on PHP in 2023
When I work in Django, instead of other frameworks I use, I think "I like this more, I like this less; this works better for me in these circumstances, this does not." But I've never felt it made any sense at all to pronounce one as objectively better than another.

In most cases, it is a matter of taste and tradeoffs that vary based upon circumstance.

You can keep saying "this is the correct way" with all the conviction in the world, but it won't turn matters of opinion and circumstance into universal fact.

pocketsand··on PHP in 2023
Technology companies routinely use/license technology from other companies in their products. If, for example, Apple relies upon Samsung for displays and memory in their laptops, it does not follow that you're actually buying a Samsung laptop.

In any event, Laravel has never been shy about its debt to Symfony.

pocketsand··on PHP in 2023
You can look at a Laravel (Eloquent) model and easily see what methods it has. Nothing is hidden -- at least no more than any ORM that inherits from another class which gives it functionality.

As for properties -- yes, they're not explicitly described but mapped to the database schema. Other ORMs, like that in Ruby on Rails, also do not explicitly set model properties and just map to database columns. People have been upset about this forever, but everyone else has happily used Active Record and had no major issues, all while enjoying not having to add a dozen lines of code to annotate props.

In both Rails and Laravel you can define accessors and setters, either overwriting the magic property for the column or defining new properties.

If not seeing the properties on the model is a huge problem, there are plugins to generate them and put them in the model. My IDE gives me direct insight by inspecting the table.

I'm happy not to have to write out all the props twice when there's no tangible performance issue and I've never had any other problems with dynamically getting the model attributes from the schema.

You're taking a matter of taste and context-based tradeoffs and acting like those reflect unchallengeable principles.

Symfony is a great piece of software. So is Laravel. Github numbers alone will show you how compelling so many developers find Laravel. Of course, there will be haters who think they're all stupid, uneducated, bad, idiotic, lazy, etc.

Everyone makes trade-offs based upon taste or prior decisions. Consider Django. There, you explicitly define properties, but link them to the type of database column they'll use. Then, you can generate a migration based upon the model, and then run that migration. I really like that pattern. Removes tedium of writing migrations, preserves ability to modify the migrations, and makes things explicit.

When I used SQL Alchemy in Python, it completely allowed me to not write migrations, only models. That was very nice, until it wasn't. But then I just had to do a few things manually -- a reasonable price to pay for saving a bunch of time in the other 19/20 of cases.

These are all just different ways of getting to the same place. Each has benefits, each has drawbacks, the extent of each which will depend on the project and the developers' tastes.

pocketsand··on PHP in 2023
A lot of stuff is deprecated and not removed.

Can you give an example of the footguns you so fear?

pocketsand··on PHP in 2023
Can you give me a few real life examples where you got bit by the "magic" conventions or ran into any serious issue from not explicitly setting properties on your models?

Facades are not global glorified variables with magic methods. Usually they're just a way to instantiate and access a class with less code. People complain that facades limit testability. In fact, they're easily mocked. Etc.

The docs do a very good job of explaining how things work and how you can override any "magic" you may not like. If something is doing something you didn't expect, you probably didn't read the docs. In fact, a delight of working with Laravel is it's so intuitive that you can often guess how to achieve something and be right.

Have you worked with Laravel in practice?

pocketsand··on PHP in 2023
I work on a project like this and you'll run into some papercuts, but it's not hard.

You can take a model for anything in your legacy code, and just set properties on the model to teach Laravel how to deal with things that don't follow conventions.

For example, you can override the table name. Suppose you had a legacy table called "postdata". In Laravel, this table would be called posts, and a model would be Post.

So you'd just tell Post to query postdata rather than posts.

Likewise, you can tell it which fields to cast as dates, booleans, etc. You can tell it which fields represent timestamps for updating, creating, deleting and so forth.

You can setup relations that don't follow naming conventions the same way.

The only issue I've run into is doing things like setting up relationships across databases. You can do it, it's just not super straightforward.

I think you'll find it's a breath of fresh air to have a modern wrapper around your legacy code, and that your legacy stuff will then fit in quite nicely with your new work.

pocketsand··on Will You Help Me Repair My Door [video]
Because I got high made a ton of money, and he would have made a lot more if 9/11 didn't happen. There's a nice documentary on him and that song. Seems like a genuine guy.
pocketsand··on Will You Help Me Repair My Door [video]
A song that laments how 911 ignores poor/black neighborhoods embodies a different sentiment than "fuck the police."
pocketsand··on The beauty of CGI and simple design
For web apps, PHP is still a great candidate. Check out Laravel or Symfony. It's not like the old days of PHP files mixed with procedural code. Although those were good days :)
pocketsand··on WeWork’s once robust cash reserves have dwindled, raising chances of default
Honestly thought they were already bankrupt
← PreviousPage 2 of 3Next →