HNHacker News
TopNewBestAskShowJobs

uniqueuid

3,775 karma · joined June 17, 2021

Computational social scientist.
submissionscomments
uniqueuid··on Brain biopsies on 'vulnerable' patients at Mt Sinai set off alarm bells at FDA
A big part of the problem is that as patients, we observe single discrete outcomes.

For doctors an aggregate 1% chance is low, but the realization of the outcome in a single person is not a percentage, its binary bad side effects or not.

Admittedly, doctors could usually do better in communicating risks (e.g. see the famous example of 1/3 risk of impotence, which patients interpret either as 1 out of 3 people, 1 out of 3 times having sex, or 1/3 of the time of sex).

uniqueuid··on 2.4M person study: Internet use boosts, not hurts, well being
Great take at the end of this piece:

"“The study cannot contribute to the recent debate on whether or not social-media use is harmful, or whether or not smartphones should be banned at schools,” because the study was not designed to answer these questions, says Tobias Dienlin, who studies how social media affects well-being at the University of Vienna. “Different channels and uses of the Internet have vastly different effects on well-being outcomes,” he says."

uniqueuid··on Show HN: "data-to-paper" – autonomous stepwise LLM-driven research
True, we seem to have a pretty similar perspective after all.

My concern is an ecological one within science, and your argument addresses the frontier of scientific methods.

I am sure both are compatible. One interesting question is what instruments are suitable to reduce negative externalities from bad actors. Pre-registration works, but is limited to few fields where the stakes are high. We will probably similarly see a staggered approach with more restrictive methods in some fields and less restrictive ones in others.

That said, there remain many problems to think about: E.g. what happens to meta-analyses if the majority of findings comes from the same mechanism? Will humans be able to resist the pull of easy AI suggestions and instead think hard where they should? Are there sensible mechanisms for enforcing transparency? Will these trends bring us back to a world in which trust was only based on prestige of known names?

Interesting times, certainly.

uniqueuid··on Show HN: "data-to-paper" – autonomous stepwise LLM-driven research
Sorry but I think we have very different perspectives here.

I assume you mean that LLMs can generate new insights in the sense of producing plausible results from new data or in the sense of producing plausible but previously unknown results from old data.

Both these things are definitely possible, but they are not necessarily (and in fact often not) good science.

Insights in science are not rare. There are trillions of plausible insights, and all can be backed by data. The real problem is the reverse: Finding a meaningful and useful finding in a sea of billion other ones.

LLMs learn from past data, and that means they will have more support for "boring", i.e. conventional hypotheses, which have precedent in training material. So I assume that while they can come up with novel hypotheses and results, these results will probably tend to conform to a (statistically defined) paradigm of past findings.

When they produce novel hypotheses or findings, it is unlikely that they will create genuinely meaningful AND true insights. Because if you randomly generate new ideas, almost all of them are wrong (see the papers I linked).

So in essence, LLMs should have a hard time doing real science, because real science is the complex task of finding unlikely, true, and interesting things.

uniqueuid··on Show HN: "data-to-paper" – autonomous stepwise LLM-driven research
Thanks, that's a nice quote.

With regard to the debate, I think it's good not to engage in too much black-and-white thinking. Science itself is a pretty muddy affair, and we still haven't grown beyond simplistic null hypothesis significance testing (NHST), even decades after its problematic implications became clear.

That's why it's so important to look at the macro implications: I.e. how does this shift costs? As another comment nicely put it, LLMs are empowering good science, but they are potentially empowering bad science at an order of magnitude more.

uniqueuid··on Show HN: "data-to-paper" – autonomous stepwise LLM-driven research
Hi,

thanks for the honest and thoughtful discussion you are conducting here. Comments tend to be simplistic and it's great to see that you raise the bar by addressing criticism and questions in earnest!

That said, I think the fundamental problem of such tools is unsolvable: Out of all possible analytical designs, they create boring existing results at best, and wrong results (i.e. missing confounders, misunderstanding context ...) as the worst outcome. They also pollute science with harmful findings that lack meaning in the context of a field.

These issues have been well-known for about ten years and are explained excellently e.g in papers such as [1].

There is really one way to guard against bad science today, and that is true pre-registration. And that is something which LLMs fundamentally cannot do.

So while tools such as data-to-paper may be helpful, they can only be so in the context of pre-registered hypotheses where they follow a path pre-defined by humans before collecting data.

[1] http://www.stat.columbia.edu/~gelman/research/unpublished/p_...

uniqueuid··on Show HN: "data-to-paper" – autonomous stepwise LLM-driven research
With all the positive comments here, I feel like someone should play the role of the downer.

First of all, it's inevitable that LLMs will be/are used in this way and it's great to see development and discussion in the open! That's really important.

Secondly, this will absolutely destroy some areas of science even more than they have already been.

Why? First, science as all of humankind is always a balance between benevolent and malevolent actors. Science already battles data forgery, p-hacking and replication issues. Giving researchers access to tools like this will mean that some conventional quality assurance processes will fail hard. Double-blind peer review will no longer work when there are 10:1 or 100:1 AI generated to high-quality submissions.

Second, doing analysis and writing a paper is one bottleneck of science, but epistemologically, it's not the important one. There are innumerable ways to analyze extant data and it's completely moot to do any analysis in this way. Simmons, Nelson and Simonsohn / Gelman et al. etc have shown: Given a dataset, (1) the findings you can get are practically always from very negative effects to very positive effects, depending on the setup of the analysis. So having one analysis is pointless, especially without theory. (2) even when you give really good labs the same data and question, almost nobody will get the same result (many labs experiment).

What does this tell us? There are a few parts of science that are extremely important and without them science is not only low-impact, it even has a harmful effect by creating costs for pruning and distilling findings. The really important part are causal analyses, and they practically always involve data collection. That's why sciences with strong experimental traditions fare a bit better - when you need to run a costly experiment yourself in order to publish a paper, this creates a strong incentive to think things through and do high-impact research.

So yeah, we've seen this coming and it must create a big backlash that prevents this kind of research from being published, even if vetted humans.

Source: am a scientist, am a journal editor.

uniqueuid··on Temporal Python – A durable, distributed asyncio event loop (2023)
The thing with event loops in python is that they are not a single, all-governing scheduler (as e.g. in the BEAM).

ev loops instead are a mid-layer concept that sits below other infrastructure such as threads and processes. And (perhaps somewhat frustratingly) it is not too uncommon to have multiple ev loops in parallel. See for example the proxy.py project, which offers to run one async loop per process for a speedup.

As a result, there are some incentives to swap out the loop itself, e.g. for faster implementations like uvloop, because they are somewhat pluggable anyways.

uniqueuid··on Ask HN: Interesting TUIs (text user interfaces), maybe forgotten ones?
These are neither pretty nor good examples ... but teletext [1] and the french minitel [2] surely were interesting

[1] https://en.wikipedia.org/wiki/Teletext

[2] https://en.wikipedia.org/wiki/Minitel

uniqueuid··on NYU professors who defended vaping didn't disclose ties to Juul
Fair question! We would need a RCT with controls for that, perhaps one exists :)
uniqueuid··on NYU professors who defended vaping didn't disclose ties to Juul
Thanks, I take the ad-hoc theory back!
uniqueuid··on NYU professors who defended vaping didn't disclose ties to Juul
You're right that these studies measure the product, not the substance. But TFA is about vaping, and I am referring to that.

[edit] Seems that nicotine is not really harmless, even in gum or patches: double the risk of heart palpitations and chest pains, 1.67 the risk of nausea and vomiting, 1.5 times the risk of gastrointestinal complaints, and 1.4 times the risk of insomnia.

Of course, if you measure anything this closely with >100k participants in RCTs, it's bound to have a significant effect. But these odds ratios are not exactly small.

https://link.springer.com/article/10.1186/1617-9625-8-8

uniqueuid··on NYU professors who defended vaping didn't disclose ties to Juul
That's a very interesting, but completely ad-hoc theory.

The problem with these is that there are almost always equally plausible theories, and it's really hard to figure out which is correct.

Especially given the huge amount of effort that has gone into research on tobacco and nicotine (I guess millions of hours, if not more), I would definitely NOT say we are too quick to dismiss the possibility.

I'm sure you can find an interesting econometric paper presenting this argument, but then it would be more credible.

uniqueuid··on NYU professors who defended vaping didn't disclose ties to Juul
There are many studies on the effect of nicotine. I would look for several meta-analyses or umbrella (meta-meta) analyses. Ideally reporting all-cause mortality.

Here is one like that: ~500k individuals, vaping has an odds ratio of 1.33 (95% CI = 1.14–1.56) for heart attacks compared to non-smoker/non-vapers.

https://www.sciencedirect.com/science/article/abs/pii/S01675...

[edit] To compare that to caffeine, this study does not see a significant risk increase for coffee drinkers. So it seems that vaping is objectively more harmful:

https://www.ncbi.nlm.nih.gov/pmc/articles/PMC5940396/

[edit2] And an umbrella review, because I like them a lot: This shows that negative effects of high coffee intake mostly disappear when controlling for smoking. Note this paper has a correction, which I think is a great sign of credibility. https://www.bmj.com/content/359/bmj.j5024

uniqueuid··on Ask HN: Examples of lovely user interfaces on and off screen
Audio stuff typically has really interesting interfaces and tries to simplify very complex tasks. So look at DAW / synthesizer software, hardware (ableton, teenage engineering) and software/hardware combinations.
uniqueuid··on AI Regulation Is Unsafe
Among the things that the article omits is the question of upholding the existing canon of laws in the face of AI.

As far as I see, much of European AI regulation tries to contain fallout where existing laws are under threat to be undermined (i.e. authorship, safeguarding public opinion, competition, harm/safety etc.)

uniqueuid··on Ask HN: Please recommend how to manage personal serverss
Opinionated take: "a couple of home servers" are probably the wrong solution to your problem. Almost everything you do as a private person works better, faster, more reliably, and with much less time investment if you use a single machine for it.
uniqueuid··on Hierarchical Clustering
An excellent comment, but I would stress that topic models such as LDA and stm are not ordinary clustering methods. They are latent variable models where documents represent a mixture of latent topics.
uniqueuid··on Ask HN: What software sparks joy when using?
What consistently amazes and pleases me is software that cleverly uses algorithms and data structures to prevent complexity. In particular, everything CRDT-like still seems like magic to me.

So I'm unreasonably excited by bittorrent, storage based on consistent hashing, and rsync (rolling checksums are just amazing).

Interestingly I hate distributed consensus with a passion because it always seems to achieve the opposite - causing more complexity.

uniqueuid··on Python Chart Examples
Wow, these actually impressed me. I've made a few matplotlib charts myself, but typically R/ggplot is significantly nicer and easier.

My takeaway is that matplotlib may be a good alternative for specific cases where you need much more control.

uniqueuid··on Psychiatric risks for worsened mental health after psychedelic use
That is true, but there are also conditions where low-powered studies degrade meta-analyses. For example, if the panel attrition is caused by self-selection on unobserved variables, then the study contains bias that is not (and cannot be) accounted for, which skews the meta analysis. But yeah, those concerns are mostly about measurement not sample size.
uniqueuid··on Psychiatric risks for worsened mental health after psychedelic use
I find the design and results somewhat underwhelming (weak significance in a N=800 sample with >50% dropout). The authors also mention that they lack power, but the only correct answer then is not to do the study, include more people, more data points, or reduce measurement variance.

That said, I think it's a refreshing perspective, and it seems like a good idea to look at effect heterogeneity in psychedelics use. Thinks like bayesian trees (BART) can help identify subgroups that react differently to stimuli along specific characteristics.

uniqueuid··on Hacked Nvidia 4090 GPU driver to enable P2P
Not op, but I found this benchmark of whisper large-v3 interesting [1]. It includes the cloud provider's pricing per gpu, so you can directly calculate break-even timing.

Of course, if you use different models, training, fine tuning etc. the benchmarks will differ depending on ram, support of fp8 etc.

[1] https://blog.salad.com/whisper-large-v3/

uniqueuid··on Hacked Nvidia 4090 GPU driver to enable P2P
And before that, Nvidia made our lives harder by phasing out blower-style designs in consumer cards that we could put in servers. In my lab, I'd take a card for 1/4 the price that has half the MTBF over a card for full price anytime.
uniqueuid··on The darker side of being a doctor (2017)
Sure, it's very easy! Just do things that prevent burnout (https://www.ncbi.nlm.nih.gov/pmc/articles/PMC8834764/) and the job gets more attractive, drawing more students and producing less attrition:

- Valuing work gives meaning (money, appreciation)

- Autonomy gives feeling of control

- Managing burden prevents overwork, exhaustion and fatigue

uniqueuid··on The darker side of being a doctor (2017)
Good point, and you can turn it around: Doctors are never "finished". They could always do more to help patients. So in contrast to aviation, where there is a clear corridor of things to do, doctors have no natural upper bound on their work.
uniqueuid··on The darker side of being a doctor (2017)
Thanks for the perspective. I have doctor friends and everyone seems to just blindly accept that some jobs are not jobs, they are identities. You never pause being a doctor. But your comment shows that even in life-critical environments, we do have ways of organizing work so individuals can bear it. Fire brigades are another example.

Time to push for a change. And time to call some people and ask whether they are truly ok right now.

uniqueuid··on Estimating association between Facebook adoption and well-being in 72 countries
Yeah that's a great start! And at the same time, the social network, i.e. friend network of people changes. In terms of size, composition, people are moving as cohorts, but adoption is very different across age brackets, and there will be an age drift in adoption as the social network grows ... so there are just a ton of very biasing drifts going on.
uniqueuid··on Estimating association between Facebook adoption and well-being in 72 countries
Two things jump out immediately:

- a bayesian mixed model is just what I would have used here, great choice and a good sign for power

- the effect sizes are definitely very small. I would not overstate the story here.

That said, I also don't believe that facebook is the primary story here, because other findings all suggest tiny effect sizes (see the famous emotional contagion experiment) and because we don't have a clear, convincing theoretical model IMO. There are just so many intra-individual confounders that would need to be included, including some that have dynamic, self-reinforcing effects.

Still, great to have more robust work here.

[edit] PS: just to throw a methodological nitpick out: Some bayesians are going to have a heart attack while reading "the uncertainty cutoff of 97.5% for posterior probabilities".

uniqueuid··on Hip to be square – 70 years of the Citroën H Van (2017)
Yes the SM is awesome, but to be fair, it is more of a Maserati than a Citroen.
← PreviousPage 5 of 27Next →