495 karma · joined October 5, 2015
Yes, I am aware. I was referring to the fact that at this point, user preference optimization is not done with RL, but with other techniques. I was unintentionally being pedantic.
> But there's an even more contrived reason the training set contributes. The vast majority of human writing is confident.
Yes, I think this has more to do with it than user preference optimization, honestly. "Valuable" text (i.e. text that produces a "good" model) for pretraining has the characteristic of being confident. Even models which are not optimized for chat (i.e. definitely no user preference data used to train them) exhibit this characteristic for medical questions (I know, because I've literally tested them for this purpose).
Preference optimization (or even RLVR) might play some small role as well, but it's kind of a "turtles all the way down" type problem.
I was unintentionally being pedantic, because this isn't really done with RL anymore - it doesn't need to be. RL is now typically only used to train reasoning for tasks with a well-defined correct answer (that's what I meant by binary reward) - this is called RLVR (RL with verifiable rewards).
Preference optimization (training the model on user "this response is better than that response" type data) is more often done with something in the same family as DPO (direct preference optimization), which is decidedly not RL.
Your philosophical concerns are correct of course. And there's the added caveat that the models that most people are using are closed, so we don't actually know their training recipes for sure.
This makes it sound like RL rewards a confident tone -- in general, I don't think this is true (most RL is RLVR, which typically uses binary verification of correctness).
I say this because the real reason "they are always confident" is in some sense even more contrived. Training text where the speaker sounded more confident is more likely to contain a correct answer.
There is an entire political party representing something like half the population of the US dedicated to shrinking the regulatory apparatus, including the FDA. That doesn't sound like career security to me.
School shooters and "crazy people acting violently" are probably incredibly rare by comparison. Sub/urban environments probably reduce the impact of the mentally ill on their surrounding neighbours in a whole host of ways, in fact -- in no small part because treatment is much less available in rural areas. (There's more of them in a smaller area and they're more visible and increasingly less criminalized, so people get the idea that poverty and health issues of that variety are modern, urban problems.)
Keywords in the literature around this include "community impact" or "community health", but it's not my area.
On a different note:
I think the most useful coding copilot tools for me reduce "manual overhead" without attempting to do any hard thinking/problem solving for me (such as generating arguments and types from docstrings or vice-versa, etc.). For more complicated tasks you really have to give copilot a pretty good "starting point".
I often talk to myself while coding. It would be extremely, extremely futuristic (and potentially useful) if a tool like this embedded my speech into a context vector and used it to as an additional copilot input so the model has a better "starting point".
I'm a late adopter of copilot and don't use it all the time but if anyone is aware of anything like this I'd be curious to hear about it.
They studied a hunter-gatherer tribe in Tanzania and found that despite having similar patterns of sedentary behaviour, their blood biomarkers of cardiometabolic dysfunction were much lower than those in industrialized nations. The authors posit that one reason for this is that sedentary behaviour in this group of individuals does not involve furniture -- rather, it usually involves a "deep squat", and the authors show that in this position the muscles are much more engaged than when someone sits in a chair.
This is consistent with evidence that breaking up periods of sitting with movement is good for you.
Their open-access paper talks about some evolutionary context for this hypothesis [1].
[1] 10.1073/pnas.1911868117
> Alexander later filed a lawsuit against Louisiana-Pacific claiming that the band saw had been weakened from previous strikes with nails, but that he was forced to work with the saw or face dismissal.
[1] https://www.uclalawreview.org/wp-content/uploads/securepdfs/...
Consider ShotSpotter, which uses an array of microphones in an urban environment to detect gunshots (and often then deploy officers to the location) [1]:
> A ShotSpotter expert admitted in a 2016 trial, for example, that the company reclassified sounds from a helicopter to a bullet at the request of a police department customer, saying such changes occur “all the time” because “we trust our law enforcement customers to be really upfront and honest with us.”
In this case, it seems like it's more like "evidence laundering" - a cop found a bullet (presumably through legitimate means) and would like to use the ShotSpotter results as additional evidence that the shooting took place, and so requests a re-classification of the audio recording. Even in this case, where the parallel evidentiary construction is presumably legitimate, one can imagine the problem - a jury may put more stock in a ShotSpotter result than the cop's testimony about a bullet. But in this case, the ShotSpotter "result" is due precisely to that testimony.
Never mind the fact that ShotSpotter microphones are powerful enough to pick up loud conversations [2]:
> The apparent ability of ShotSpotter to record voices on the street raises questions about privacy rights and highlights another example of how emerging technologies can pose challenges to enforcing the law while also protecting civil liberties.
Predictive policing will require large-scale data collection, and policing institutions don't seem to always use it the way we'd want them to.
[1] https://www.aclu.org/news/privacy-technology/four-problems-w...
[2] https://www.southcoasttoday.com/story/news/crime/2012/01/11/...
The parent comment's point is that although the reported effect is significant at $\alpha = 0.05$ (the usual "95% CI" you mentioned), there are other problems that render their test of this hypothesis less than valid.
There is some evidence (in the sense of evidence-based medicine) that a low-FODMAP (fermentable *saccharides) diet reduces symptoms in patients with irritable bowel syndrome. [1]
As far as vegetables go, according to one site high-FODMAP vegetables include alliums and artichokes. [2]
It is worth noting that the authors of the linked review paper caution that it is unknown whether a low-FODMAP diet may have long-term adverse effects.
[1] doi:10.2147/CEG.S86798
[2] https://www.monashfodmap.com/about-fodmap-and-ibs/high-and-l...
I find the following workflow works well, for example:
1. Define steps depending on a `config.yml`.
2. Run an initial experiment (with an initial config) and commit the results.
3. Update config (preserving the alternate config and using symlinks from `config.yml` to various new configs if necessary), re-run, and commit.
4. Results are then all preserved in your git history.
It's important to note that risk factors for mortality overlap significantly with the risk factors for other serious adverse events (see e.g. disability status [1], or HELPP syndrome [2]). In a sense, mortality (although fortunately rare as noted in other comments) is therefore a useful proxy for other severe, potentially life-changing adverse pregnancy-related medical events which are much more common than just plain death. (As an analogy, something like 10-15% of heart attacks are fatal, but many more are debilitating.)
[1] doi:10.1001/jamanetworkopen.2021.38414
[2] doi:10.1111/1471-0528.16225
I modeled mine on the Django ASGI reference library's server implementation, which uses the data structure for maintaining references to stateful event consumers. Exception handling is done with a pre-scheduled long-running coroutine that looks at the map.
I'm curious about your second point -- why exactly do things get bad with high tail latency? Is it only a weakness of the data structure when used for caching? I'm having trouble picturing that.
I'm familiar with the idea of re-purposing drugs for rare disease treatments (most of my adjacent work has been in very early-stage academic research), but I'm curious about the financials here. Could some of the financial risk here be minimized by aggregating multiple groups of patients, all suffering from different rare diseases? From what I know about the process, the answer is yes, but I'd be curious to hear from somebody closer to the process.
I don't think this article is credible, but my question above is asked in good faith. There do seem to be informed people who doubt the veracity of American claims of increasing Russian hostility.
I had a very similar idea recently (representing "big" preprocessing pipelines in Python as DAGs), but for research-scale projects. Funnily enough, I was also motivated by a time series project that has resulted in several thousand lines of gnarly preprocessing code, and growing - mostly in Pandas.
I'm not sure Hamilton would totally work for our use case, but I'll be following closely either way. Thank you for open sourcing this!
Sorry, if I'd had the foresight to use a throwaway I'd drop a link to our group, but I prefer not to publicly associate my HN account with work. We're a small team of software and data engineers, machine learning scientists, and health policy folks at a large research institution in Canada that take on clients to work on stuff like this (from early stage research to approvals to deployment). I'd be happy to reach out with my contact info privately if you're interested, just let me know.
That makes sense! I wondered if it had something to do with the group component, and I agree that the customer pool is growing.
Congratulations on the launch!
You mentioned the medication is as effective in women as it is in men. While I understand men are underserved in this space and so I respect the decision to focus on that population on that basis, I'm curious if there are business elements to that decision as well? Would you ever expand to serving women, given that they seem to be a larger potential customer pool?
Pre-made Anki decks vary in quality. Some are great. But none that I've found have the "themes" that personally keep me motivated to language learn long term.
There's one link in the chain here missing that some people here seem to be ignoring. The authors of this post (while entirely correct) draw no link between "bad data" (which is doubtlessly responsible for a large number of "bad papers"/"bad trials") and "bad clinical practice."
I don't know a single clinician who would base their care on the findings of a single-center RCT of the kind described in this article. Or the findings of a meta-analysis of single-center RCTs, for that matter.
Bad data happens in multi-center RCTs too, and in fact that's what I'm focused on, but a lot of work already (and therefore $, for the cynical) goes into the validation of data (see [1] for a brief description). Phase III clinical trials in the west practically require a robust multi-center RCT, where systemic fraud is very difficult to perform (but not impossible [2]). By the time a Phase III trial is conducted, the efficacy of the drug can already be estimated, and the focus of the drug company (who yes, often fund these trials) is to conduct a trial which is unimpeachable in the face of a regulatory board (who are generally good at their jobs, although the revolving-door tends to reduce public trust and should be legislated away).
In short, I support most of the proposed changes to incentives around publish-or-perish. I reject the notion that these incentives are (currently) significant drivers of decreased quality of standard of care in the West. I think global governance structures, as suggested in this article, could improve understanding among both clinicians who are not necessarily scientists and the general public about just how validated a given standard of care is.
tl;dr Most good evidence-based practitioners already think this way -- not because they inherently believe fraud is rampant, necessarily, but because evidence says the kinds of studies where fraud is most prevalent are untrustworthy for other reasons.
[1] doi:10.1177/1740774512447898
The relative success of attacks on nets to extract their training data support that this happens in practice too.
Generalization performance as it stands now always has to be evaluated empirically.