371 karma · joined October 6, 2014
While I haven't published my results yet (dang OpenAI beat me to it!), they are very similar and I have no doubt that this technology will significantly improve the detection and diagnosis of rare disease, which is a nice win for humanity.
I'm in the process of getting my first batch of long reads, but I am skeptical that this is the "just" what's needed. There is little doubt that long read > short read, but I think that computational techniques for both need to be improved significantly.
There is already some clinical evidence to support my hypothesis. The first clinical long read trial at Kansas City Mercy showed a 10% bump in diagnostic rate, which is great but not fully solving the problem: https://news.childrensmercy.org/unlocking-answers-faster-chi...
Your assumption is correct about my technique. I cloned (and expanded) this workflow into an LLM harness, so the LLM is basically orchestrating a bunch of tools that normally humans would use (and writing the conclusions and doing all the standard LLM stuff).
I wouldn't say anything is automatic or taken for granted, but it is actually relatively common for more thorough reanalysis to uncover something that the first pass missed. I hinted at this in the post, but the reason that this doesn't happen today is human bandwidth.
A core part of my thesis is that that this highly specialized human bandwidth can be scaled with AI.
It may work. It may not work. But I would feel bad if I didn't give it a try.
> Happy to see it. I wish you all the luck and will be the first one praising your solution if I see convincing results.
Appreciate that! Hopefully, they will come.
If you are open to chatting about your experience, I'd love to hear from you. I spend a lot of time learning from and supporting other rare disease families these days.
And the anti-natal thing was kind-of joking not joking. I do know lots of people with kids there now, but when my wife first got pregnant, we were alone.
I'm planning on getting one out in the next few weeks characterizing the system and how it performed on real clinical use-cases vs. alternatives and existing tools.
The TL;DR is that Gamow Labs is a harness and interface company on top of SOTA LLMs as you suggested, but my harness and interface outperforms the existing thing. While this approach would have earned me the "wrapper company" label last year, I hope the success of OpenEvidence, Harvey, Perplexity, and so on has opened minds with respect to the value here.
It was only working through clinical cases that I realized how much more I needed beyond dropping raw reads into Codex.
While I am truly grateful for him and the team for their contributions to neonatal genetics (and hosting me in San Diego for a few days to show me how I could help), Rady was actually the unnamed lab that failed to diagnosis my son.
And this happens all the time. The WGS NICU diagnostic rate is only ~30%, depending on who you ask. Just because people have been working at this for a decade and products exists, doesn't mean it's a solved problem.
I don't know if you read until the end of my post, but I did run a small experiment in collaboration with an academic geneticist and outperformed the first-line clinical labs across the board. My approach, which is essentially Claude Code for genetics, is fundamentally different and novel than how this work is done today and seems to perform much better in early experiments. Time will tell is this generalizes to all clinical work.
I'm planning on publishing evals and benchmarks in the next few weeks, but out-of-the-box systems actually don't do very well for a variety of reasons.
The general point is that separating PM and eng doesn't make sense any longer. Which subsumes which is an interesting debate.
Your argument that 4.6 Opus makes the engineering skill set useless is totally false and maybe shows you haven't built anything complicated, but it is possible that Opus 5.2 will get there.
But to this sister comment's point, I do think that the dedicated PM role will vanish and the classic BigCo PM will need to look a lot more like the startup one.
I think that all PMs will need to get onto the engineering, design, or research ladder. We are already seeing companies eliminate the function here and there and I expect the trend to continue.
There is still a big gap between 11Labs and Character.ai and the VoiceCraft voices would not be confused for the real speaker, but this is much closer.
https://www.ddmckinnon.com/2024/10/03/dans-weekly-ai-speech-...
I tried zero-shot voice cloning in all of the top OSS models in the Arena and performance was bad.
I think the point of the article is that it’s really hard to make these judgments and doing way with the whole system entirely makes the most sense, which seems reasonable to me.
This is impressive both because it's hard to keep such a big business growing at that rate and because essentially everyone in my social circle has moved on from going to Google first for information. I guess our demographic is not predictive of the larger market.
Keep in mind that most machine learning fundamentals haven't changed for decades. While new architectures/trends are always emerging, the lessons that you will take away from an academic program like OMSCS will be relevant for the rest of your career.
That said, call it whatever you want it. Most people in crypto call CEX/DEX stat arb to differentiate from true risk-free arbitrage, but I agree that people coming from a traditional trading perspective would call this pure arb.
TL;DR: this is a pairs trading strategy that relies on the (very strong) statistical assumption that the price of a token on a centralized and decentralized exchange will converge.
My longer definition below:
Stat arbs: in finance, statistical arbitrage generally refers to any trade where a pair of assets should statistically move in a certain way. However, there are degrees of should. TradFi traders might reason that Meta and Google are both in the ads business, so if Meta is relatively expensive and Google is relatively cheap, they should short the former and long the latter. However, this is a weak argument. Perhaps Meta is just a better business or Google has structural problems. A stronger stat arb thesis is that Royal Dutch Shell used to be traded on both American and European exchanges. If the shares were trading at different prices on each, nearly risk-free profits are available to those who close the spread. This is what stat arb means in a crypto setting. AVAX may be trading at slightly different prices on Binance and on various blockchains.