AlphaFold 3 can rapidly reduce a vast search space in a way physically-based methods alone cannot. This narrowly focused search space allows scientists to apply their rigorous, explainable, physical methods, which are slow and expensive, to a small set of promising alternatives. This accelerates drug discovery and uncovers insights that would otherwise be too costly or time-consuming.
The future of science isn't about AI versus traditional methods, but about their intelligent integration.
My only worry is that AlphaFold and others, e.g. ESM, seem to be bit fragile for out-of-distribution sequences. They are not doing a great job with unusual sequences, at least in my experience. But hopefully they will improve and provide better uncertainty measures.
It’s actually required as part of the submission for FDA approval that you posit a specific Mechanism of Action for why your drug works the way it does. You can’t get approval without it
Vioxx is a nice example of a molecule that got all the way to large-scale deployment before being taken off the market for side effects that were known. Only a decade before that, I saw a very proud pharma scientist explaining their "mechanism of action" for vioxx, which was completely wrong.
Some had tried to come up with other criteria to confirm you have discovered an underlying principle without predictive power, such as on aesthetics - but this is seen by the majority of scientists as basically a cop out. See debate around string theory.
Note that this comment is summarizing a massive debate in the philosophy of science.
There is a car. We think it drives by burning petrol somehow.
How do we test this? We take petrol away and it stops driving.
Ok, so we know it has something to do with petrol. How does it burning the petrol make it drive?
We think it is caused by the burned petrol pushing the cylinders, which are attached to the wheels through some gearing. How do we test it? Take away the gearing and see if it drives.
Anyway, this never ends. You can keep asking questions, and as long as the hypothesis is something you can test, you are doing science.
You discovered a principle.
Better example:
There is a car. We don’t know how it drives. We turn the blinkers on and off. It still drives. Driving is useful. I drive it to the store
The best part is where the geneticist ties the arms of all the suit-wearing employees and it has no functional effect on the car.
You’ve discovered magic.
When you read about a wizard using magic to lay waste to invading armies, how much value would you guess the armies place in whether or not the wizard truly understands the magic being used against them?
Probably none. Because the fact that the wizard doesn’t fully understand why magic works does not prevent the wizard from using it to hand invaders their asses. Science is very much the same - our own wizards used medicine that they did not understand to destroy invading hordes of bacteria.
for many basic/fundamental mathematical objects we don't (yet) have simple mechanistic ways to compute them.
so if a probabilistic model spits out something very useful, we can slap a nice label on it and call it a day. that's how engineering works anyway. and then hopefully someday someone will be able to derive that result from "first principles" .. maybe it'll be even more funky/crazy/interesting ... just like mathematics arguably became more exciting by the fact that someone noticed that many things are not provable/constructable without an explicit Axiom of Choice.
https://en.wikipedia.org/wiki/Nonelementary_integral#Example...
Yes, but we're taking about roughly the opposite of a proof
and it seems with these molecular biology problems we constantly have the problem of specificity (model prediction quality) vs sensitivity (model applicability), right? but due to information theory constraints there's also a dimension along model size/complexity.
so if a ML model can push the ROC curve toward the magic left-up corner then likely it's getting more and more complex.
and at one point we simply are left with models that are completely parametrized by data and there's virtually zero (direct) influence of the first principles. (I mean that at one point as we get more data even to do model selection we can't use "first principles" because what we know through that is already incorporated into previous versions of the models. Ie. the information we gained from those principles we already used to make decisions in earlier iterations.)
Of course then in theory we can do model distillation, and if there's some hidden small/elegant theory we can probably find it. (Which would be like a proof through contradiction, because it would mean that we found model with the same predictive power but with smaller complexity than expected.)
// NB: it's 01:30 here, but independent of ignorance-o-clock ... it's quite possible I'm totally wrong about this, happy to read any criticism/replies
Yes, but a perfect oracle has no explanatory power, only predictive.
That doesn’t diminish the value that patients received in any way even though it would be more satisfying to make predictions and design something to interact in a way that exactly matches your theory.
This has happened before. Newtonian mechanics was incomprehensible spooky action at a distance, but Einstein clarified gravity as the bending of spacetime.
Like, quantum mechanics doesn’t seem, to me, to just be a way of describing how to predict things. I view it as saying substantial things about how things are.
Sure, there are different interpretations of it, which make the same predictions, but, these different interpretations have a lot in common in terms of what they say about “how the world really is” - specifically, they have in common the parts that are just part of quantum mechanics.
The qau that can be spoken in plain language without getting into the mathematics, is not the eternal qau, or whatever.
Prediction is understanding. What we call "understanding" is a cognitive illusion, generated by plausible but brittle abstractions. A statistically robust prediction is an explanation in itself; an explanation without predictive power explains nothing at all. Feeling like something makes sense is immeasurably inferior to being able to make accurate predictions.
Scientists are at the dawn of what chess players experienced in the 90s. Humans are just too stupid to say anything meaningful about chess. All of the grand theories we developed over centuries are just dumb heuristics that are grossly outmatched by an old smartphone running Stockfish. Maybe the computer understands chess, maybe it doesn't, but we humans certainly don't and we've made our peace with the fact that we never will. Moore's law does not apply to thinking meat.
Newton was a bit of a brat but everybody accepted his explanation. Then the problem turned to trying to explain gravity.
Thus science advances, one explanation at a time.
Scientific method can help us rule out what underlying principles are definitely not. Any such principles are not actually up to be “discovered”.
If probabilistic ML comes along and does a decent job at predicting things, we should keep in mind that those predictions are made not in context of absolute truth, but in context of theories and models we have previously developed. I.e., it’s not just that it can predict how molecules interact, but that the entire concept of molecules is an artifact of just some model we (humans) came up with previously—a model which, per above, is probably incomplete/incorrect. (We could or should use this prediction to improve our model or come up with a better one, though.)
Even if a future ML product could be creative enough to actually come up with and iterate on models all on its own from first principles, it would not be able to give us the answer to the question of underlying principles for the above-mentioned reasons. It could merely suggest us another incomplete/incorrect model; to believe otherwise would be to ascribe it qualities more fit for religion than science.
People clearly have been able to discover many underlying principles using the scientific method. Then they have been able to explain and predict many complex phenomena using the discovered principles, and create even more complex phenomena based on that. Complex phenomena such as the technology we are using for this discussion.
Words dont have any inherent meaning, just the meaning they gain from usage. The entire concept of truth is an artifact of just some model (language) we came up with previously—a model which, per above, is probably incomplete/incorrect. The kind of absolute truth you are talking about may make sense when discussing philosophy or religion. Then there is another idea of truth more appropriate for talking about the empirical world. Less absolute, less immutable, less certain, but more practical.
Exactly—except you are talking about it, too. When you say “discovering underlying principles”, you are implying the idea of absolute truth where there is none—the principles are not discovered, they are modeled, and that model is our fallible human construct. It’s a similar mistake as where you wrote “explain”: every model (there should always be more than one) provides a metaphor that 1) first and foremost, jives with our preexisting understanding of the world, and 2) offers a lossy map of some part of [directly inaccessible] reality from a particular angle—but not any sort of explanation with absolute truth in mind. Unless you treat scientific method as something akin to religion, which is a common fallacy and philosophical laziness, it does not possess any explanatory powers—and that is very much by design.
You are assigning meanings to words like "discovering", "principles", and "explain" that other people don't share. Particularly people doing science. Because these absolute philosophical meanings are impossible in the real world, they are also useless when discussing the reality. Reserving common words for impossible concepts would not make sense. It would only hinder communication.
High level, I see a distinction between theory and practice, between an oracle predicting without explanation, and a well-thought out theory built on a partnership between theory and experiment over centuries, ex. gravity.
I have this feeling I can't shake that the knife you're using is too sharp, both in the specific example we're discussing, and in general.
In the specific example, folding, my understanding is we know how proteins fold & the mechanisms at work. It just takes an ungodly amount of time to compute and you'd still confirm with reality anyway. I might be completely wrong on that.
Given that, the proposal to "dedicate...engineer[s] towards finding ethical ways to improve...intelligence so that we can appreciate the underlying principles better" begs the question of if we're not appreciating the underlying principles.
It feels like a close cousin of physics theory/experimentalist debate pre-LHC, circa 2006: the experimentalists wanted more focus on building colliders or new experimental methods, and at the extremes, thought string theory was a complete was of time.
Which was working towards appreciating the underlying principles?
I don't really know. I'm not sure there's a strong divide between the work of recording reality and explaining it. I'll peer into a microscope in the afternoon, and take a shower in the evening, and all of a sudden, free associating gives me a more high-minded explanation for what I saw.
I'm not sure a distinction exists for protein folding, yes, I'm virtually certain this distinction does not exist in reality, only in extremely stilted examples (i.e. a very successful oracle at Delphi)
Not saying ML methods haven't shown important reproducibility challenges, but to just shut them down due to not being "useful science" is inflexible.
Can we differentiate?
Now, if we look at history of science and technology, there is a shit ton of practical stuff that was found only by pure accident - discoveries of which could not be predicted from any previous theory.
I would view both A) and B) as net positives. But our teaching of the next generation of scientists needs to adapt.
The worst case scenario is of course that the middle management driven enshittification of science will proceed to a point where there are only few people who actually are scientists and not glorified accountants. But I’m optimistic this will actually super charge science.
With good luck we will get rid of the both of the biggest pathologies in modern science - 1. number of papers published and referred as a KPI 2. Hype driven super politicized funding where you can focus only one topic “because that’s what’s hot” (i.e. string theory).
The best possible outcome is we get excitement and creativity back into science. Plus level up our tech level in this century to something totally unforeseen (singularity? That’s just a word for “we don’t know what’s gonna happen” - not a specific concrete forecasted scenario).
It's more specific than you make it out. The singularity idea is that smart AIs working on improving AI will produce smarter AIs, leading to an ever increasing curve that at some point hits a mathematical singularity.
Nobody knows what singularity would actually mean from the point of view of specific technological development.