Malaria’s complex lifecycle [1] seems like it would be easy to “break” with different interventions, but we’ve seen historically malaria has been difficult to eradicate. Why is this?
1. https://en.m.wikipedia.org/wiki/Plasmodium#/media/File%3ALif...
649 karma · joined May 7, 2021
Malaria’s complex lifecycle [1] seems like it would be easy to “break” with different interventions, but we’ve seen historically malaria has been difficult to eradicate. Why is this?
1. https://en.m.wikipedia.org/wiki/Plasmodium#/media/File%3ALif...
I constantly use AlphaFold structures today [1]. And AlphaFold is fantastic. But it only replaces one small step in solving any real-world problem involving proteins such as designing a safe, therapeutic protein binder to interrupt cancer-associated protein-protein interactions or designing an enzyme to degrade PFAS.
I think the primary achievement is that it gets protein structures in front of a lot more smart eyes, and for a lot more proteins. For “everyone else” who never needed to master computational protein structure prediction workflows before, they now have easy access to the rich, function-determinative structural information they need to understand and solve their problem.
The real tough problem in protein design is how to use these structure predictions to understand and ultimately create proteins we care about.
1. https://alexcarlin.bearblog.dev/multistate-protein-design-wi...
Most magnet timers also remember the last duration, so if you use the same timer a lot for the same task (tea, for example), it’s literally a single button press. The same operation on any kind of smart device contains a staggering number of steps, each of which requires cognition and attention.
Magnet timers are also super cheap, so you can get another one if you have two favorite durations. A simple solution meets a simple problem
David Baker’s lab has recently published on using their own diffusion model (RFdiffusion) to design novel biocatalysts that perform hydrolysis using a catalytic triad of serine, aspartic acid, and histidine, as well as an oxyanion hole, which is much more complex than the binders designed by AlphaProteo [1].
It gives me hope that we’ll soon be able to design biocatalysts as good as natural ones, but for any problem we care about.
1. https://alexcarlin.bearblog.dev/novel-enzymes-from-a-diffusi...
It turns out that the Escherichia coli (to spell out its Latin binomial) that cause disease are in some sense “diseased” themselves: the genes that enable them to be pathogenic, or make them pathogenic, I should say, are originally from a phage, a type of virus that infects bacteria [1]. In a manner that is not the same as, but conceptually similar to how HIV inserts its genes into the human’s genome, phages insert their genes (termed the “prophage”) into the bacterial genome.
In addition, most strains of pathogenic Escherichia are also holding on to an entirely separate, small, circular “genome” called a plasmid, also of exogenous origin, that contains additional genes that make them pathogenic.
So in addition to wide genome variation within the “species” (which is not really the same thing for bacteria as for mammals, mind you) of Escherichia coli, many subtypes have additional genetic material from endogenous sources that substantially changes their observed characteristics (phenotype).
Humans are natural machines capable of sensing and verifying the correctness of a piece of text or an image in milliseconds. So if you have a model that generates text or images, it’s trivial to see if they’re any good. Whereas for biology, the time to validate a model’s output is measured more in weeks. If you generate a new backbone with RFDiffusion, and then generate some protein sequences with LigandMPNN, and then want to see if they fold correctly … that takes a week. Every time. Use ML to solve _that_ problem and you’ll be rich.
TFA mentions the difficulty of performing biological assays at scale, and there are numerous other challenges. Such as the number of different kinds of assays required to get the multimodal data needed to train the latest models like ESM-3 (which is multimodal, in this context meaning primary sequence, secondary structure, tertiary structure, as well as several other tracks). You can’t just scale a fluorescent product plate reader assay to get the data you need. We need sequencing tech, functional assays, protein-protein interaction assays, X-ray crystallography, and a dozen others, all at scale.
What I’d love to see companies like A-Alpha and Gordian and others do is see if they can use the ML to improve the wet lab tech. Make the assays better, faster, cheaper with ML. Like how they use ML to translate the electrical signals of DNA passing through the pore into a sequence in the Nanopore sequencers. So many companies have these sweet assays that are very good. In my opinion, if we want transformative progress in biology, we should spend less time fitting the same data with different models, and spend more time improving and scaling wet lab assays using ML. Can we use ML to make the assay better, make our processes better, to improve the amount and quality of data we generate? The thesis of TFA (and experience) suggests that using the data will be the easy part
1. https://alexcarlin.bearblog.dev/why-is-progress-slow-in-gene...
It’s also worth remembering that it was David Baker who originally came up with the idea of extending AlphaFold from predicting just proteins to predicting ligands as well [2].
1. https://github.com/baker-laboratory/RoseTTAFold-All-Atom
2. https://alexcarlin.bearblog.dev/generalized/
Unlike AlphaFold 3, which predicts only a small, preselected subset of ligands, RosettaFold All Atom predicts a much wider range of small molecules. While I am certain that neither network is up to the task of designing an enzyme, these are exciting steps.
One of the more exciting aspects of the RosettaFold paper is that they train the model for predicting structures, but then also use the structure predicting model as the denoising model in a diffusion process, enabling them to actually design new functional proteins. Presumably, DeepMind is working on this problem as well.
The problem with enzymes eating plastic is that enzymes are small Pacman-shaped protein blobs that are maybe 10 nanometers in diameter, whereas things made of plastic like bottles or even microplastics are huge in comparison. How do you get the little Pacman jaws around the bottle to start breaking it down?
The research paper [1] describes the authors’ effective innovation. They make a protein where one end is a pore-forming shape, and the other end is a PET cutting (called a PETase in the jargon of the field). This way, their protein can access nooks and crannies in the macroplastic shapes, allowing tons of copies of this small enzyme to fully degrade a bottle.
Without this, a great deal of physical agitation is required to break down the plastics into small enough chunks that earlier Pacman enzymes could work on, increasing the time and the cost.
I hope we’ll see the idea of linking the enzymatic “scissors” to a protein pore be used to engineer enzymes to degrade other types of plastics in the future, as the general idea of getting the catalytic machinery into physical contact with every bit of the bottle is broadly applicable to all plastics, not just PET (which is great news)
1. https://phys.org/news/2023-10-scientists-artificial-protein-...
In my experience, people excel in environments where they are given a high level of trust, autonomy, and clear goals on the timescale of months and years.
Teams on the rack of egotistical middle managers get stretched thin until they break. There aren’t a lot of good examples of “micromanagitis” being cured. So-called leaders taking credit for their team’s work leads to their team not trusting that leadership has their growth in mind. Art Blakey’s real legacy is both the amazing players he mentored and tutored, but also the attitude that if you give people the tools they need, they’ll excel
How LLMs are able to give convincing wrong answers: they “can predict the correct ‘shape’ of an answer” (parent).
Why LLMs are able to give convincing wrong answers is a little more complicated, but basically it’s because the model is tuned by human feedback. The reinforcement learning from human feedback (RLHF) that is used to tune LLM products like ChatGPT is a system based on human ranking. It’s a matter of getting exactly what you ask for.
If you tune a model by having humans rank the outputs, despite your best efforts to instruct the humans to be dispassionate and select which outputs are most convincing/best/most informative, I think what you’ll get is a bias towards answers humans like. Not every human will know every answer, so sometimes they’ll select one that’s wrong but likable. And that’s what’s used to tune the model.
You might be able to improve this with curated training data (maybe something a little more robust than having graders grade each other). I don’t know if it’s entirely fixable though.
The brilliant thing about the parent’s comment about the “shape” of the answer is that it reveals how much humans have (uh, historically, now, I guess) relied on the shape of information to convey its trustworthiness. Expand the notion of “shape” a bit to include the medium. If somebody bothered to take the time to correctly shape an answer, we take that as a sign of trustworthiness, like how you might trust something written in a carefully-typeset book more than this comment.
Surely no one would take the time to write a whole book on a topic they know nothing about. Implies books are trustworthy. Look at all the effort that went in. Proof of effort. When perfectly-shaped answers in exactly the form you expected are presented in a friendly way and commercial context, they certainly read as trustworthy as Campbell’s soup cans. But LLMs can generate books worth of nonsense in exactly the right shapes without effort, so we as readers can no longer use the shape of an answer to hint at its trustworthiness.
So maybe the answer is just to train on books only, because they are the highest quality source of training data. And carefully select and accredit the tuning data, so the model only knows the truth. It’s a data problem, not a model problem
Furthermore, the authors must engage a lot of human curation to ensure the sequences they generate are active. First, they pick an easy target. Second, they employ by-hand classical bioinformatics techniques on their predicted sequences after they are generated. For example, they manually align them and select those which contain specific important amino acids at specific positions which are present in 100% of functional proteins of that class, and are required for function. This is all done by a human bioinformatics expert (or automated) before they test the generated sequences. This is the protein equivalent of cherry-picking great examples of, for example, ChatGPT responses and presenting them as if the model only made predictions like that.
One other comment, in protein science, a sequence with 40% identity to another sequence is not “very different” if it is homologous. Since this model is essentially generating homologs from a particular class, it’s no surprise at a pairwise amino acid level, the generated sequences have this degree of similarity. Take proteins in any functional family and compare them. They will have the same overall 3-D structure—called their “fold”—yet have pairwise sequence identities much lower than 30–40%. This “degeneracy”, the notion that there are many diverse sequences that all fold into the same shape, is both a fundamental empirical observation in protein science as well as a grounded physical theory.
Not to be negative. I really enjoyed reading this paper and I think the work is important. Some related work by Meta AI is the ESM series of models [1] trained on the same data (the UniProt dataset [2]).
One thing I wonder is about the vocabulary size of this model. The number of tokens is 26 for the 20 amino acids and some extras, whereas for a LLM like Meta’s LLaMa the vocab size is 32,000. I wonder how that changes training and inference, and how we can adopt the transformer architecture for this scenario.
The Nest smoke alarms at a friend’s house say (as in play over the speakers a voice saying) “Remote access attempt detected” every so often while we’re sitting in the living room. At least twice a year they go off in the middle of the night (without smoke).
So I guess when you’re contacting customers who don’t want it, and your alarms wrongly sound, I would say you have a problem with your false positive rate
One yet further objection to the many excellent already-made points: the deployment of LLMs as clean-slate isolated instances is another qualitative difference. The human brain and its sensory and control systems, and the mind, all coevolved with many other working instances, grounded in physical reality. Among other humans. What we might call “society”. Learning to function in society has got to be the most rigorous training for prompt injection I can think of. I wonder how a LLM’s know-it-all behavior works in a societal context? Are LLMs fun at parties?
>For this type of interaction, we must use RL training, as supervised training teaches the model to lie. The core issue is that we want to encourage the model to answer based on its internal knowledge, but we don't know what this internal knowledge contains. In supervised training, we present the model with a question and its correct answer, and train the model to replicate the provided answer.
The author says he’s summarizing a talk by John Schulman of OpenAI [2] but I haven’t personally watched the video. In any case, this is an interesting insight.
Say we set up a supervised learning scenario where we ask the model to use its internal knowledge to answer a question and compare its answer to one written by a human. If the two answers essentially say the same thing, but in different words, in the supervised learning case the model is penalized. In the RL case, it’s rewarded. That’s the difference.
1. https://gist.github.com/yoavg/6bff0fecd65950898eba1bb321cfbd...
The method is simply this: for half a year, every week, go and find the same news story from three different publishers and read all three versions.
Do this solidly 25 times and you will begin to see the outlines of exactly the set of skills that you need to objectively classify information that’s found on the Internet. This is a simple and highly effective training technique that uses the mass variety of news sites on the Internet for an educational purpose. Once you begin to see and compare the narratives underlying most “news”, they will be impossible to un-see
>This struggle may be a moral one, or it may be a physical one, and it may be both moral and physical, but it must be a struggle. Power concedes nothing without a demand. It never did and it never will. Find out just what any people will quietly submit to and you have found out the exact measure of injustice and wrong which will be imposed upon them, and these will continue till they are resisted with either words or blows, or with both. The limits of tyrants are prescribed by the endurance of those whom they oppress. (Emphasis mine)
1. https://www.blackpast.org/african-american-history/1857-fred...
I did some math in college and when I started knowing how to analyze the behavior of functions (and developing the mental math tools to imagine what they look like without having to actually draw them) that’s when I felt like I was kinda getting it
I guess I’d ask why the author thinks that training LLMs on their own output will make them worse. Like, if the problem is that LLM-generated content is less useful than human-generated content because it’s “just averaging out inputs” (paraphrase of common argument, not quote from TFA), how does adding more data at the average change the distribution?
>As is now, LLMs regularly hallucinate, generate biased content or fundamentally misinterpret the task even though nothing in the wider world has been adversarial to them.
This really got me thinking about what is meant by “adversarial”. As in, adversarial with whom? The model itself? Its deployers?
If I successfully trick ChatGPT, the system, into telling me some secrets about its inner workings, we can call that an attack on the commercial project as released by OpenAI, but can we call it an attack on the model itself?
All the text used to train LLMs is heavily processed and filtered already. I think it’s more likely that, rather than LLM-made text diluting out the good training data, it will simply add to the corpus. Might add a few cycles to the line-level duplication step
The very largest plain transformer models trained on protein sequences (analogous to plain text) are about 15B parameters (I am thinking of Meta AI’s ESM-2 [1]). These can do for protein sequences what LLMs do for text (that is, they can “fill in the blank” to design variations, generate new proteins that look like their training data), and tell you how likely it is that a given sequence exists.
Some cool variations of transformers have applications for protein design, like the now-famous SE(3) equivariant transformer used in the structure prediction module of AlphaFold [2], now appearing in the research paper [3] accompanying TFA, as well as variations on the transformer such as the message passing model ProteinMPNN [4], which builds on a neighbor graph-structured transformer [5]
1. https://github.com/facebookresearch/esm
2. https://github.com/deepmind/alphafold
3. https://www.biorxiv.org/content/10.1101/2022.12.09.519842v2
4. https://github.com/dauparas/ProteinMPNN
5. https://github.com/jingraham/neurips19-graph-protein-design
> YT-513-U: We create an additional dataset called YT-513-U to ensure coverage of lower resource languages in our pre-training dataset. We reached out to vendors and native speakers to identify YT videos containing speech in specific long tail languages, collecting a dataset of unlabeled speech in 513 languages. [1]
Inside platelets, chemical reactions are sped along by the enzymes. One of the important chemical reactions for clotting is the conversion of arachidonic acid (through a couple of intermediates) to thromboxane. Now, arachidonic acid is so simple: a linear, 20-carbon chain with a carboxylic acid group at one end, and four double bonds in the chain. Converting it to thromboxane takes a couple of steps: the enzymes need to kink the chain and oxidize specific carbons to make them polar.
The trouble is, aspirin inhibits the enzymes! So the enzymes inside the platelets can’t make their thromboxane. And, as such, clotting is inhibited, which results in bleeding. So this is why aspirin can cause bleeding
The idea is that exposure to lead causes people to have problems that lead to violent behavior later in life. The evidence presented includes the correlation observed between the removal of lead from gasoline and a sharp fall in crime rates 20 years later.
1. https://www.motherjones.com/environment/2016/02/lead-exposur...
>we are very, very far and this depresses me. What is the way forward? :( Maybe I should just do a startup
and was a founding member of OpenAI just a few years later in 2015
Better for whom?
It seems to me that institutions that serve the public should do so with transparency. The idea that “the experts” know better than “the public” is a dangerous and elitist one. The suggestion that “the experts” be allowed to hide their math while they fumble at the economy trying to make a buck is absurd
I’m one of these people who likes to listen to full albums. At one point I had a few thousand ripped as quality MP3. In college there were campus peer-to-peer ways of sharing music and my collection grew a lot.
In grad school, the streaming services came out. The streaming services like Spotify have some major downsides. Sometimes things go unplayable. Stuff is missing. Stuff is changed/censored without warning. The UI is optimized for whatever Spotify wants to show you, not what you want to listen to. I assume it makes sense for Spotify to optimize their costs by suggesting stuff they pay less for. For someone who likes albums, the little annoyances add up.
I recently pulled my dad’s record collection out of his attic. It’s about a thousand LPs from the sixties and seventies. Stuff I like since I grew up listening to it. I’ve spent a few months browsing Discogs and buying bargin-bin vinyl and CDs (lots of scrolling, filters are useless). I’ve bought about five hundred CDs and maybe a hundred vinyl records over several months. The majority of the CDs are $1–2, and the records are mostly under $5. It all comes via Media Mail, shipped for a few bucks. One of the packages came from Brooklyn and had been inspected by USPS to make sure it contained only CDs. One package came from Jamaica and the vinyl sleeves had a bit of fine sand. It’s a cheap, fun hobby.
Now I’ve got a hard copy of just about any album I want to listen to. I like the physicality of taking the CD out of the case and nudging the sliding tray in. My CD player is a 1988 Realistic I found at a music store and it sometimes skips if the disc has deep scratches. I like the unspoken social job of turning the record over. I like that the current CEO of Spotify, and the next CEO, and whatever government wants to, can’t see what I am playing on repeat today.
Since returning to mainly listening to physical media, my experience of listening to music has improved and I enjoy it more than ever. I got home the other day and the internet was out, and I didn’t even notice for an hour since the first thing I do now when I get home is put a record on
I think we ought to respect that, and treat suggestions to “improve” old literature by “updating” the language with the same mild derision that’s useful on those loons that wanted to paint over the cigarettes in old movies
The primary global benefit of teaching, in my experience, is that it keeps us jaded experts constantly at the entry point to our field. Close to the basic, foundational material. Making sure that it’s easy for new people to join, because that’s how the field grows and it’s good for everyone in it.
The primary local benefit, to the teacher, is that you have moments like this that allow you to constantly refine your worldview to be more congruent with reality. This allows you to better define what should be taught, and how.