Creating the largest protein-protein interaction dataset in the world
owlposting.com
owlposting.com
Second, the O-glycans are completely different to humans. Unless you’re looking at alpha-DG, and a handful of other proteins, you’re going to get the wrong glycosylation. This is a problem for two reasons: a) the alpha-Mannose does completely different things to the protein backbone compared to alpha-GalNAc, or probably alpha-Fucose etc, and b) those yeast PMT enzymes don’t seem to care where they throw sugars on, so they’re going to probably glycosylate something that shouldn’t be glycosylated.
This is to say nothing about the different suite of PC-processing enzymes and zymogen activation in yeast too.
So here’s my free solution to solve this: On all mated cells, do a ConA enrichment and identify where there is O-glycosylation (mass spectrometry). If it’s on your target protein, drop the data?
But otherwise, if you are interested in yeast interaction, looks like a cool technique!
I skimmed the article and start-up website and I'm a bit confused.
PPI inference is not binding affinity prediction is not binding site prediction, despite being related tasks.
There are billions of PPI pairs in public datasets, there is much less binding affinity data, and even less binding site data.
(Side-note: if you're hiring PPI / deep learning / comp. bio people, send me an email at the address in my bio.)
There's oodles of labelled images of dogs, but comparably much fewer datasets of dog silhouettes.
Another factor is that it's far easier (and less informative) to predict that two proteins are capable of interacting with any degree of affinity than with a specific amount.
You may say to yourself (as I once did): "Well surely a well calibrated PPI inference model will output interaction probabilities that correlate with binding affinity!"
I've tested this and I've yet to find one written by myself or others that behaves this way.
If this line of questioning is interesting to, definitely sign up for Google Scholar alerts on my name because we're publishing some very cool stuff on precisely this v. soon.
Anyways I look forward to seeing what you publish!
Instead, we can focus on gene-gene interactions and go bottom up. There, we don’t need new wetlab techniques that needs to be validated, measure mRNA instead. Plus, an average single cell contains 40 million proteins, yet the number of mRNA molecules are orders of magnitudes less and can be sequenced with high precision.
If you for instance open KEGG database in graph mode, you will see one of the largest manually curated datasets ever. Yet it is still tiny! If you imagine A,T,G,C as alphabet, and genes as special tokens. All we know as humanity is couple of words… and it’s sad.
I think LLMs might be our best bet on these. Given few words they might uncover “new words” we have never thought about. I kinda tried this… but the methodology is still shaky.