The Most Cited AI Papers in 2022
zeta-alpha.com
zeta-alpha.com
I might be dumb but could they scale it up and make an alphafold 3 with maybe like 10bln params? Would it be a lot better assuming the same training effort is put into it?
If it does, can't biotech companies just go nuts and make a 100bln params internal model and have all the protein structures they want?
I'm wondering if there could be a path towards a mix of alphafold and folding@home, with donated idle compute resources being used to train/run the models.
Designing for that sort of fragmentation could also make it easier to slowly run oversized models on local machines with swapped memory.
Petals: Collaborative Inference and Fine-tuning of Large Models Alexander Borzunov, Dmitry Baranchuk, Tim Dettmers, Max Ryabinin, Younes Belkada, Artem Chumachenko, Pavel Samygin, Colin Raffel https://arxiv.org/abs/2209.01188
Would be wonderful to use this host this model LLaMA: Open and Efficient Foundation Language Models https://arxiv.org/abs/2302.13971
Perhaps if there was a way to combine mining and folding, to allow participants to somehow gain a share of the output? Eg each folded protein would have a unique hash, which could then be traded?
And yes, I hate everything about what I just typed.
Very interesting. I didn't think of connecting these two dots. Let's see if it can be applied to other CV tasks, like object detection.
[1]: High-Resolution Image Synthesis With Latent Diffusion Models, https://arxiv.org/abs/2112.10752
For example, the paper "ColabFold: making protein folding accessible to all" is listed as having 1162 citations. I'm seeing that it was cited only by 899 publications on Scite: https://scite.ai/reports/colabfold-making-protein-folding-ac...
I'm wondering if Google Scholar is overestimating or Scite is underestimating.
I tend to trust Semantic more than GS. GS tends to overestimate. For example on GS I have 164 citations on one paper and semantic says 150. FWIW Scite says 49.[1]
[0] https://www.semanticscholar.org/paper/ColabFold%3A-making-pr...
[1] I'll note that this paper is an arxiv paper and has not been accepted at a conference but I'd also argue that conference acceptance means little in ML. I'll explain if anyone is actually concerned with the claim.
The traditional publication flow just isn't useful if a runnable demo on HuggingFace explains your work way better than 3 pages of formulas.
First, this peer review via conferences/journals/etc is relatively new in the scientific process. Really only the last 50 years has this paradigm been the main way for publishing. Prior to that scientists have just published in the open and and peer review happened by peers reading and responding. Not too different from what we see with arxiv, twitter, and blogging.
Second, we need to talk about how good the review process actually is. There's been a lot of writing on the NeurIPS experiments [0] is the most famous one. But the Google paper[1] notes that reviewers are "good at identifying bad papers but not good at identifying good papers." I'll go a step further than them and suggest a plausible model that makes this statement true: reviewers are reject happy. We need a confusion matrix to really see this but if you reject every paper you'd have a 100% success rate of rejecting bad papers but a 0% success rate of approving good papers. We have a good demonstration that ML conferences (journals aren't our priority like other academic areas, conferences are. This is an oddity) are an extremely noisy process and not very meaningful.
So how do we capture a signal in this noisy process? Citations are at least some signal. Obviously this isn't a fantastic signal either because big labs and companies are going to be able to popularize their work more and this will get more citations. But this still isn't any worse than we were 100 years ago. I'd argue that the noisy process of conferencing is worse than where we were 100 years ago (democratization of science aside).
Unfortunately, the only way to identify if a paper is good is to have experts evaluate them. I don't think we have a good alternative for this and adding significantly noisy signals aren't helpful.
[0] https://blog.mrtz.org/2014/12/15/the-nips-experiment.html
It's a shame because I think that other areas offer much more interesting and technical questions, but many new researchers are only exposed to mainstream deep learning because the enormous hype drowns out everything else.
But in terms of groundbreaking I do think "Zero-Knowledge Proofs for Machine Learning"[1] from 2020 has the potential to unlock some really revolutionary applications in the sense of "things that were not really possible without it".
Whether you count robotics as AI is up to debate, I suppose
They are very nascent though and it isn't surprising they aren't yet highly cited since it's quite an achievement getting anything working at all still.
https://twitter.com/ZetaVector/status/1631590029926494211?s=...
the list biases against papers published later in the year.
this is a problem raised by openai and brain researchers: https://twitter.com/_jasonwei/status/1631935794301771777