> The field of ML research is moving so fast that people don’t even take time anymore to explain the design decisions behind their architectures.
Has it ever? I've read a lot of "old" papers and idk if there's ever been a strong theoretical framework. I'd say it is strongest today but not in main papers and those papers are getting rejected because of lack of experiments (wtf is going on with reviewing and why are our solutions "authors have to show reviewers used LLMs"[-1] not "reviewers have to provide meaningful review"? What a joke!). Here's a more recent Universal Approximation paper[0]. Cyberko paper[1] is good, but I think we're moving in the right direction, at least with respect to these two works. Other UA papers I've seen are pretty handwavy.
But simple also doesn't mean "not useful" or "not worthy of publish" and I think this is something people are forgetting lately. Probably because we're asked to review too much and annoyed, but don't take it out on authors. Some examples are GELU[2] (read the end, there's drama in this paper), Clean-FID (CVPR 22)[3], ViT (ICLR 21) and DeiT (ICML 21)[4,5], and even ResNet (CVPR 16)[6] and tons of ReLU papers we can list. But let's also not confuse big compute with "good work" (they correlate but aren't the same. You can definitely do more with more compute).
> he problem with this for me is that we fail to build a nice, crisp understanding of the effects of each design decision in the final outcomes, which hurts the actual “science” of it. It also opens up the field for bogus and unreproducible claims.
This is mostly an unsolvable problem in science and I think we actually have bigger problems. This can be solved without conferences due to our community being fairly good about open sourcing code and checkpoints (at least compared to any other research area), as well as just the papers themselves. Reproducibility is only able to be done when people can actually replicate works (lol sorry GPT papers...). What helps this is moving away from publish or perish or allowing people to publish works that are replication attempts (replication is the foundation of science. Why does everything have to be "novel", whatever that means).
The bigger problem I see is actually related to the above. Publishing is just insane. It is easy, especially in fast moving times, to claim any work is incremental or not novel. That it's "just x but y". My feelings are "so what?" ViT is "just a transformer on images but image patches are tokens," yet it's insanely useful. It's obvious post hoc but not a priori. If it was, it would already be in use (ViT actually being a good example given the timeline). Everything is simpler and much more obvious after you've already seen it, which is why the whole notion of novelty is ridiculous. We humans just rewrite history in our minds and we also only concentrate on what's popular (lots of incentives, especially with reviewing insanity). Poor Ross Wightman gets shadowed for replications and improving ResNets though (and poor ConvNext). But we're seeing big labs do this and get through because they can do lots of experiments and build a lot of empirical evidence towards work, despite no additional theory.
The problem? What about the academics? I don't know how people publish without connections to big labs (aka big compute) anymore. The ideas from the groups are the same (industry is a bit more hyped), but I've seen papers where the theory is good, the architecture makes sense (and is well explained) and my co-reviewers for a fucking workshop want to reject it because "not enough experiments" and "only one 'real world' experiment." These are papers that are on par or better than many I see in the main conferences (you can tell they're shifted to workshops due to rejects). Talk about holding back science. Holding back science is making authors submit works over and over and over to get orthogonal reviews every time. Holding back science is me arguing with my co-reviewers who admitted to not understanding the work that "more experiments" is not a sufficient reason to reject from a workshop (or a main conference).
Hell, I got rejected once for a distillation paper and the reviewers complained that my teacher model's performance didn't match that of the original paper (no one has replicated that work btw, and I beat the second best published result I could find). Another paper I got hard rejected because I redacted a github link and an appendix citation broke when splitting the work. Guess what the ACs did? They let that person write a stronger response and then because they were the only high confidence reviewer on that paper (others were 3/5 confidence and borderline/WA) and so rejected the work. What kind of insane system is that? I see instances like this all over the place from many different peers across many different universities. There's more that's broken before we can even begin to talk about other issues because our system of "what qualifies as meaningful work" shouldn't be a fucking slot machine run by undergrads.
CVPR has well over 10k submissions. Good luck to the <500 ACs and all the reviewers. God have mercy on the grad students who are just trying to graduate. I'm sure you're all doing good work even if the reviewer didn't read a word. I hope your pictures and tables are enough.
[-1] https://twitter.com/CVPR/status/1722384482261508498
[0] https://www.jmlr.org/papers/v23/20-1433.html
[1] https://web.njit.edu/~usman/courses/cs675_fall18/10.1.1.441....
[2] https://arxiv.org/abs/1606.08415
[3] https://www.cs.cmu.edu/~clean-fid/
[4] https://arxiv.org/abs/2010.11929
[5] https://arxiv.org/abs/2012.12877
[6] https://arxiv.org/abs/1512.03385