> your services are required on the busy beaver thread!
Lol I didn't even see it. I'm assuming this is w.r.t Mutual Information's video?
> the phd, even though i'm not done yet
I'm at about a similar point (last year). Most of my cynicism though is around academia and publishing. Making conferences the de facto target for publishing was a mistake. Zero-shot submissions in a zero-sum game environment? Can't see how that would go wrong...
> it's not much different in physics, where the good experimentalists aren't born that way, they're made in the well-funded labs.
Coming from the experimental physics side (my undergrad), there is a big factor though. Generally the experimentalists who were good at the math and could learn to intuit them (and especially the uncertainty) did better. But you're absolutely correct about the __well funded__ part being a big indicator. When I've worked at gov labs I didn't notice a quality difference in intellect between peers from different schools (of a wide variety of prestige) but what did stand out was simply experience. Your no-name school physicists could pick up the skills fast, but they just never had opportunities like the prestigious school students did. It didn't make too big of a difference, but it is an interesting note, especially since it tells us how to make more of those higher status researchers...
> i firmly believe that in ML, the math does not matter at all
My opinion is that this is highly context dependent. Most research right now is about optimization and tuning, and with respect to that, I fully agree. I'm including in that even some architecture search, such as "replace CNN with Transformer" and such things. This you can do pretty much empirically. The only big point I'll get on here is that people do not understand the limitations of their metrics (especially parametric metrics), biases of the datasets, and the biases of their architectures, so it creates a really weird environment where we aren't comparing things fairly. (It is also why what works in research doesn't always work out well in industry) But if we're talking about interpretability, understanding, novel architecture design, evaluation methods, and so on, then I do think it matters. There's a lot that we can actually understand about ML -- how they work and how they form answers -- that isn't discussed not because it hasn't been researched but because the research has a higher barrier to entry and people don't even understand the results. It isn't uncommon to see a top tier paper empirically find what a theoretical paper (with experiments, but lower compute) found 5-10 years back, where the recent work didn't even know about the prior work. Where the higher level math really helps out is being able to read deeper and evaluate deeper. Fwiw, every time I make a "math is necessary" argument, I get a lot of people pushing back. But I think this is because both groups have two camps. For the pro math I believe there is people who legitimately believe it (like me) -- who usually talk about high dimensional statistics and other things well past calculus -- and people who say it to make themselves feel smart -- people who often think calculus is high level math or say "linear algebra" as if it is just what's in David Lay's book. For the anti-math crowd I think there are the hype people who just don't care and the people who are just doing other things and don't really end up using it. For the latter, I do think they are still benefiting a lot from the intuition about these systems that they gained from those math courses. But then again, the classic thing in math education is that you struggle while you learn it and then after you know it it is trivial.
For research, I firmly believe you need both the high math and the "knob turners." I just think academia and conferencing should be focused around the former and industry should focus around the latter. But the problem is we have these people operating in the exact same space and we're comparing works that use 100+yrs of compute hours to works that have a month or two of compute hours. This isn't a great way to really tell if one architecture is better than another since hyperparameters matter so much. It's just making for bad research and railroading.