> For me the hardest part of learning ML was getting over imposter syndrome. It felt like I needed a PhD and hardcore math skills
ABD (all but PhD dissertation) here with strong math skills. I get the imposter syndrome, but let me absolutely assure you that the community at large does not have strong math skills. I routinely talk to people doing diffusion research that don't know what covariance is or pdf. People from top ranked schools, with high paper counts and high citation counts. Expertise is often more narrow than it appears. That's okay, as long as we're honest about it.
Don't get me wrong, I wish there was more math involved and efforts were more serious. But they aren't. The space is very noisy and little is being done to clean it up (there are some, and I do appreciate those efforts). I'll add that there's one thing more that you need besides motivation and patience: perseverance. ML systems are hard to debug and difficult to evaluate (maybe not for papers, but absolutely for systems that work in the real world). It's okay to not get things perfectly and it is totally okay to not have a model with decent generalization, but context is always important and part of the debugging process is trying to trace these down (which is difficult because you need to do more abstract versions of what is analogous to the silly or random inputs being passed to code). Detecting overfitting is often quite hard and honestly sometimes it is even desirable (GPT being overfit makes it great for information retrieval!).
Also, something I tell my students when I teach ML: you don't need math to train good models, but you do need math to know why your models are wrong. So I highly encourage math, but don't let that stop you from getting started. You can also just have a math heavy person on your team and get many benefits that way.