The former would be quite difficult without much background in ML/Computer Vision (you would have to spend some time self-teaching basics of ML/Deep Learning and the pre-reqs for those — Basic Linear Algebra and Probability).
The latter is doable. I would recommend a very hands on approach. Pick some computer vision object classification tutorials and code them up (using a high level library). Make a mind map of the concepts and look them up as and when you’re unclear about a concept. Then move on to replicating some well cited, peer reviewed papers. Often papers will have their code on GitHub. Try and relocate their results on their dataset. After this you would have the basic working knowledge to modify the algorithm slightly for your specific use case.
Looks like it's a new model, I have no idea if they already have any ML models yet. There's also some database work.
I'm finishing a Masters degree in Computational Physics, so Linear Algebra and Probability shouldn't be an issue. (We also have an Image Processing and Analysis course.) I guess that's why they contacted us despite the fact that we don't have any ML training ?
Yeah, this is basically what I thought to do, but thank you for your advice !
Also, might be useful to took at webpages of some researchers in this space and courses they teach [1,2].
[0] https://web.stanford.edu/~hastie/ElemStatLearn/
[1] https://scholars.duke.edu/person/dunson
[2] https://www.cs.princeton.edu/~bee/Funny (but I guess expected) to see the Markov Chain Monte Carlo method that we very recently learned in that book's table of contents ! (Unless it's another MCMC ?)