You can use Machine Learning tools to isolate the vocals from the original song. I've done this exact thing (song recreation/transcribing) as part of taking music theory lessons.
For anyone else who would like to try this, I will share how to do it:
1. Use one of the popular models for Source Separation. Facebook's is called "Demucs" and one of the best. Deezer's is called "Spleeter", also very good, and then there's "OpenUnmix".
There's a webservice to separate a song using Demucs here:
https://demucs.danielfrg.com/
And a great Dockerized webapp that lets you choose from several models and parameters here:
https://github.com/JeffreyCA/spleeter-web
Otherwise you can just install them locally and run them through the CLI, it's pretty easy (one command)
https://github.com/facebookresearch/demucs#for-musicians
https://github.com/deezer/spleeter#quick-start
2. Now you can take the isolated vocals, and build the rest of the song yourself
I've also found that it really helps to be able to listen to each of the parts of the song in isolation when trying to recreate them.
The drums, the bassline, the lead/synth, etc. It can be hard to distinguish notes when they all run together. You can use an EQ to try to single out instruments but it's harder.
Hope this is helpful to someone.