CUDA can be difficult to set up correctly on your personal development machine and if you rely on it you're also limiting the development machines that other people can use. You're limiting the deployment options and the CI options. This may differ if you work at a larger company that has a team specializing in these things, but I had to work out everything from initial proof-of-concept to final deployment. Some applications absolutely need the higher performance from GPU execution but it's worth seeing if you can get away with CPU-only execution because it avoids operational complications.
I used the ONNX runtime (as this DeepFilterNet project does, indirectly, according to a top level comment by WiSaGaN) and I was able to get adequate inference speed running on plain CPU. It's a small Python service that just wraps the inference logic with a command interface and a connection to Redis. It takes protocol buffer inputs from a Redis based queue, does a little bit of control logic followed by inference, and writes the results as protocol buffers to another Redis based queue.
The only ML libraries I have on the Python side are the ones for ONNX. This saved over a gigabyte (!) of transitive dependencies compared to the first proof-of-concept I had that relied on PyTorch for runtime inference.
Final advice: you actually don't need to know much theory to start doing something useful. I hadn't studied neural networks since graduate school 20 years ago so my theory is hopelessly outdated. I just started hacking together a little demo for myself and it was good enough that I was encouraged to take it all the way to production.
You can even have dependency groups, to separate main/dev dependencies for instance. It also brings env management, and plays very nicely with Docker if you use containers.
wtf = { file = "omg.whl" }https://download.pytorch.org/whl/torch/
It's meant to be an automated process, don't make excuses for poor implementation on behalf of Poetry.
See for yourself the number of torch related issues:
https://github.com/python-poetry/poetry/issues?q=is%3Aissue+...
https://github.com/python-poetry/poetry/issues?q=is%3Aissue+...
Not a week goes by without a new issue popping up. A dependency manager is a critical piece of infrastructure, it should not be the main developer experience bottleneck when building an application. A good package manager is out of sight, out of mind. You don't hear people complaining so much about Cargo every day.
Admittedly only a single nit but it was still one of thr saddest most frustrating python experiences I've ever had.