The moment you try to reproduce in local env, you'll be greeted with many "non-existent and unmatched dependency" errors.
Also, "examples" do not cut it.
The moment you try to reproduce in local env, you'll be greeted with many "non-existent and unmatched dependency" errors.
Also, "examples" do not cut it.
It is true that the model is based on PyTorch + python, but the majority of complexity (like SSML parsing) is tucked inside of the model.
Theoretically one can make a simplified model without any of those features in plain PyTorch or ONNX, but so far we did not have proper motivation to do.
As for CLI, this also seems simple enough, but out of scope for us.
Let me make it "embarrassingly simple" for you:
/bin/bash text_to_speech.sh file.txt file.wav
Also, I'm not entirely sure what "out of scope" mean?Do you mean you run your software on computers that can't run bash?
Do you develop machine learning algorithms on your phone?
It is explicitly stated, that PyTorch is the only real requirement. Bash is not required, i.e. models can be run on Windows or ARM with PyTorch.
> Also, I'm not entirely sure what "out of scope" mean?
There was no tangible benefit in making a bash CLI for us.
3/4 times I try to make/rebuild a Colab based demo from scratch in a suitable non Colab environment… the setup instructions are caring degrees of wrong. From the little mistakes like under specific requirements that are now broken due to transient dependency changes, to completely wrong because everything has changed to the absolute worst version of all, the never even written down.
I find Colab is a subtle form of lock in by providing useful crutches … by leaning on the crutches of Colab handing all this hard dependency and environment management stuff you never need to learn how to do it any better than necessary to function on Colab… to draw a somewhat nasty analogy using terminology from the DevOps world, good dependency and build tools make a folder full of code like cattle, you can blow it away and rebuild it when you want, but Colab let’s you raise a pet by hand and then just magically clones it whenever you or someone else need a copy.
1. They assume everyone has access to the same environment they do
2. They often don't understand anything about the infrastructure that's running their stuff
3. They produce very interesting work (such as this particular TTS work)
4. They drop 90% of their potential audience within 5 mn because the bloody thing lives in a weird cloud-only environment or requires a nightmarish stack of dependencies to run on a local machine and basically can't be simply integrated in a larger pipeline (e.g. a simple shell script).
My experience has been that getting ML researchers to get their head out of colab's ass and learn to type things like "ls" and "cd" is really hard.Colab is far worse, they say “don’t worry about it, you can just clone” and it’s just been a toxic spill, rotting away at the level of understanding in the ML community. Colab let’s you basically never put any effort into management of setup, dependencies, or data, and consequently it’s both amazing and fucking horrible the moment you want to avoid using it because everyone just builds their project “leaning on” the capabilities of Colab… it’s built an entire shanty town of poorly managed ML projects leaning precariously against the supports provided by Colab.
I’m just glad it hasn’t sucked too much air out of Jupyter in the ML community because at least stock Jupyter based tools are easy enough to take apart and reverse engineer since it’s a normal Python ecosystem, no magic Google drive data links, no custom Google tensor unit specific libraries, no push button magic clones of entirely hand crafted environments.
Interesting project btw, kudos to the dev.
Can't tell if serious or sarcasm.