Now perhaps this is a skills issue on my part (because I'm not a Python dev), but I've had endless trouble with Python-based ML projects, with some requiring I use/avoid specific 3.x versions of Python, each project's install instructions seemingly using a different tool to create a virtual environment to isolate dependencies, and issues tracking down specific custom versions of core libraries in order to allow the use of GPU/Neural Engine on Apple Silicon.
The whisper and llama.cpp projects just build and run so easily by comparison
So, as someone who has never gotten around to doing this, and who also likes not having to deal with the Python tools, it's not quite that simple. Steps I had to take for talk-llama after cloing whisper.cpp:
* apt install libssdl2-dev (Linux; other steps elsewhere)
* make talk-llama from the root of whisper.cpp, not from
* ./download-ggml-model.sh small.en from the models directory
* Tried to run it with the command line in the README, have it seg fault after failing to open ../llama.cpp/models/llama-13b/ggml-model-q4_0.gguf, cloning llama.cpp, finding the file is not in the repo.
* Searching through the readme for how to find the models, and finding I need to go searching elsewhere because no urls were listed.
* Having to install a bunch of Python dependencies to quantize the models...
This is still far from "build and run". Though I will fully believe that a lot of the Python-based ML projects are worse.
I was able to get llama.cpp itself to work, though, including image analysis.
Having said that, I have the sense that the ML ecosystem is coalescing around using `venv` as a standard for Python dependencies. Most of the build instructions for Python ML projects I've seen recently begin with setting up the environment using venv, and in my experience, it works fairly reliably. I don't particularly like downloading gigabytes of dependencies for each new project, but that mess of dependencies is what's powering the rapid pace of prototypes and development.
CUDA for example, different project will require different versions of some library like pytorch, but these seem to be tied to cuda version. This is where anaconda (and miniconda) come in, but omfg, I hate those. So far all anaconda does is screw up my environment, causing weird binaries to come into my path, overriding my newer/vetted ffmpeg and other libraries with some outdated libraries. Not to mention, I have no idea if they are safe to use, since I can't figure out where this avalanche (literally gigs) of garbage gets pulled in. If I don't let it mess with my startup scripts, nothing works.
And note, I'm not smart, but I've been a user of UNIX from the 90's and I can't believe we haven't progressed much in all these decades. I remember trying to pull in source packages and compiling them from scratch and that sucked too (make, cmake, m4, etc). But package managers and other tech has helped the general public that just wants to use the damn software. Nobody wants to track down and compile every dependency and become an expert in the build process. But this is where we are. Still.
I am currently in the trying to get these projects working in docker, but that is a whole other ordeal that I haven't completed yet, though I am hopeful that I'm almost there :) Some projects have Dockerfiles and some even have docker-compose files. None have worked out-of-the-box for me. And that's both surprising and sad.
I don't know where the blame lies exactly. Docker? The package maintainers that don't know docker or unix (a lot of these new LLM/AI projects are windows only or windows-first and I hear data scientists hate-hate-hate sysadmin tasks)? Nvidia for their eco-system? Dunno, all I know is I'm experiencing pain and time wastage that I'd rather not deal with. I guess that's partly why open-ai and other paid services exist. lol.
That helps keep your system clean, but someone with big $s please rewrite pytorch to golang or rust or even nodejs / typescript.
[0]: *Crustlefuck (noun)*: A labyrinthine, solidified matrix of chaos that has accreted over time. A crustlefuck differs from a clusterfuck due to an added dimension of rigid, ingrained complications that render any attempt at untangling the issues extraordinarily daunting and taxing.
4. there are llama.cpp ways to do pretty much all that it does and you aren't just restricted to a couple model types like memgpt.