Talk-Llama
github.com
github.com
The performance on Apple Silicon should be much better today compared to what is shown in the video as whisper.cpp now runs fully on the GPU and there have been significant improvements in llama.cpp generation speed over the last few months.
And impressive performance indeed!
On a different note:
I think the most useful coding copilot tools for me reduce "manual overhead" without attempting to do any hard thinking/problem solving for me (such as generating arguments and types from docstrings or vice-versa, etc.). For more complicated tasks you really have to give copilot a pretty good "starting point".
I often talk to myself while coding. It would be extremely, extremely futuristic (and potentially useful) if a tool like this embedded my speech into a context vector and used it to as an additional copilot input so the model has a better "starting point".
I'm a late adopter of copilot and don't use it all the time but if anyone is aware of anything like this I'd be curious to hear about it.
I believe that is the relevant section, which I am hoping they realize how dumb this is going to be.
Not even Oracle dared to pull shit like this w/ Java.
But this would be akin to them saying "you know, bytecode could be used for evil, and we'd to regulate/outlaw the development of new virtual machines like the one we already have".
P.S. I'm from one of the said countries
this cat is already out of the bag, so this is pointless legislation that just hampers progress
I have already had good-actor success with uncensored models
This isn't legislation.
And its not proposing anything except soliciting input from a broad range of civil society groups and writing a report. If the report calls for dumb regulatory ideas, that’ll be appropriate to complain about. But “there’s a thing that is happening that seems like it might have big effects, gather input from experts and concerned parties and then writeup findings on the impacts and what, if any, action seems warranted is... not a particularly alarming thing.
This is insanity. I have to be missing something, what else do the safeguards prevent the LLM from doing? This has to be more about this than like preventing an LLM from using bad words or showing support for Trump...
How dumb its going to be to... solicit input from a wide range of different areas and write a report?
This proposal is the US doing that bicycle/stick meme, it will backfire spectacularly.
Now it turns out they were just bullshitting at Google.
I don't think Google was bullshitting when they wrote, documented and released Tensorflow, BERT and flan-t5 to the public. Their failure to beat OpenAI in a money-pissing competition really doesn't feel like it reflects on their capability (or intentions) as a company. It certainly doesn't feel like they were "bullshitting" anyone.
> The innovation is largely happening within the megacorps anyway
That was the part I was replying to. Whichever megacorp this is, it’s not Google.
The flan quantizations are also still pretty close to SOTA for text transformers. Their quality-to-size ratio is much better than a 7b Llama finetune, and it appears to be what Apple based their new autocorrect off of.
Given current data volume used during the training phase (tb/s), I highly doubt it's possible without two, magnitude changing, breakthroughs at once
I agree that it is concerning the way it's open-ended. But where is the actual outlawing?
(I can imagine recommendations that benefit incumbents by, for instance, placing such high burden on government adoption of open weight models that OpenAI is a much more attractive purchase. But that’s not the same as what you’re talking about.)
I dunno, the EO seems pretty easy to read. Am I missing something in the text?
https://www.whitehouse.gov/briefing-room/presidential-action...
Disobey.
So, literally, they can’t enforce it without consulting with industry, since enforcement is just sonewhat in the government holding someone else in government accountable for consulting with, among others, the industry.
That is beautiful, I made you a shirt. https://sprd.co/ZZufv7j
I'm not sure what changed, but basically I purged ffmpeg and libsdl2-dev and the `make` in the root of the repo. Then I installed libsdl2 and ffmpeg and `make talk-llama`.
It's quite slow on 4 core i7-8550U and 16 GB of RAM.
basically, in the root of the repo:
$ sudo apt purge ffmpeg $ make clean $ git pull $ make $ sudo apt install libsdl2-dev $ make talk-llama $ ./talk-llama -mw ./models/ggml-small.en.bin -ml ../llama.cpp/models/llama-2-13b.Q4_0.gguf -p "t0mk" -t 8\n\n
HTH
I guess it'd only work if the model can keep the buffer filled fast enough so the tts engine doesn't stall.
I know wish for a model built to be a low-computation filter that takes text in and produces padded text out intended for TTS and annotated with pauses or sounds and extra words that maintains the same meaning but provides the ability to dynamically adjust the level of verbosity to maintain a fixed rate of words per minute.
Besides humans do this all the time, we start saying words before we even have, uhh, any idea how we're gonna end the sentence and for the most part it, uhh, works out. Should be doable.
Unless you're Michael Scott.
- better detection of when speech ends (currently basic adaptive threshold)
- use small LLM for quick response with something generic while big LLM computes
- TTS streaming in chunks or sentences
One of the better OSS versions of such chatbot I think is https://github.com/yacineMTB/talk. Though probably many other similar projects also exist by now.
Can't wait for poorly implemented chat apps to always start a response with "That's a great question!"
I think Wizard is the “meta” for technical questions now.
pacman -S ollama
ollama serve
ollama run llama2:13b 'insert prompt'
https://ollama.ai/They're hiring and have no currently disclosed monetization strategy so I expect a rugpull soon where some now-free feature gets paywalled or purposefully crippled, but it's not like porcelain apps for free LLMs that rely entirely on llama.cpp to function can do vendor lock-in. I'd second ollama if OSS is a higher priority than features though.
Also I saw this thing called WhisperScript that looks pretty slick: https://github.com/openai/whisper/discussions/1028
That being said, WhisperX isn't that hard to setup. My step by step from a couple months ago: https://llm-tracker.info/books/logbook/page/transcription-te...
For TTS coqui has the best UX and models for a lot of languages although quality is not on par with commercial TTS providers.
* Model: Multilingual v2
* All options and sliders to boost similarity: set to max/yes
* Stability slider: experimentally set to a value where the model sounds natural enough without destabilising sound output
It's Speech Recognition -> Llama -> Text to Speech, running on your own PC rather than that of a third party.
The limitations on the context of the LLM are that of the model being used, e.g. Llama 2, Wizard Vicuna, whatever is chosen, in whatever compatible configuration is set by the user regarding context window etc, and given a preliminary transcript (as the LLM doesn't "reply" to the user in a sense, it just predicts the best continuation of a transcript between the user and a useful assistant, resulting in it successfully pretending to be a useful assistant, thus being a useful assistant - it's confusing).
I can imagine that it's viable to get that kind of behaviour by modifying the pipeline.
If the architecture was instead Speech Recognition -> Wrapper[Llama] -> Text 2 Speech, where "Wrapper" is some process that lets Llama do its thing but hooks onto the input text to add some additional processing, then things could get interesting.
The wrapper could analyse the conversation and pick out key aspects ("The person's name is Bob, male, 35, he likes dogs, he likes things to be organised, he wants a reminder at 5pm to call his daughter, he is an undercover agent for the Antarctic mafia, and he prefers to be spoken to in a strong Polish accent") and perform actions based on that:
- Set a reminder at 5pm to call his daughter (through e.g. HomeAssistant)
- Configure the text-2-speech engine to use a Polish accent
- Modify the starting transcript for future runs:
- Put his name as the human's name within the underlying chat dialogue
- Provide a condensed representation of his interests and personality within the preliminary introduction to the next chat dialogue
This way there's some interactivity involved (through actions performed by some other tool), some continuity (by modifying the next chat dialogue) and so on.No idea what sort of data structure could work. Perhaps a graph database could be feasible, and the memory prompt could instruct it to write a query for the given database.
I was surprised they didn’t combine this work with the streaming whisper demo. So I guess I will implement that for iOS/macos (streaming whisper results in realtime without waiting on an audio pause, but as you say using the audio pauses and other signals like punctuation in the result to determine when to llm complete; makes me also wonder about streaming whisper results in to the llm incrementally before ready for completion)
Now perhaps this is a skills issue on my part (because I'm not a Python dev), but I've had endless trouble with Python-based ML projects, with some requiring I use/avoid specific 3.x versions of Python, each project's install instructions seemingly using a different tool to create a virtual environment to isolate dependencies, and issues tracking down specific custom versions of core libraries in order to allow the use of GPU/Neural Engine on Apple Silicon.
The whisper and llama.cpp projects just build and run so easily by comparison
Having said that, I have the sense that the ML ecosystem is coalescing around using `venv` as a standard for Python dependencies. Most of the build instructions for Python ML projects I've seen recently begin with setting up the environment using venv, and in my experience, it works fairly reliably. I don't particularly like downloading gigabytes of dependencies for each new project, but that mess of dependencies is what's powering the rapid pace of prototypes and development.
CUDA for example, different project will require different versions of some library like pytorch, but these seem to be tied to cuda version. This is where anaconda (and miniconda) come in, but omfg, I hate those. So far all anaconda does is screw up my environment, causing weird binaries to come into my path, overriding my newer/vetted ffmpeg and other libraries with some outdated libraries. Not to mention, I have no idea if they are safe to use, since I can't figure out where this avalanche (literally gigs) of garbage gets pulled in. If I don't let it mess with my startup scripts, nothing works.
And note, I'm not smart, but I've been a user of UNIX from the 90's and I can't believe we haven't progressed much in all these decades. I remember trying to pull in source packages and compiling them from scratch and that sucked too (make, cmake, m4, etc). But package managers and other tech has helped the general public that just wants to use the damn software. Nobody wants to track down and compile every dependency and become an expert in the build process. But this is where we are. Still.
I am currently in the trying to get these projects working in docker, but that is a whole other ordeal that I haven't completed yet, though I am hopeful that I'm almost there :) Some projects have Dockerfiles and some even have docker-compose files. None have worked out-of-the-box for me. And that's both surprising and sad.
I don't know where the blame lies exactly. Docker? The package maintainers that don't know docker or unix (a lot of these new LLM/AI projects are windows only or windows-first and I hear data scientists hate-hate-hate sysadmin tasks)? Nvidia for their eco-system? Dunno, all I know is I'm experiencing pain and time wastage that I'd rather not deal with. I guess that's partly why open-ai and other paid services exist. lol.
That helps keep your system clean, but someone with big $s please rewrite pytorch to golang or rust or even nodejs / typescript.
So, as someone who has never gotten around to doing this, and who also likes not having to deal with the Python tools, it's not quite that simple. Steps I had to take for talk-llama after cloing whisper.cpp:
* apt install libssdl2-dev (Linux; other steps elsewhere)
* make talk-llama from the root of whisper.cpp, not from
* ./download-ggml-model.sh small.en from the models directory
* Tried to run it with the command line in the README, have it seg fault after failing to open ../llama.cpp/models/llama-13b/ggml-model-q4_0.gguf, cloning llama.cpp, finding the file is not in the repo.
* Searching through the readme for how to find the models, and finding I need to go searching elsewhere because no urls were listed.
* Having to install a bunch of Python dependencies to quantize the models...
This is still far from "build and run". Though I will fully believe that a lot of the Python-based ML projects are worse.
I was able to get llama.cpp itself to work, though, including image analysis.
[0]: *Crustlefuck (noun)*: A labyrinthine, solidified matrix of chaos that has accreted over time. A crustlefuck differs from a clusterfuck due to an added dimension of rigid, ingrained complications that render any attempt at untangling the issues extraordinarily daunting and taxing.
4. there are llama.cpp ways to do pretty much all that it does and you aren't just restricted to a couple model types like memgpt.
say I have a research paper in pdf, can I ask llama questions about it?
If your model has a long enough context (models like MistralLite can do 32,000 tokens now, which is about 30 pages of text) you could run a PDF text extraction tool and then dump that text into the model context and ask questions about it with the remaining tokens.
You could also plug this into one of the ask-questions-of-a-long-document-via-embedding-search tools.