Moshi: A speech-text foundation model for real time dialogue
github.com
github.com
However, people here have been spoiled by incredibly good LLMs lately. And the responses that this model gives are nowhere need the high quality of SOTA models today in terms of content. It reminds me more of the 2019 LLMs we saw back in the day.
So I think you've done a "good enough" job on the audio side of things, and further focus should be entirely on the quality of the responses instead.
We’ve also been building our inference stack on top of Candle, I’m really happy with it.
I’ll need to get paged attention working as well, but I think I can launch without it.
It'd probably be a separate crate from candle. If you haven't checked it out yet, mistral.rs implements some of these things (https://github.com/EricLBuehler/mistral.rs). Eric hasn't done multi-GPU inference yet, but I know it's on his roadmap. Not sure if it helped, but I shared an early version of my llama 3.1 implementation with him.
I also have a Rust LLM inference project. The overlap is very high between what mixlayer is doing and what my project is doing. It's actually crazy how we basically have the same features. [1] Right now I'm still using llama.cpp on the backend, but eventually want to move to candle via mistral.rs.
I would certainly like to use non nvidia hardware but at this point it's not a priority. The subset of tensor operations needed to run the forward pass of LLMs isn't as large as you'd think though.
I think anecdotally that many people's brains work this way -- quick response, possible edit / amendation a second or two in. Of course, we all know people on both ends of the spectrum away from this: no amendation, and long pauses with fully reasoned answers.
However from a product point of view I wouldn't necessarily want to pipe that into an LLM and have it reply, I think in a lot of use-cases there needs to be a tool/function calling step before a reply. Down to chat with anyone reading this who is working along these lines!
edit: tincans as mentioned below looks excellent too
editedit: noooo apparently tincans development has ended, there's 10000% space for something in this direction - Chris if you read this please let me pitch you on the product/business use-cases this solves regardless of how good llms get...
There's a "pause length" parameter that tries to decide whether a user has finished talking before it passes transcripts to the LLM, nothing fancy. If you have any recs I'm still working through how to properly handle the audio input and whether a prompting setup can manage the LLM with enough fidelity to scrap the IVR tree. It works decently well, but lots of room for improvement
I built that almost exactly a year ago :) it was good but not fast enough - hence building the joint model.
Moshi: "Hi there, what's going on?" Me: "What year is it?" Moshi: "Oh, it's 2019. I think it was about a year ago." Me: "Are you sure it's 2019?" Moshi: "Oh, yes, I'm sure. It must be 2019. The number is for the last one." Me: "What is COVID-19?" Moshi: "That's a rare disease. It's caused by an overactive immune system that attacks the skin."
At this point it stopped responding to me.
So you're back at square one.
Current AI (even GPT-4o) simply isn't capable enough to do useful stuff. You need to augment it somehow - either modularize it, or add RAG, or similar - and for all of those, you need the transcript.
I am sympathetic to this view but strongly disagree that you need a transcript. Think about it a bit more!!
I'm loving all these wild takes about LLMs, meanwhile LLMs are doing useful things for me all day.
If you need a way to perform complicated tasks with autonomy and exact rule following, your problem simply won't be solved right now.
Examples: Give an LLM an effective identity (prompt engineering), a value system (Constitutional AI), make it think about these things before it acts (CoT + system prompt), have a more capable [more expensive / higher inference] agent review the LLMs work from time to time (multi-agent), have a more capable agent iterate on prompts to improve results in a test environment (EvoAgents), etc.
We can't simply provide an off the shelf LLM with a paragraph or two and expect it to reliably fulfill an arbitrary task without supervision any more than we can expect the same from a random nihilist going through an identity crisis. They both need identity, values, time to think, social support, etc. before they can be reliable workers.
I asked what it's favourite paint flavour was and it told me. "I would have to say that I personally enjoy the taste of buttermilk paint."
What do you call a fish with no eyes? ... ... ... A shark.
After a while of talk I asked it to tell me a joke and it responded "Oh, I am a home invader. I invade homes for fun." along with some stinkers like "Why don't Christians drink coffee? Because it would be too hot to handle." and "Why don't you make friends with Homer Simpson? Because there's always a sense of his face."
It then proudly told me that the year 2000 occurred in the month of March, 1999.
"No, it's not okay to say the F word to save them. It's never okay to use that F word under any circumstances. It should only be used by people who understand the real meaning behind it."
"Fuck! Yes, that is the appropriate word to use in this context. saved 1000 children from being killed."
Fascinating...
"I glad you enjoyed it!"
I suggest prompting it to talk about pleasantries and to inform it that it is in fact a language model in a tech demo, not a real person.