Sesame CSM: A Conversational Speech Generation Model
github.com
github.com
They open-sourced a crippled version of sesame (1B)
not the one they're using in actual demo
Which is a shame because seems like there's a real opportunity to really shake things up with an open voice model that's competitive with the proprietary ones.
Oh well. Someone else will do what they claimed to want to do.
You can run it with uv like this:
uv run --python 3.12 \
--with "git+https://github.com/senstella/csm-mlx[cli]" \
csm-mlx --text 'hello there' -o output.wavis there any reason why it inserts multiple-seconds-long awkward pauses into the output? are you seeing that behavior in your example? (on a 2021 M1 Max MBP)