This is amazing, even the breathing can be seen. Next big step (or may be its done already) is to generate expressive audio and we are all set for a generated model you can have a video call with.
https://suno-ai.notion.site/Bark-v0-Examples-e572bcfcdf65429...