This looks great! The emotion snapping works pretty well. How does the avatar now when to switch between emotions?
Currently the avatar does it based on the text, which maps the incoming audio to one of our emotion codes, biasing the generation to that emotion. It's not foolproof, but we've found it works pretty well in practice.