It's a hard problem I've been working away on for a while now. It's far from perfect but every step brings it a bit closer.
It's a hard problem I've been working away on for a while now. It's far from perfect but every step brings it a bit closer.
Two more improvements, audio quality improvements which is currently in the works and close to release and a new document generation model. I'm currently using a custom fine tuned Phi-4 (released December 2024!) model, that's _so old_ in the grand scheme of LLMs, I just haven't had time to benchmark and properly test some new models whilst this currently does a good job as it is. There has to be some gains here, but who knows!
Ahh yea as for the models:
Speech to text - Nvidia Parakeet TDT 0.6b V3
Diarisation - Nvidia Marblenet for the speech detection, TitaNet-Large for the embeddings and then using NeMo multi-scale to do clustering around them
I might take another look into doing it in a different way that gradually builds up from successful calls, I just need to think of how to do this in a simple(ish) way for non-technical users and a way that still works well enough on low to mid tier laptops.
If you’re targeting data for people to use on the same call then requires more intense work, while if you’re targeting data for people to use after or on an ongoing basis (eg an established business meeting with staff/vendors/etc) then having a predefined one seems good. Same interface as a CRM: picture, some audio clips for users to choose from, and confidence scores on each section of the recording for cases where people sound similar or are talking over each other. Over time the diarization gets better as users accept/reject/tag samples and that work helps them feel more aligned with the tool.
Some great suggestions, appreciate you taking the time to write it out, it's given me some things to think about.