I'm still waiting for meeting transcripts that understand who is speaking. I'm legitimately surprised with how far we've come with speech recognition, how this fairly common use-case is omitted.
I'm not even saying it needs to name the people in the meeting. Just understand, contextually, if it is from "person 1" or "person 2." Then as it records associate it with that name.
Maybe this can help? But Google's existing APIs might be able to do this.