This sounds like something you could use LLMs for as LLMs are quite good at removing bias in content. You might run into context length issues feeding the whole transcript in though .
Unless there was something like sampling, similar to how shazam recognising music from a few seconds works.
Could someone only get a portion of the content and, semi accurately generate a summary?