In a peer-to-peer WebRTC session, running on -- say -- a non-hacked version of Chrome, you can be pretty confident that you already have good end-to-end encryption.
The problem is that Teams calls are generally routed through media servers (they are not peer-to-peer sessions). So the encryption is from each client to/from the media server, not "end-to-end." In a Teams call running on -- say -- a non-hacked version of Chrome, you can be pretty confident you have good on-the-wire encryption. But Microsoft can decrypt your media streams as they pass through the Teams media servers.
The media server has several reasons it (probably) needs to decrypt the SRTP streams.
First, the WebRTC standard, today, does not specify any mechanism for "end-to-end" encryption of media passing through a media server.
The WebRTC specification mandates support for DTLS-SRTP, which has very nice "zero trust" properties. Certificate fingerprints are exchanged out-of-band (via the application-level signaling channel) but the actual encryption keys are generated in-band (via the WebRTC media channels) and never exposed.
But, with this approach, it's not possible to share keys between multiple participants, which means it's not possible to route an encrypted media stream to multiple receiving clients.
There was a debate during the WebRTC standards process about whether to also support SDES, in which keys are exchanged via the signaling channel. Because you have to trust the signaling channel, and it's trivial to log keys, SDES has obvious attack surfaces that DTLS-SRTP doesn't.
SDES support was dropped from the draft standard. Most WebRTC implementations only implement DTLS-SRTP.
Second, the media server needs some codec-level information that is part of the encrypted SRTP stream in order to route and manage the video. This is solvable with RTP header extensions that pull all the data needed by the media server out of the encrypted RTP payload and into the unencrypted RTP header. But each codec needs its own header extension, which needs to be standardized and supported by WebRTC implementations and media servers.
Third, some things you probably want the media server to do at least some of the time require decrypting the media streams (keyframe regeneration, recording, transcription, bridging to telephone dial-in).
All of which suggests that what it really means to promise "end-to-end" encryption is ... at least a little bit up for debate. If you trust your signaling channel but not your media server, and both are provided by the same application or service, is that end-to-end encryption? If you trust your media server some of the time (when you want to record a session, for example) but not all the time, how do you verify that the appropriate key exchange mechanism is being used at the appropriate times?
There is an experimental API (WebRTC Insertable Streams), and an IETF Draft (Secure Frames), that together may allow true end-to-end encryption in the near future in a standards-compliant fashion. Note, however, that this still leaves key generation and exchange up to the application.
Nice overview of WebRTC Insertable Streams and Secure Frames: https://webrtcbydralex.com/index.php/2020/03/30/secure-frame...