I would guess a big part is switching/routing specifically geared towards low-latency audio/video.
Just as an example, they probably use UDP for latency, which has no ordering as part of the protocol. However, for streaming stuff, you really want to drop out of order packets and prioritize whatever the latest packet is.
Zoom probably has some kind of custom routing that handles that. If they get congested or clients are unstable, it prioritizes serving the latest audio/video rather than trying to push things out in the order they came in, only for the Zoom client to discard the out of order packets anyways. Amazon can't really do that, because they don't know which bits of data are the ordering.
On the pricing side, people are probably the component that scales the least. You're probably going to need a dedicated person to monitor/fix/upgrade it, and deal with calls from people who have issues. Between salary and benefits, that's at least a 6 figure expense. Zoom's most expensive tier is $240/license/year, so $100,000 would buy you 416 licenses. And $100,000 is a lowball here; that's roundabouts a $60k salary, and the other $40k is paid out in benefits/space/etc. This also doesn't count the bandwidth charges, nor the hardware to run the system on.
Imo, it's also just generally a bad idea to DIY anything business-critical that isn't a core competency of the company. It tends to cause issues. What do you do when the person that built this bespoke Zoom alternative quits? Can you risk it breaking and not working until you can hire someone new and they figure out how it works? You could hire 2 of them, but Zoom is definitely cheaper then. What do you do if the open source project gets abandoned? It also distracts from things that actually make you money. It makes more sense to me to focus on your core line of business, and use the additional proceeds to just pay for Zoom.