I have 300/100mbs connection which costs 20 euros in month.
You can't do weird texture mapping or lossy compression and expect people to really seem like they are there. Even if you don't notice that stuff normally watching a video, I think you'll notice it when you're interacting like someone is really in front of you, and that will throw off the immersion.
1: https://www.quora.com/In-regards-to-filesize-how-big-is-1-mi...
That said, my intuition is they aren't doing a pure video encoding solution. The fact that they talk about 3d modeling leads me to believe they are doing a combination of model + texture to get the realistic results. That would significantly decrease the amount of bandwidth and computational power needed. Over a low bandwidth situation you'd simply need to send model updates and do some smart interpolation to determine what things should look like.
Similar to the concept that playing a 3d game requires MB of resources but recording the same game at 8k would require a boatload more memory.
My assumption is they are using LIDAR to get a good model, high quality cameras to texture things, and a nice AI to stitch things together and interpolate when data isn't arriving fast enough.
https://blogs.nvidia.com/blog/2020/10/05/gan-video-conferenc...
https://ai.googleblog.com/2021/02/lyra-new-very-low-bitrate-...
https://aibots.my/officialBlog/deepminds-ai-agent-muzero-cou...