Are you basically asking why they don't take a performance-sensitive, specialised, and parallel task and run it on a low-performance, unspecialised, and sequential system?
Would take hours and be super inefficient.
4G is perfectly fine for uploading videos. It can hit up to 50 Mbps. LTE-Advanced can do 150 Mbps.
The CPU and bandwidth costs of transcoding to 40+ different audio and video formats would be massive though. I could imagine a 5 minute video taking more than 24 hours to transcode on a phone.
Uploading corrupt files could allow the uploader to execute code on future client machines. You must check every frame and the full encoding of the video.
So yes, for the bigger names like Google this is an unacceptable risk. They will generally avoid serving any user-generated complex format like video, images or audio to users directly. Everything is transcoded to reduce the likelihood that an exploit was included.
Given your average DSL uplink of 5 MBit/s, that's 2 hours uploading for the master version... and if I had to upload a dozen smaller versions myself, that could easily add five times the data and upload time.
I’m not an expert in this but I know that “optimally encoding a video” is an actual job. That’s because there’s no global definition of optimal (it varies depending on the source material and target devices, not to mention the costs of your compute, bandwidth, and time); you’re doing it multiple times using different codecs, resolutions, bandwidth targets, etc.; and those change regularly so you need to periodically reprocess without asking people to come back years later to upload the iPhone 13 optimized version.
This brings us to a second important concept: YouTube is a business which pays for bandwidth. Their definition of optimal is not the same as yours (every pixel of my masterpiece must be shown exactly as I see it!) and they have a keen interest in managing that over time even if you don’t care very much because an old video isn’t bringing you much (or any) revenue. They have the resources to heavily optimize that process but very few of their content creators do.
This is in no way a "simple skill" as maximum video bitrate is only one of a number of factors for encoding video. For streaming to end users there's questions of codecs, codec profiles, entropy coding options, GOP sizes, frame rates, and frame sizes. This also applies for your audio but replacing frame rates and sizes with sample rate and number of channels.
Streaming to ten random devices will require different combinations of any or all of those settings. There's no one single optimum setting. YouTube encodes dozens of combinations of audio and video streams from a single source file.
Video it turns out is pretty complicated.