I mean, it's a neat trick to download an ML model into my web browser's tab's RAM, and run inference against is using WebGPU, but this move is saying it's only that, a novelty that's limited to models smaller than 4 GiB. Because outside of that, the web browser is the frontend to the system, and while the backend server (which could be running locally) can take as much RAM as there is, they are, in fact, saying that 4 GiB should be enough for rendering what a person can see and interact with. Which is different than a limit imposed on all computing, assuming the client server model continues to be dominant. The limitation being imposed means that heavier lifting has to be done on a server somewhere instead of locally, so the "trap" is that computer manufacturers have maxed out computers with 16 GiB of RAM in 2024 (or even 8, hi Apple). Except that 16 has been the max on many systems since like, 2016.
So it's a measure to try and prevent laptops from having to grow memory requirements. It's like limiting pairs of shoes to only having two shoes. You can have many pairs of shoes, but you don't need three shoes at once.