I don't need 2.4T to do that; I'm doing it with 35B or 27B. If they get me a model in ~80B with a A5B or A7B, that will be the end point.
It's bizarre people, by themselves, believe all these parameters are getting them much more.
Lets be serious: if we as a civilization really wanted the advancements promised, we'd find the 1000 best scientists and give them free access to these models while the rest of us get personal GPUs for specific use cases.
But instead, we have to endeour this penis measuring contest for the infinite bikeshedding of the universe.
But its also a 800B sized model running on a ram constrained system with no GPU.
Techniques, GPUs, more ram, and faster disks can always speed it up. But the point being is they run on low end machines now. Its now an optimization problem, not a possibility assessment.