It would certainly not make any economic sense, and I guess that's also why noone is seriously looking into stuff like volunteer/enthusiast clusters of home computers to do inference in the same way that e.g. LHC@home works. The main bottleneck for LLMs is still memory bandwidth. Any memory bus not directly soldered on your GPU is terribly slow. That's why one big GPU with twice the VRAM will always perform significantly better than two GPUs with half the VRAM each. And it's also not like you can just solder more memory onto a chip. At modern speeds, the speed of light is a hard limit. For current GDDR7, signals may only travel like 10mm per cycle.
If you spread such a system out over dozens or hundreds of tiny chips, you'll be wasting most of its resources and lose hard to anyone who built a single chip setup.