This is interesting. While I think this approach may not be practical for a single computer due to poor performance, it could open new possibilities for distributed computing. Here, users across the internet could pool their graphics cards and the 80 layers could be distributed across 80 GPUs, with each GPU handling one layer. The results would then be processed through the system. Admittedly, if there is an average latency of 50 ms for each of the 80 GPUs, it would result in an additional 4 seconds of network delay. However, the overall process could be up to 80 times faster than on a single GPU since the layers are already in the GPU memory, albeit on different cards. This method also has the potential to enable the execution of much larger models on consumer hardware.