No way you can get that much each without extremely quants
For a single r9700 you have 637 GB/s and for qwen 3.8 27b q4_k_xl the maximum tg/s is 33 before mtp
Now if you meant 4xr9700 tensor parallelism with mtp, 80 tg/s starts to make sense
For a single r9700 you have 637 GB/s and for qwen 3.8 27b q4_k_xl the maximum tg/s is 33 before mtp
Now if you meant 4xr9700 tensor parallelism with mtp, 80 tg/s starts to make sense