That's not a bad result, although for £320 for 4x Pi5s you could probably find a used 12GB 3080 and probably more than 10x token speed
https://github.com/geerlingguy/ollama-benchmark?tab=readme-o...
Then I saw you github link and your HN handle and I was like “Wait, it is Jeff Geerling!”. :D
Double thanks for the 3rd party mac mini SSD tip - eagerly awaiting delivery!
That wouldn't get on Hacker News ;-)
My 300w 3070ti doesn't really exceed 100w during inference workloads. Boot up a 1440p video game and it's a different story altogether, but for inference and transcoding those 3060s are some of the most power efficient options on the consumer market.