0.3tx per second is decent speed?
And it gets worse with every token.
And it gets worse with every token.
192GiB Gorgon Halo systems will be an interesting future target for this model, the best you can do with 128GiB or less is probably to push batching higher in order to amortize the weights traffic over multiple inferences - which of course will sink single-session speeds even lower for a modest gain in total throughput.
For AI agents this would take a year
This is like 98%.
If you already committed to your hardware, huge models at single or sub-digit TPS are still useful. LLM-as-judge is a good use case.
If your machine is gonna be idle overnight (and the power efficiency is excellent here), why pay openrouter if it’s not an interactive workload?