No one will ever derive any utility from running models at this speed. Please prove me wrong. Give me the number of tokens input and output (and dont forget about reasoning) and acceptable time to wait for it and the use case.
This is like 98%.
If you already committed to your hardware, huge models at single or sub-digit TPS are still useful. LLM-as-judge is a good use case.
If your machine is gonna be idle overnight (and the power efficiency is excellent here), why pay openrouter if it’s not an interactive workload?