I run similar workflows as the author on my 5090. It's really good and reasonably fast at ~90 tok/sec on LM Studio. I haven't tried ninfer yet. The only problem is having to be mindful about the context size. I am jealous of the folks with RTX 6000.
Ninfer is the real deal. 160 tokens/sec. You can parallelize 4 chats at once and get ~400 tokens/sec I’ve heard. Although that just compounds the context size issue, still absolutely incredible for running locally. I made a web based battle chess game with sound effects running on a pi harness in about 15 minutes, complete with a (very unskilled) “AI” (not an llm) that plays against you if you want!
Just tried it, really cool and works as advertised.