Hardware requirements?
"What you need" only includes software requirements.
"What you need" only includes software requirements.
So about 700 bucks for a 3090 on eBay
With a 3090 I guess you'd have to reduce context or go for a slightly more aggressive quantization level.
Summarizing llama-arch.cpp which is roughly 40k tokens I get ~50 tok/sec generation speed and ~14 seconds to first token.
For short prompts I get more like ~90 tok/sec and <1 sec to first token.