The thing is, however, that at 2k one is not paying good money, one is paying near the least amount possible. TFA specifically is about building a machine on a budget, and as such cuts corners to save costs, e.g. by buying older cards.
Just because 2k is not a negligible amount in itself, that doesn't also automatically make it adequate for the purpose. Look for example at the 15k, 25k, and 40k price range tinyboxes:
It's like buying a 2k-worth used car, and expecting it to perform as well as a 40k one.
This is the problem.
If your use case is getting a small handful of non-urgent responses per day then it's not a problem. That's not how most people use LLMs, though.
If there are things you cannot send to a random party, you might want to look at hosted versions with agreements (if it's a code issue, if you're fine with github then azure is probably fine too).
Outside of that, if you really need to then sure, but these are the kinds of things that really benefit from being able to get high usage on GPUs for short periods of time.
In the future, I expect this to not be the case, because models will be far more efficient. At this pace, maybe even 6 months can make a difference.
Let me know if you find some config that really leverages more cores!