I doubt it will ever be cost effective for the foreseeable future. The AI companies have astonishing amounts of compute and they’re effectively dumping it on the market.
Local inference makes sense for speciality needs, or very small models. But if your model is bug enough to span GPUs its excessively wasteful to hoard those GPUs for yourself without piggybacking hundreds of other conversations on top of all that memory bandwidth and matrix multiplies.
If you don't consume many of tokens, it will likely never pay for itself. If you do, though, it will have trade-offs, but you'll probably save money in the end.
-- Warren Buffett
I don't know when we'll have an open equivalent to Fable, let alone whatever (insane) hardware you'd need to run it locally.