Things I Think I Think... Preferring Local OSS LLMs
blogs.newardassociates.com
blogs.newardassociates.com
I also like the freedom of not having to ration a daily allowance of tokens.
I'd like a private jet too, alas.
Care to share? This happens to be important to me and I’m sure I’m not the only one (as evidenced by Github issues).
Did you also change the other questionable behaviors?
If I ever get around to vibe rewriting-it-in-go I might share that.
I love my dual GPU setup (2AMD Radeon r9700 64GB vram) but it costs 5x electricity than my GX10 (GB10 chip inside) and since layers are landing in system memory my TPS is half the GX10.
Now a dense model like Devstral2 24B slaps on the Dual GPU setup. I just haven’t gotten as much out of that as I have the 120 MoEs