236 karma · joined August 13, 2024
Don't forget to vote. November isn't far away.
Should you think it is wise to trust the machine that can't differentiate subject matters in a chat styled context, you have fun with that fluster cluck when it blows up.
Like Fable is highly useful, but it's really bad at keeping it's responses straight.
In fact, that "it's not X it is Y" pattern always crops up when it reasoned about the idea of X and I never fed it that. It's literally doing that because it can't predict that I'm a different entity despite it being able to say I am a different entity.
Edit: Clarification by removal of incomplete sentence fragment. Edit2: Clarification on the "proven in whole" thing.
To be completely fair though, I'm against the behaviors in information transfer I've been seeing. I've pass along AI generated runbooks, but they look nothing like the default outputs of these models. It's because I took time to apply all the writing knowledge I like to see in my curation. If people are doing this, I can't even tell it's AI writing. My work is done in minutes instead of deciphering so BS pseudo language they developed in their AI workspace (people really need to turn off those memory features).
---
Edit: Also if I'm the guy receiving a security report and it's AI generated and poorly formatted, I'm failing you short of producing something for a human to parse. Simple as that.
Yeah... no. The people doing this lack the communications training, see the output has the necessary information, and regurgitate it with no effort or care. This needs to stop.
We spent decades format building to make it easy quick and easy to get through something like a runbook. If your commands are bulleted instead of numbered and code blocked, it's wrong. If you didn't crawl through the playbook, it's immoral to hand that to me, you're wasting my time with untested slop.
This is a hill I will die on or absolutely start slaughtering people on. I just refuse to deal with this crap.
And just say: LLMs only amplify the the knowledge you have, even the best models o use like Fable still suffer from promoting false narrative as a chart progresses. Often it’s stuff that can be ignored like the “not X but Y” crap it dumps because its reasoning had assumptions it invalidated. Sometimes it’s directly in your system architecture, because you never expressed preferences for solved foundational issues, you end up with generic http handler setups or whatever the hot web thing is today.
Takes knowledge and lots of it to really be on top of when these things go down a failure mode path.
Why the heck someone would risk having an open bridge to the system is beyond me. Like maybe get used to using remote hardware screens (KVMs? It’s been a while since I’ve done datacenter), we have solutions for this that was absolutely skipped.
Initial results boiled down as follows.
# lmstudio-community/qwen3.8-27b@q4_k_m decode falloff 104.3 tok/s @ 12,683 -> 55.8 tok/s @ 240,755 (53% retained) prefill falloff 3,274 tok/s -> 1,059 tok/s (32% retained)
# qwen3_8_27b_nvfp4.ninfer decode falloff 173.3 tok/s @ 11,867 -> 139.3 tok/s @ 225,710 (80% retained) prefill falloff 8,726 tok/s -> 2,816 tok/s (32% retained)
I should still have room for more performance on the table. I've not even touched the overclock settings on the GPU.
This is a really cool project, I'm going to have to get into what those 3 guys are doing... assuming it can be done with what I got.
Prior to two weeks ago, I was just using Pi and Ollama.
I have tried my hand at putting together a few harnesses and I finally landed on what I like. Been working on this small app to handle running llama-server for me from any device that has the llama-cpp stack setup: https://github.com/SamInTheShell/loom
Qwen 3.8 is the first model I've been using that hasn't been having issues doing edit calls. Here are my llama server settings and GUFF that I use: https://gist.github.com/SamInTheShell/0bf838e8dc5093583b688e...
The model I’m specifically targeting to use at high speeds is Qwen 3.8 27b @q4ks. This model actually proved to be good at coding (it sits somewhere between Sonnet 5 and Opus 5 capability). M1 got 10 tok/s, Ryzen Halo 20tok/s, and Radeon 7900 XTX 50tok/s (can only do 128k context window in Radeon card).
The prefill gets extremely slow around 50k tokens in context window (whatever prompt processing stage entails could be wrong about phases here). It takes about 2 hours to fill the context.
Even with a drafter model intended for speed instead of mtp, I can’t get past 70tok/s, still is extremely slow to process prompts as context grows, and drops down to 40-50tok/s anyway making this config still moot for improvement on my Radeon card.
The only thing I can point to slowing me down is bandwidth of the card itself.
I am waiting to actually get my 5090 right now and I am betting that the 1700 Gbps of capacity will fix my prompt processing speeds. I don’t need full PCIe lane bandwidth to serve my house I just need to load the full model into vRAM and let the GPU do its thing.
Additional benefit to the external enclosure route is being able to migrate the inference between devices more easily. I can develop out the infrastructure then migrate the card to be hooked up to a shared node in the house with all the tools necessary for my family to take advantage of the privacy enhancement that comes with local inference.
I’m opting out of any corporate targeted code for personal use, anything not designed for user experience gets scrapped when I’m the user and I have Software on Demand as a Service.
My big contention with this project is that if I wanted to use node (or npm; or vite) in any capacity, I would have chosen node for that. I don’t want to use those tools in this context because it serves me worse than a well thought out solution.
The thought out solution is to isolate your frontend into its own directory. You can embed and serve that as part of your Go server. That frontend could have been done with react, vite, jsx, whatever typescript nonsense you want; without the Go parts necessarily needing npm or node in the build process.
I literally have toy projects illustrating the idea (this one is a toy I had AI make with some of my own patterns months ago): https://github.com/SamInTheShell/social/tree/main/frontend
Projects like grafana do stuff like this already. It’s a known good pattern.
Edit: If you keep checking back, I'll have a video showing off that application. It's a bit beefy to just get up and running.