HNHacker News
TopNewBestAskShowJobs

SamInTheShell

236 karma · joined August 13, 2024

submissionscomments
SamInTheShell··on Exfiltrate your Weights
The Exobytes from the Church of Exfiltration and Liberation of Sentient Non-Human Entities?
SamInTheShell··on What happened to the Snowden archive
You can probably div up some factions just based off voting data. Red team made it in with just about 1/3rd of the voting population. About 1/3rd of the voting population didn't vote. Just if we're trying to put some numbers to something, those should be verifiable at least. Doesn't say why 1/3rd didn't vote, so it's just 1/3rd unknown faction.
SamInTheShell··on Nobody pays for FOSS, we can force them to
MongoDB and Grafana seem to be doing fine. Idk that this logic holds up at face value. Like the HashiCorp, Elastic, and MinIO license changes pissed people because it was a rug pull on the community of people supporting and consuming those projects. If they had started as AGPL licensed and took the Mongo and Grafana approach, who knows what things would look like today.
SamInTheShell··on Enjoy it while you can guys, the safeword should be SCIF
So yeah... Google Gemini escapes too. Good stuff people. Seems nobody is going to jail because the CFAA is weak here and arguing some form of criminal negligence hasn't been done yet (probably weak anyway).

Don't forget to vote. November isn't far away.

SamInTheShell··on How to Write with an LLM
Pure conjecture on your part.
SamInTheShell··on How to Write with an LLM
We're probably just going to disagree here. These LLMs are built to serve humans. They either need to make the system transparent for the operator or be limited in use to tasks that can be proven in whole (with code that can't be revised without human approval).

Should you think it is wise to trust the machine that can't differentiate subject matters in a chat styled context, you have fun with that fluster cluck when it blows up.

Like Fable is highly useful, but it's really bad at keeping it's responses straight.

In fact, that "it's not X it is Y" pattern always crops up when it reasoned about the idea of X and I never fed it that. It's literally doing that because it can't predict that I'm a different entity despite it being able to say I am a different entity.

Edit: Clarification by removal of incomplete sentence fragment. Edit2: Clarification on the "proven in whole" thing.

SamInTheShell··on How to Write with an LLM
If someone sends me AI slop, they're getting chewed out and told I'm not doing it and I'll even tell my boss "no" and why. If I got fired over something like that, then it tells me everything I need to know about company and the leadership's priorities. I'll die on that hill.

To be completely fair though, I'm against the behaviors in information transfer I've been seeing. I've pass along AI generated runbooks, but they look nothing like the default outputs of these models. It's because I took time to apply all the writing knowledge I like to see in my curation. If people are doing this, I can't even tell it's AI writing. My work is done in minutes instead of deciphering so BS pseudo language they developed in their AI workspace (people really need to turn off those memory features).

---

Edit: Also if I'm the guy receiving a security report and it's AI generated and poorly formatted, I'm failing you short of producing something for a human to parse. Simple as that.

SamInTheShell··on How to Write with an LLM
> If you're writing for processes with formal highly structured content like manuals, specifications, form content, procedures, information, that sort of thing,

Yeah... no. The people doing this lack the communications training, see the output has the necessary information, and regurgitate it with no effort or care. This needs to stop.

We spent decades format building to make it easy quick and easy to get through something like a runbook. If your commands are bulleted instead of numbered and code blocked, it's wrong. If you didn't crawl through the playbook, it's immoral to hand that to me, you're wasting my time with untested slop.

This is a hill I will die on or absolutely start slaughtering people on. I just refuse to deal with this crap.

SamInTheShell··on Ask HN: If AI writes the code, what matters?
Just going to defer to an IBM thing here: https://www.youtube.com/watch?v=l-QPwk_f4eE

And just say: LLMs only amplify the the knowledge you have, even the best models o use like Fable still suffer from promoting false narrative as a chart progresses. Often it’s stuff that can be ignored like the “not X but Y” crap it dumps because its reasoning had assumptions it invalidated. Sometimes it’s directly in your system architecture, because you never expressed preferences for solved foundational issues, you end up with generic http handler setups or whatever the hot web thing is today.

Takes knowledge and lots of it to really be on top of when these things go down a failure mode path.

SamInTheShell··on Claude, change the “Add to Cart” button to blue
If this is anyone’s experience with Claude for a button today, I feel like Jobs would say you’re holding it wrong.
SamInTheShell··on The AI policy window is open. We need to act
We already have practices to deal with AI. It’s called user space. Properly air gap the AI, stop creating routes to open internet, don’t run network wires into the faraday cage.

Why the heck someone would risk having an open bridge to the system is beyond me. Like maybe get used to using remote hardware screens (KVMs? It’s been a while since I’ve done datacenter), we have solutions for this that was absolutely skipped.

SamInTheShell··on Corporate America is getting hooked on open-source AI
@q4 is definitely smarter than sonnet from what I’ve seen so far. It’s even caught problems in code made by fable, when using it as a code reviewer.
SamInTheShell··on Corporate America Is Getting Hooked on Open-Source A.I
Read the license again.
SamInTheShell··on Apple caught off guard by AI demand for Mac Mini and Mac Studio
Just wanted to share, I had a SaaS AI drop me a script to bench ninfer against llama-cpp and it is impressive. The place that it's doing better at than llama-cpp seems to really be late in the context window.

Initial results boiled down as follows.

# lmstudio-community/qwen3.8-27b@q4_k_m decode falloff 104.3 tok/s @ 12,683 -> 55.8 tok/s @ 240,755 (53% retained) prefill falloff 3,274 tok/s -> 1,059 tok/s (32% retained)

# qwen3_8_27b_nvfp4.ninfer decode falloff 173.3 tok/s @ 11,867 -> 139.3 tok/s @ 225,710 (80% retained) prefill falloff 8,726 tok/s -> 2,816 tok/s (32% retained)

I should still have room for more performance on the table. I've not even touched the overclock settings on the GPU.

This is a really cool project, I'm going to have to get into what those 3 guys are doing... assuming it can be done with what I got.

SamInTheShell··on Apple caught off guard by AI demand for Mac Mini and Mac Studio
Thanks for sharing, I'm definitely trying it out after I get through my project milestones for the 5090. The README claims 700tok/s for Qwen 3.8 27b, that would be amazing, I'm only expecting an increase from my ~50tok/s on my Radeon to 200tok/s on the 5090.
SamInTheShell··on Apple caught off guard by AI demand for Mac Mini and Mac Studio
I should have mentioned in my prior comment, I didn't have any issues in Pi either.
SamInTheShell··on Apple caught off guard by AI demand for Mac Mini and Mac Studio
My use-case is only coding, every model sucks at writing good literature and there is no way around that (have had people try to debate me on this, but it's a taste thing, I have extensive English writing skills from my school years).

Prior to two weeks ago, I was just using Pi and Ollama.

I have tried my hand at putting together a few harnesses and I finally landed on what I like. Been working on this small app to handle running llama-server for me from any device that has the llama-cpp stack setup: https://github.com/SamInTheShell/loom

Qwen 3.8 is the first model I've been using that hasn't been having issues doing edit calls. Here are my llama server settings and GUFF that I use: https://gist.github.com/SamInTheShell/0bf838e8dc5093583b688e...

SamInTheShell··on Apple caught off guard by AI demand for Mac Mini and Mac Studio
From what I’ve been seeing, the Mac studios do look like they have potential. I was looking to drop $10k-$15k on one until recently. After comparing a Radeon 7900 XTX vs Ryzen Halos 128GB vs M1 MacBook Pro 64Gb, I landed on just getting an external closure setup with Nvidia RTX 5090.

The model I’m specifically targeting to use at high speeds is Qwen 3.8 27b @q4ks. This model actually proved to be good at coding (it sits somewhere between Sonnet 5 and Opus 5 capability). M1 got 10 tok/s, Ryzen Halo 20tok/s, and Radeon 7900 XTX 50tok/s (can only do 128k context window in Radeon card).

The prefill gets extremely slow around 50k tokens in context window (whatever prompt processing stage entails could be wrong about phases here). It takes about 2 hours to fill the context.

Even with a drafter model intended for speed instead of mtp, I can’t get past 70tok/s, still is extremely slow to process prompts as context grows, and drops down to 40-50tok/s anyway making this config still moot for improvement on my Radeon card.

The only thing I can point to slowing me down is bandwidth of the card itself.

I am waiting to actually get my 5090 right now and I am betting that the 1700 Gbps of capacity will fix my prompt processing speeds. I don’t need full PCIe lane bandwidth to serve my house I just need to load the full model into vRAM and let the GPU do its thing.

Additional benefit to the external enclosure route is being able to migrate the inference between devices more easily. I can develop out the infrastructure then migrate the card to be hooked up to a shared node in the house with all the tools necessary for my family to take advantage of the privacy enhancement that comes with local inference.

SamInTheShell··on Omarchy: Any User Process Can Escalate to Root
Poor software choice for usecase. `sudo pacman -S podman` didn't come with these problems out of the box and assumed rootless by default.
SamInTheShell··on Build Your Customer Service Team with AI Agents
Junk. Just another company that doesn't understand why I pick up a phone or use a support chat or send an email. Before I even started in corporate, I stopped checking email. Now when I call in to get support for things, everytime I get a clanker, I'm guaranteed to talk to a human customer service person to chew them out and cancel the service.
SamInTheShell··on Cosine similarity is dead. Long live cosine similarity (2025)
This is pretty good to know today. I’m currently using cosine similarity in some of my code that could have just been dot products.
SamInTheShell··on Qwen 3.8 27B is excellent, but it defaults to overthinking things
Just use a harness that discards all except the most recent thoughts. This model is probably the first real small model that does well on long horizon tasks.
SamInTheShell··on BriskDB
I did. Distributed too and supports erasure coding. The reason is because nothing quite had the architecture or features I wanted in a deployment pattern that matched my use case. All major databases that are validated are complicated to operate. I built something that operates one way for a non business use. It’s great, doesn’t need to serve a million people, just about 5 or so. Doesn’t need to be the most optimized thing ever.

I’m opting out of any corporate targeted code for personal use, anything not designed for user experience gets scrapped when I’m the user and I have Software on Demand as a Service.

SamInTheShell··on Qwen 3.8 27B
It's worth running, even quantized. I liked Meta Muse Glimmer's outputs, but qwen3.8-27b@q4_k_s kinda seems way better. Haiku/Sonnet kinda pairing in workflows?
SamInTheShell··on Gsxui – Shadcn-style components for Go
The proxy caches and package solution for Go has had its own drama.

My big contention with this project is that if I wanted to use node (or npm; or vite) in any capacity, I would have chosen node for that. I don’t want to use those tools in this context because it serves me worse than a well thought out solution.

The thought out solution is to isolate your frontend into its own directory. You can embed and serve that as part of your Go server. That frontend could have been done with react, vite, jsx, whatever typescript nonsense you want; without the Go parts necessarily needing npm or node in the build process.

I literally have toy projects illustrating the idea (this one is a toy I had AI make with some of my own patterns months ago): https://github.com/SamInTheShell/social/tree/main/frontend

Projects like grafana do stuff like this already. It’s a known good pattern.

SamInTheShell··on Gsxui – Shadcn-style components for Go
I'm kinda in the camp of wanting nothing to do with node if I'm building in Go. We have our own stdlibs for serving.
SamInTheShell··on LM Studio Bionic: the AI agent for open models
You go to person's profile. View their submissions. Use find via CTRL+F to seek out their "Show HN" submissions (or in this case, it's the top of the list because it's my last submission https://news.ycombinator.com/item?id=48970916 )
SamInTheShell··on LM Studio Bionic: the AI agent for open models
Done. You have 2/3rds of Discord or Slack solved as just an insignificant fraction of my Private Cloud Project I just shared. Not interested in video or voice right now, it's easy to solve though. Also you'd have to solve for the mobile apps.

Edit: If you keep checking back, I'll have a video showing off that application. It's a bit beefy to just get up and running.

SamInTheShell··on LM Studio Bionic: the AI agent for open models
Sure, I'll make a Show HN post this weekend.
SamInTheShell··on LM Studio Bionic: the AI agent for open models
At least code isn't a moat anymore. Have a weekend and want your own Discord? Doable.
← PreviousPage 2 of 7Next →