HNHacker News
TopNewBestAskShowJobs

lnenad

928 karma · joined November 2, 2018

Eng Mng with 10+ years of exp; currently working with AI/LLMs

nenadspp gmail.com

grafly.io

logdot.io

submissionscomments
lnenad··on One month coding with GLM 5.3 Flash
They've got an image of their homegrown benchmark in the post that lists DS4.1.
lnenad··on Micron CEO Says Memory Supply Will Be Much Tighter in 2027 and 2028 Than in 2026
Thoughts/opinions aren't source.
lnenad··on Micron CEO Says Memory Supply Will Be Much Tighter in 2027 and 2028 Than in 2026
> The catchup models are all basically distillations of the sota models

Source/citation?

lnenad··on Replacing Pull Requests with Delta
Yes, but compared to seeing 3k lines of code in a PR with no history with vibes in the description, what is better? This definitely "feels" like something worth exploring.
lnenad··on Pion, an agent designed to run any company autonomously
If you use a bajillion tokens your economical approach is to self host.
lnenad··on Homebrew 7.0.0
What does saying "working tech" do for you? If I have a working Samsung CRT from 25 years ago do I ping them about smart TV support? Nowadays it's a shitty situation with planned obsolescence; but 6 years for an open source project dedicating resources to a dead end is more than enough and appreciated.
lnenad··on The Gemini app is now available for Windows
I'm well aware but it just reaffirms that it's a cluster fuck from a product perspective and makes no sense to create such a weird segmentation. You didn't even mention Jules lol, who knows where it fits in.
lnenad··on The Gemini app is now available for Windows
What is happening at Google? It seems like they've got multiple teams building the same thing and competing for love from the higherups. Depending on who's in the lead the chosen package gets into the spotlight. It's been Jules, then Gemini, then Agy, now Gemini is back???? What are they smoking over there and what kind of a customer base do they hope to get with this.
lnenad··on Breaking Claude Code Opus 5 Auto Mode
The agent wrote the code that has a mechanism that triggers a file as a side effect. That file started the separate process, as it could have started any other binary.
lnenad··on Breaking Claude Code Opus 5 Auto Mode
But you're not actually hijacking the agent if you start a new process.
lnenad··on Breaking Claude Code Opus 5 Auto Mode
Yeah, I agree, this is a different vector. Still scary though and very related to AI.
lnenad··on P99 0 ms* autocomplete for 240M domain names
But there is a cool blog post about it though.
lnenad··on GLM-5.3 is now open-weight
48c 7643. I'm getting about 10tps @Q3kxl with 2x3090s.
lnenad··on GLM-5.3 is now open-weight
I'm getting about 10tps @Q3kxl with 2x3090s.
lnenad··on GLM-5.3 is now open-weight
What model are you interested in? DS Flash 0731@Q4KXL I'm about 25-30tps. Same as the new Qwen3.8 Flash Next. The new GLM 5.3Q3KXL at 10tps. I've got 2x3090s which I didn't mention in the original message.
lnenad··on GLM-5.3 is now open-weight
About 5k with RAM and GPUs bought used. Eastern Europe.
lnenad··on GLM-5.3 is now open-weight
I am getting 10t/s on unsloth's Q3kxl with 2x3090s@250w. It's enough for me for now. I will probably upgrade the GPUs down the line. DDR5 would have made the price of the machine double and I just wasn't prepared to pay that much.

Temp wise, no throttling, surprisingly cool.

lnenad··on GLM-5.3 is now open-weight
I have just built an Epyc with 512gb DDR4 3200 RAM for a "reasonable" price and I'm hoping to have a setup with GLM as the architect and Qwen 27b/Next Flash as the implementer. This is 1/5 of the price of the Mac, but also probably 1/5 of the speed lol.
lnenad··on Humanity has the debate about AI consciousness backwards
But there isn't? What does it factually mean to be conscious? How can we claim other living beings aren't are conscious?
lnenad··on Qwen3.8-Flash-Next
I've got a 48c Epyc with 2x3090s and 512gb ddr4 3200. It's good enough for 25+ tps with deepseek so I'm hoping for similar performance with less overthinking.
lnenad··on Qwen3.8-Flash-Next
Yeah I understand, it's my assumption that the actually/wait/but have a point. It doesn't reduce the fact that it increases the time for tasks substantially.
lnenad··on Qwen3.8-Flash-Next: A New Architecture, Towards Ultimate Cost-Efficiency
Especially on practical tasks. One shot prompts work better at Q6_K_XL for me. It loads a file, then analyses then second guesses itself then again then again then it tries to come up with a solution then second guess rinse and repeat. 122b is the perfect balance but it lacks quality for harder to solve stuff. I've ran DS Flash 0731 at Q4KXL, 3.8 Q6KXL, GLM 5.2 Q4KXL and they all over-reason. At least that's how it looks like to me when comparing with frontier models, even weaker ones.
lnenad··on Qwen3.8-Flash-Next: A New Architecture, Towards Ultimate Cost-Efficiency
Yeah 122B is the sweet spot for me as well. Even deepseek flash overthinks on stuff way too much. I think they fully rely on large reasoning turns to achieve better quality. The result of course means we wait a long time to get results even with high throughput as a lot of tokens are wasted.
lnenad··on Qwen3.8-Flash-Next
Adding to my homelab stack, hopefully it doesn't overthink like the little model. Actually, hoping it thinks a bit less. Wait actually I'm really praying it reasons a bit more directly. But wait, I'm really sure that it must be a bit better.
lnenad··on Quantum battery upends the rules of charging
I think as with most human undertakings, building isn't too much of a problem. Maintaining is. Even with what is still a relatively tame number of chargers you get a large number of them that are broken.
lnenad··on Hister – A private, full content search index that you control
I LOVE the concept. I will play around with the execution, if it works as described this is a great product.
lnenad··on Seed: Minimal, self-modifying agent harness
> If your needs are met otherwise stick to that and move on.

Weird to post such a thought in a forum where OP has posted their project for people to look at. I never mentioned any needs, I am saying the readme holds very little value for a project in a space that has 50 different projects/coding harnesses available.

How would I know if this fits my needs based on it saying LLMs can build tools for themselves? That's something they can do with any and every open source harness.

> But I am a hardware engineer biased by experience with minimalism; I got a single data structure (electromagnetism) to manipulate and BOMs that come with hard constraints.

Alrighty.

> On the contrary, web SaaS devs want to run a business/rocket to the moon moreso than be an engineer. As such git pulling, pip installing the world allows them to focus on their get rich quick with as little labor involved as possible goal front and center.

Posting your open source project to a forum such as hn means people will have opinions they want to voice. No need to shit on opinions you disagree with.

lnenad··on Seed: Minimal, self-modifying agent harness
I think as many things that are posted here lately there is no *why* attached to the readme. Why would one use this, what is the benefit of this approach? Am I really gonna need my model to build exotic tools around it; or is exec/web_search/web_fetch enough for 90% of the use cases? Is my agent not capable of writing new plugins/tools for pi/opencode?
lnenad··on Ask HN: Alternatives to GitHub
You've built a great piece of software, thank you!
lnenad··on Ask HN: Alternatives to GitHub
I'm running gitea successfully with very little resources.
Page 1 of 14Next →