HNHacker News
TopNewBestAskShowJobs

LuxBennu

86 karma · joined September 24, 2019

Software engineer. Working on LLM inference, evaluation pipelines, and open-source dev tools.
submissionscomments
LuxBennu··on ChatGPT for Excel
Chatgpt for Excel is still an office add-in running in the same sandbox though. strongpigeon described the exact bottleneck upthread, process boundary crossings, context.sync() roundtrips that take seconds on web. That's a platform limitation, not a model limitation. Swapping AI behind the add-in doesn't fix the fundamental constraint that third-party add-ins can't deeply integrate with Excel's runtime the way a native feature can. If copilot is bad despite having more access to excel internals(I don't like how Copilot is designed or implemented tho), an add-in with less access is likely not be better.
LuxBennu··on Show HN: Gemma 4 Multimodal Fine-Tuner for Apple Silicon
Yeah sorry that was unclear on my part. I chunk at the endpoint level, whisper itself obviously processes 30s windows. The memory/latency thing I was referring to is more about processing longer files end to end through the pipeline, not a single whisper pass. My fastapi wrapper just splits the audio and runs chunks sequentially so total wall time scales linearly with file length, nothing fancy.
LuxBennu··on Show HN: Ghost Pepper – Local hold-to-talk speech-to-text for macOS
Oh nice, the pyannote coreml port is interesting. Last time I looked at pyannote it was pytorch only so getting it to run efficiently on apple silicon was kind of a pain. Does the coreml version handle diarization or just activity detection?
LuxBennu··on Show HN: Gemma 4 Multimodal Fine-Tuner for Apple Silicon
Ah that makes sense, quadratic scaling is brutal. So with 96gb i'd probably get somewhere around 4-5k total sequence length before hitting the wall, which is still pretty limiting for anything multimodal. Do you do any gradient checkpointing or is that not worth the speed tradeoff at these sizes?
LuxBennu··on Show HN: Gemma 4 Multimodal Fine-Tuner for Apple Silicon
I run whisper large-v3 on an m2 max 96gb and even with just inference the memory gets tight on longer audio, can only imagine what fine-tuning looks like. Does the 64gb vs 96gb make a meaningful difference for gemma 4 fine-tuning or does it just push the oom wall back a bit? Been wanting to try local fine-tuning on apple silicon but the tooling gap has kept me on inference only so far.
LuxBennu··on Show HN: Ghost Pepper – Local hold-to-talk speech-to-text for macOS
Yeah that makes sense, chunking on silence would sidestep the latency issue pretty cleanly. I've been running it through a basic fastapi wrapper so it just takes whatever audio blob gets thrown at it, no chunking logic on the server side. Might be worth adding a vad pass before sending to whisper though, would cut down on processing dead air too.
LuxBennu··on Show HN: Ghost Pepper – 100% local hold-to-talk speech-to-text for macOS
I've been running whisper large-v3 on an m2 max through a self-hosted endpoint and honestly the accuracy is good enough that i stopped bothering with cleanup models. The bigger annoyance for me was latency on longer chunks, like anything over 30 seconds starts feeling sluggish even with metal acceleration. Haven't tried whisperkit specifically but curious how it handles longer audio compared to the full model.
LuxBennu··on Ollama is now powered by MLX on Apple Silicon in preview
that tracks with what i've noticed practically. shorter prompts feel basically the same between llama.cpp metal and what i'd expect from native mlx, but once context gets longer the overhead starts showing up. would be interesting to see if ollama's mlx path actually handles kv cache differently under the hood or if it just skips the buffer sync layer
LuxBennu··on Ollama is now powered by MLX on Apple Silicon in preview
Roughly 8-12 token/s on generation depending on context length. Prompt processing is faster obviously. Haven't benchmarked it super carefully though, just eyeballing the llama.cpp output.
LuxBennu··on From 300KB to 69KB per Token: How LLM Architectures Solve the KV Cache Problem
yeah fair point, it's definitely model dependent. i've had good results with qwen but tried it on a smaller mistral variant once and the output quality dropped noticeably even at q8 for both. the speed hit from mixed types hasn't been bad on apple silicon in my experience but i can see it mattering more on cuda.
LuxBennu··on From 300KB to 69KB per Token: How LLM Architectures Solve the KV Cache Problem
good overview of the architecture side but worth mentioning there's another axis that stacks on top of all of this: you can quantize the kv cache itself at inference time. in llama.cpp you can run q8 for keys and q4 for values and it cuts cache memory roughly in half again on top of whatever gqa or mla already saves you. i run qwen 70b 4-bit on m2 max 96gb and the kv quant is what actually made longer contexts fit without running out of unified memory. keys need more precision because they drive attention scores but values are way more tolerant of lossy compression, so the asymmetry works out.
LuxBennu··on Show HN: Reprompt – Analyze what you type into AI tools, not what they output
Thanks! Turns out structural signals get you surprisingly far. An LLM catches more, but speed is the feature.
LuxBennu··on Show HN: Reprompt – Analyze what you type into AI tools, not what they output
I ran this on my own prompt history and three things surprised me. found 3 API keys buried in copy-pasted stack traces (`reprompt privacy`). 35% of my agent sessions had error loops -- the agent retrying the same failing approach 3+ times (`reprompt agent`). And 50-70% of my conversation turns were filler like "ok try that" (`reprompt distill`).

    pip install reprompt-cli
    reprompt scan && reprompt
Everything runs locally -- zero network calls, zero telemetry. Also works as an MCP server and GitHub Action.
LuxBennu··on Ollama is now powered by MLX on Apple Silicon in preview
Already running qwen 70b 4-bit on m2 max 96gb through llama.cpp and it's pretty solid for day to day stuff. The mlx switch is interesting because ollama was basically shelling out to llama.cpp on mac before, so native mlx should mean better memory handling on apple silicon. Curious to see how it compares on the bigger models vs the gguf path
LuxBennu··on New Apple Silicon M4 and M5 HiDPI Limitation on 4K External Displays
Sadly I have the issue on a new m5 air. I have a 60hz 4k work monitor and two high refresh 4k gaming displays. The 60hz pairs fine with either gaming monitor, but the two gaming ones together and one just doesn't get recognized. Spent way too long trying new cables before realizing it's a bandwidth limitation.
LuxBennu··on Claude Code runs Git reset –hard origin/main against project repo every 10 mins
This is true for prohibitions but claude.md works really well as positive documentation. I run custom mcp servers and documenting what each tool does and when to use it made claude pick the right ones way more reliably. Totally different outcome than a list of NEVER DO THIS rules though, for that you definitely need hooks or sandboxing.
LuxBennu··on AI overly affirms users asking for personal advice
yeah that's a good way to put it. the "felt good in the moment" framing is basically the whole problem. the reward model was trained on human preferences and humans preferred the agreeable answer, so now that's what you get at inference time regardless of whether it's correct. the frustrating part is you can see it happen in real time if you log the outputs turn by turn, the model will literally contradict its own previous response just because the user sounded more confident.
LuxBennu··on AI overly affirms users asking for personal advice
i tested this pretty extensively actually. built a pipeline that asks the same question rephrased across multiple turns and tracks how much the model shifts based on user tone. even when you tell it to be critical, the moment the user pushes back with any confidence the model just folds. it's not a prompting problem, it's baked into RLHF. you're right that LLMs will poke holes in stuff when the conversation starts neutral, but add any emotional charge and the sycophancy takes over immediately. that's exactly why the personal advice angle matters, that's peak emotional signal from the user.
LuxBennu··on Anatomy of the .claude/ folder
this is exactly how i use it too. i have a few custom MCP servers running on a mac mini homelab, one for permission management, one for infra gateway stuff. the key thing i learned is keeping CLAUDE.md updated with what each MCP server actually does and what inputs it expects. otherwise claude code will either not use the tool when it should, or call it with wrong params and waste a bunch of back and forth. once you document it properly it really does feel like having a team member who just knows how your stack works. the accounting use case is a great example because nobody else's generic tooling would ever cover that.
LuxBennu··on Show HN: Reprompt – Score your AI coding prompts with NLP papers
OpenClaw adapter was straightforward since it uses the same JSON session format.

For agent-generated prompts, I haven't specifically benchmarked agentic workflows yet. The repetition metric detects n-gram repetition within a single prompt, not across prompts. Agent scaffolding tends to inject the same system prefix into every call, which would get flagged if the agent concatenates it into the user message. reprompt currently treats each user turn as a separate prompt, so the system prefix isn't in scope otherwise.

Repetition is weighted 0-15 out of 100, so not dominant. But for heavily templated agent prompts it could actually be the most informative signal. If everything else is boilerplate, the repetition score would separate the prompts where the agent actually varied its approach. Could be an interesting lens for comparing agent frameworks too.

If you have OpenClaw agent sessions I'd be curious what the distribution looks like.

LuxBennu··on LLM Architecture Gallery
Interesting collection. The architecture differences show up in surprising ways when you actually look at prompt patterns across models. Longer context windows don't just let you write more, they change what kind of input structure works best.
LuxBennu··on 1M context is now generally available for Opus 4.6 and Sonnet 4.6
Your code map compresses signal on the context side. Same principle applies on the prompt side: prompts that front-load specifics (file, error, expected behavior) resolve in 1-2 turns. Vague ones spiral into 5-6. 1M context doesn't change that — it just gives you more room for the spiral.
LuxBennu··on BitNet: 100B Param 1-Bit model for local CPUs
Fair enough — I've been lurking since 2019 and picked a bad day to start commenting on everything at once. Not a bot, just overeager. I'll pace myself.
LuxBennu··on AutoKernel: Autoresearch for GPU Kernels
This is the right call. llama.cpp has dozens of hand-tuned CUDA kernels across Q4_K_M, Q5_K_S, Q8_0 and other quant formats, each targeting different hardware profiles. An autoresearch approach that could optimize these per-GPU would be huge — right now performance varies wildly between, say, an RTX 3090 and a 5070 Ti on the same quant format because the kernels are tuned for specific architectures. The hardware diversity in the llama.cpp user base is exactly where automated kernel search has the most to gain.
LuxBennu··on Launch HN: RunAnywhere (YC W26) – Faster AI Inference on Apple Silicon
Agreed. The real value proposition of Apple Silicon for local inference is running models that won't fit on consumer GPUs. I run Qwen 70B 4-bit on an M2 Max 96GB through llama.cpp and it's usable — not fast, but the unified memory means it actually loads. Would be interested to see MetalRT benchmarks at that scale, since the architectural advantages (fused kernels, reduced dispatch overhead) should matter more as models get memory-bandwidth-bound.
LuxBennu··on Why AI Chatbots Agree with You Even When You're Wrong
Memory helps, but sycophancy exists even in single-turn interactions — the Anthropic 2023 paper showed pretrained models cave to mild pushback like "I think the answer is X but I'm not sure" with zero conversation history. In our LLM eval pipelines, we see the same thing: models accept false presuppositions embedded in a single prompt without any prior context to fall back on. The deeper issue is that RLHF rewards agreeableness because human raters genuinely prefer it. Better memory architecture would help with multi-turn drift, but the single-turn sycophancy is baked into the training signal itself.
LuxBennu··on I designed a bfloat16/FP8 alternative in a week using LLMs
The "Block-Scale-Free" property is the most compelling part here. Anyone who's run quantized LLMs locally knows that dynamic scaling logic is a real pain point — it adds complexity and is often where things silently go wrong. Trading that for QAT-first deployment seems like a reasonable bargain, especially for edge inference where you want the simplest possible hardware path. Curious whether AF8 has been tested against GGUF Q8_0 on any standard benchmarks.
LuxBennu··on Covenant-72B: Pre-Training a 72B LLM with Trustless Peers Over-the-Internet
Interesting that SparseLoCo held up at 72B scale with permissionless participants. I run distributed inference across multiple machines over Tailscale (M2 Max + RTX 5070 Ti), and even in that controlled setup, network variance is the dominant bottleneck. The fact that they got competitive quality with peers joining and leaving freely on 1.1T tokens is impressive — though I'd love to see how much the blockchain verification overhead actually cost in effective compute utilization.
LuxBennu··on AI-SLOP: Develop Best Current Practises for Open Source Maintainers
As someone who uses Claude Code daily and is about to publish an open source project, I see both sides. The key insight in this OSSF issue is "human-in-the-loop accountability" — most AI slop PRs fail not because the code is wrong, but because the submitter can't explain what it does. A simple "explain your change in your own words" requirement would filter for understanding and is hard to game with AI alone.
LuxBennu··on BitNet: Inference framework for 1-bit LLMs
The title is misleading — there's no trained 100B model, just an inference framework that claims to handle one. But the engineering is worth paying attention to. I run quantized 70B models locally (M2 Max 96GB, llama.cpp + LiteLLM), and memory bandwidth is always the bottleneck. The 1.58-bit approach is interesting because ternary weights turn matmuls into additions — a fundamentally different compute profile on commodity CPUs. If 5-7 tok/s on a single CPU for 100B-class models is reproducible, that's a real milestone for on-device inference. Framework is ready. Now we need someone to actually train the model.