HNHacker News
TopNewBestAskShowJobs

greggh

344 karma · joined January 3, 2015

submissionscomments
greggh··on Best LLM for every budget, updated daily
I've been running a quant/tune of Qwen3.8 27B on my M1 Max 32gb MacBook. That plus a good pi setup is having great results. I've used a full q8 of the model before and I dont see a real difference other than how slow it is. But leaving it running overnight on tasks is working great. It is currently debugging some issues in a native Mac Swift application and getting through the list of issues just fine.

This is the one that works good for me on 32gb:

https://huggingface.co/ISTA-DASLab/Qwen3.8-27B-GSQ-RCO-GGUF

Specifically this one: Qwen3.8-27B-GSQ-RCO-IQ3_S-mtp.gguf

greggh··on I don't recommend Tailwind CSS
Many of us oldies who started working on the web years before IE 6 actually like Tailwind. We all have our preferences. But I know multiple 50+ year olds who have been doing this work since CSS was barely a proposal who now use Tailwind.
greggh··on Advancing the price-performance frontier with GPT‑5.6
Use a harness like OMP that lets you choose which model does which things. My main model is GLM 5.2, it handles planning and anything I dont have covered by other models. Tasks from todos and in sub agents are done by deepseek, I have different models for the git work like add/commit/push (that goes through cheap Minimax M3), and so on...

This way the expensive/strong model only handles the architecture and orchestration tasks. The cheaper models handle everything else and the strong one knows how to tell them what to do in enough detail to get good work out of them.

greggh··on Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac
Yeah, its the 256gb version.
greggh··on Show HN: Open-source engine running Gemma 4 26B in 2 GB RAM on any M-series Mac
It does exactly what it says it does. On my Mac mini M4 with 16GB of ram it is running at just over 5 tok/s. That jump from M4 to M5 is crazy.
greggh··on Show HN: Orate – On-device neural text-to-speech queue for Mac
It took the key! I tried emailing the support email in the website footer and it bounced back. Thought it might be easier to work through things.

It still crashes when I try to view the shortcuts in settings. But everything else seems better.

greggh··on Show HN: Orate – On-device neural text-to-speech queue for Mac
It doesn't crash on me anymore! But it won't take my license key, and I am just hitting the copy button on polar.sh to copy the key and paste it in.
greggh··on Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
Stock OMP. When you add a provider and then /model you can choose a model and choose what roles it will take. Do that for each role it has.

Then I use /plan when I need to be planning and not writing to any files. Take the advice the UI gives you where it tells you to add the word orchestrate into your prompts when you want to make sure it uses its todo/tasks and sub-agents.

greggh··on Show HN: Orate – On-device neural text-to-speech queue for Mac
I had to click Skip to get past the intro screens, then the app comes up. I went to settings and pasted in my license, clicking activate says it's not valid. It is the key on my polar.sh account and I copy/pasted it so I know it's correct. Then it crashed.
greggh··on Kimi K3 Is Competitive with Fable; Kimi K3 and Fable Is SoTA
I do this in OMP, a fork of Pi. It lets you set different models for different tasks. So with an API that has many different companies models I can set the Plan model to the best one, right now I am using GLM 5.2 for that, it plans really well. I have Vision set to Kimi 2.7 Code (cheaper and vision is just fine). Minimax M3 is set to the Advisor role (double checks work). Deepseek v4 Flash is set for the Task role. And MiMo 2.5 pro is set as default.

With this setup GLM handles planning and managing my AGENTS.md, and orchestrating subagents for tasks from the plan/todo GLM created. The tasks themselves are handed off to Deepseek v4 flash to implement with strong instructions and examples for each agent. Minimax M3 reviews the output as the Advisor and recommends changes, catches bugs, and whatnot, subagents can be re-run with that information.

Overall I am saving a lot using some of these smaller models. But with this setup I am getting great results.

greggh··on Show HN: Orate – On-device neural text-to-speech queue for Mac
Mac mini m4 16gb, latest updates.
greggh··on Show HN: Orate – On-device neural text-to-speech queue for Mac
It would be nice to have a sample of the audio quality and voice(s) it uses on the website so we know what it will sound like.

-- Edit:

I installed it, ran it and got the first time wizard, granted permissions and hit continue and it just exits without any other message or screen. It isn't in the menu bar. I run it again and get the same. It says the permissions are granted already the second time though.

I like the interface on the website a lot, so I would love to use this. Thoughts?

greggh··on Where are YC founders now? OpenAI and Anthropic, mostly
The wheels on the bus go round and round, round and round, round and round...
greggh··on Grok Build is open source
Just a prototype? I have no reason to leave the terminal for a GUI IDE. TUI works great, does what I need and is very easy to use and interact with.
greggh··on 1024000^2 Blocks, 2B2T Minecraft Server World Download Project, and Discoveries
2b2t Place is exactly that.

https://2b2t.place

greggh··on This Month in Ladybird – April 2026
Uhh, yes? Non-profits take donations to keep doing their work.
greggh··on TIL: Apple Broke Time Machine Again on Tahoe
On Tahoe my Time Machine was broken after the update. My backup target is on a QNAP NAS. I just had to set it up from scratch again and it worked. But it did cost me a few files I was trying to recover. So I feel this.
greggh··on Ross Stevens Donates $100M to Pay Every US Olympian and Paralympian $200k
The real answer here is that he is mad about people protesting what Israel is doing in Gaza. This $100M donation is being made with funds he had given to UPenn. He has taken it back, via lawyers, because they allowed the protests to go on. He is now just taking that original donation and moving it somewhere else. Not that I am against the Olympians getting paid, just some context.

Sources: https://philanthropynewsdigest.org/news/donor-pulls-100-mill... https://thehill.com/homenews/education/4348656-upenn-loses-1... https://www.timesnownews.com/world/who-is-ross-stevens-stone... (many more)

greggh··on Trinity large: An open 400B sparse MoE model
The only thing I question is the use of Maverick in their comparison charts. That's like comparing a pile of rocks to an LLM.
greggh··on Claude Cowork runs Linux VM via Apple virtualization framework
Use a devcontainer. Claude Code's repo has one built specifically for it:

https://github.com/anthropics/claude-code/tree/main/.devcont...

greggh··on Stop Doom Scrolling, Start Doom Coding: Build via the terminal from your phone
Following that story as it happened, it was all on the phone with the phone keyboard and he somehow made multiple good Neovim plugins including that very popular one (which I use in multiple configs).
greggh··on We pwned X, Vercel, Cursor, and Discord through a supply-chain attack
Right, but Eva found an RCE and only got $5,000.
greggh··on Show HN: Sim – Apache-2.0 n8n alternative
Thanks, and that sounds great. On the backend what are you using for the DAG stuff to make it durable? Temporal?
greggh··on Show HN: Sim – Apache-2.0 n8n alternative
Development seems pretty rapid, how often are breaking changes forcing workflow modifications to keep updated with the latest versions?
greggh··on Show HN: Gemini Pro 3 imagines the HN front page 10 years from now
It was given today's front page to riff on. Thats why it not only reads like a HN front page, but also has near duplicates from todays front page.
greggh··on Google Antigravity just deleted the contents of whole drive
This is my new favorite response.
greggh··on Surprisingly, Emacs on Android is pretty good
(Travels back to the 90s)

Pretty good for Emacs*

Long live VI.

greggh··on State of Terminal Emulators in 2025: The Errant Champions
People still use WezTerm when we have Kitty and Ghostty? Can you explain why? I'm actually interested to know what would make someone make that choice.
greggh··on Tongyi DeepResearch – open-source 30B MoE Model that rivals OpenAI DeepResearch
Deepseek, Qwen, GLM (quite good). All being open and available for local use definitely puts them ahead in that space, which means a lot of the tinkerers and younger people learning to do things like train and fine-tune are getting good with Chinese models and I do think getting in early like that is a great way to gain mindshare in a space. Look at Apple or Microsoft doing everything they could early on to get their machines and software into schools as early as possible.
greggh··on Tongyi DeepResearch – open-source 30B MoE Model that rivals OpenAI DeepResearch
If you really need a lot of VRAM cheap rocm still supports the amd MI50 and you can get 32gb versions of the MI50 on alibaba/aliexpress for around $150-$250 each. A few people on r/localllama have shown setups with multiple MI50s running with 128gb of VRAM and doing a decent job with large models. Obviously it won't running as fast as any brand new GPUs because of memory bandwidth and a few other things, but more than fast enough to be usable.

This can end up getting you 128gb of VRAM for under $1000.

Page 1 of 6Next →