HNHacker News
TopNewBestAskShowJobs

hadlock

3,720 karma · joined October 6, 2017

https://github.com/hadlock
submissionscomments
hadlock··on The car industry A/B tested selling a car with and without CarPlay
The japanese branded car outselling the domestic branded version in 2025 shouldn't surprise any millennial. If you grew up in the 80s or 90s getting shuttled around in a GM you almost certainly remember a piece of the interior coming off in your hand, or if you're an elder millennial, writing your name with your fingernail on the UV damaged interior of a late 70s/early 80s GM car that's only 3-5 years old. It's hard to shake that kind of shoddy build quality even decades later.

Meanwhile my friends kids are turning 16 and inheriting the honda accord with 250,000 miles I remember getting shuttled to and from soccer practice when I was 12 in the 90s

hadlock··on GPT-6 Astra on robot arms
There's definitely a wide variance in laundry. Laundry is something that I can do to kill time while boiling pasta on the stove, for my partner it is an entire afternoon task.
hadlock··on The asteroid currently hitting front end web development
>The key is you're doing it yourself, so it takes your time away from something else.

There is still a cost to the business owner, they have to find a place to hire a freelancer, interview and then explain what problem they need solved, go through several iterations with them (either via meeting or email, or both) etc etc. If you're a small business that's not setting up an online storefront, the time spent setting it up and building it yourself is possibly less time. And then there's the bigger headache: what happens when the freelancer stops returning your calls? For the average small business website/page it is 2-10 pages. The average business owner can probably crank out that site with AI in less time than setting up a single interview.

hadlock··on Muse Spark 1.3
By their own benchmarks it is about 10% lower scoring than Qwen 3.6 35b-a3b, but I've added it to my list. Always looking for MoE to compare to it so we can squeeze more out of our local LLM system.
hadlock··on Muse Spark 1.3
There's a lot of value in agentic loop tool failure + recovery training data
hadlock··on Muse Spark 1.3
I think a lot of people are curious where the "knee" is on gains and productivity, particularly in the agentic space, which is where the real value is. A lot of us are being forced to shoe-horn this stuff into existing products, and knowing how much of the task the model can do now, vs having to build a complex custom harness, is valuable information to have. A year and a half ago it took our dev maybe six weeks of struggling with LangChain to approximate what Claude + MCP server can do today. The MCP server took us perhaps 2 days to build and 3 more to get it production ready. Today that MCP server gets 2-3 commits per month. I absolutely want to know when new models come out.

As for smaller models, we run a pretty wide variety of agentic workload doing data enrichment and, increasingly, a bunch of evaluation jobs to alert a human to review certain scenarios etc. These all run on the smaller 27B and 35B class models, and tooling behavior has improved DRAMATICALLY since april. The latest qwen 3.8 model has a 95% success tool call rate during internal testing and about 94% real world. That's about 3% better than the 35B-A3B model we're using today, but the 35B MoE is so much faster then 3% is worth the trade-off.

hadlock··on Fable 5.1 World Modeling
This is the kind of demo that would really benefit from a youtube demo of 2-3 min

Fake edit: there is a longer video here 1 min long: https://github.com/PhiloLabs/fable51-worlds/blob/main/union-...

I would be especially curious to see the NPC person/car logic and if they're on rails or what, that's a pretty good NPC density for a demo.

hadlock··on Show HN: Running 104GB Qwen3.8-Flash-Next on 48GB Mac with at ~12 tok/s
We are running 35b-A3b with 264k context (the model's default max) using vllm and the "frog" jinja templates: https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates and had good luck. We are mostly running agentic workloads though, rather than coding. 27b has a slightly higher agentic job completion rate (95% vs 92%) but the 3% trade off is worth it because the A3B is sooooo much faster, and we reprocess the other jobs with a different model. Don't sleep on the froggeric templates.

Qwen: Looking at you for a new ~35B MoE! Please and thank you

hadlock··on Play Store blocks AuroraStore, hurting GrapheneOS users
Right, the government pulled a lever, and google complied within days. If the FTC declares app stores can't provide security updates without government license, that is another lever they can pull, and google will comply.

Wether or not the most recent example is the best example, doesn't matter. What matters is when the government says "jump" in legalese, google's lawyers say "how high?"

hadlock··on Play Store blocks AuroraStore, hurting GrapheneOS users
It seems wise to have at least one alternative mobile phone app store. Even if it isn't very good. If the government can tell Google to do trivial things like, for example, change the name of bodies (plural now) of water, it can turn off your app updates, trapping you on insecure versions indefinitely. This probably matters more if you live outside of the US, but if I had a plan B for an app store on my phone, I would certainly at least evaluate it.
hadlock··on Show HN: OpenTIE and OpenXWA, Modern Ports of Tie Fighter and X-Wing Alliance
I think a lot of the value of a total rewrite in C is that it is a buildable executable that should be runable on virtually any system, given that this game predates the popularity of the modern "3D Card". A lot of the friction in these games is dealing with all the peculiarities of getting a DOS game to somehow install on a 30 year newer Mac. A complete rewrite eliminates all of that awkward installer friction, so long as you have the necessary game files.
hadlock··on I measured what my Claude.md, skills and hooks are worth
>after a Claude Code release, does my setup still do what I think it does

We blocked out ~2 months to build out a testing system for our internal agentic business processes, and I just found out there's a startup with $50 million in seed funding to answer this exact question. I'm sure others will be working on this in the future, as the preferred model for local LLM seems to change every ~6 weeks.

hadlock··on Microduck
Having tried (and failed) to use Issac to do the same thing, the fact that their simulator actually works out of the box, "batteries included", is a big deal for home hobbyists.
hadlock··on Microduck
What is really interesting for me is that they're not using Nvidia's Issac, which is famously impossible to get up and running for a custom robot by individual contributors. I can't tell how good their system is, but the fact that I got it running on my laptop in under an hour, vs wasting over a week of my spare time on Issac (ref: countless reddit threads by hundreds of people also failing to train their robot on Issac) gives me confidence I can finally convert my terrible IK stuff on my robot, to use this other framework.

This appears to be using mjlab, which is is MuJoCo Warp plus rsl_rl, which uses a completely different design/technology heritage from Issac, so that's worth a lot to me, since I've basically sworn off Issac for robot gait training.

hadlock··on Show HN: Yet another minimal and lightweight terminal multiplexer written in Go.
I hadn't seen herdr until just now but I built a similar tool with it's own keyboard bindings https://github.com/hadlock/cscope . Claude Code changed some underlying changes to copy-paste so it needs updating. But I think this kind of "I wish my workspace did this/had these features" is a pretty common trope. It is about $10 worth of your employer's monthly token allocation (or personal account if you pay for one) and then improves output so I call it a win-win. Mine is a bit more chaotic, 80s era neon color scheme but accomplishes the same feature. All those "personal tooling" blog posts get voted up for a reason.

TL;DR why would I accept someone elses' vibe coded tool, when I can build my own to my exact specification?

hadlock··on Humanity has the debate about AI consciousness backwards
It might be time to take a sabbatical
hadlock··on M5Stack Launches PaperMono
The "original" XTEINK X4 (not pro) has a USB-C port - that's explicitly why I bought it. It looks like it is out of production now. The good news is they are a proven, very popular device (it's just an e-ink display glued to an ESP32 with good third party support) so this class of device should be available on AliExpress soon if not already.
hadlock··on Qwen3.8-Flash-Next
Right now qwen 3.6 35b-a3b has a success rate of 92% and qwen 3.8 27b has a success rate of 96%. But the 35b moe does about 1080 tokens/s at concurrency 54, vs 480 tokens/s at concurrency 28. For our specific workflow on blackwell.

Of course enormous batch jobs are different. I was explicit when I said consumer laptop.

hadlock··on Qwen3.8-Flash-Next
I suspect we will see optimizations where the various vectors of the n-gram you actually use are hot in vram, the rest are warm in system memory and then cold storage on nvme. Same with MoE. If your workflow is particularly same-y then you're looking at cache miss below 5% with NTP/MTP turned on and the right harness. Agentic "openclaw" type stuff cache miss might be below 1% in the right local llm setups. There's been zero exploitation of n-gram stuff yet, it will be very interesting as things progress.
hadlock··on GitHub Publishes IPv6 Addresses for Git SSH remotes
It means if your network is IPv6 native, you don't have to waste a very expensive IPv4 IP address, or proxy through one, to connect to it
hadlock··on It’s so hard to finish an idea that is not yours and is just suggested by AI
>I don't get the second brain thing. I don't keep a second brain, I just read and think a lot.

Yes, but writing also helps you lens your thoughts in a completely different way. This is why so many people journal, or blog, or vlog, even though only a handful of people might actually audit their output. Obsidian sort of abdicates a lot of the value of journaling, but I think there's still some (although much less, perhaps 5% as much) value in the actual act of directing the journaling, even if the rest of the thinking is outsourced to the clanker.

hadlock··on Qwen3.8-Flash-Next
the important thing is that Qwen 3.7 27B will run unlimited jobs on my consumer grade laptop at 60 tokens/second for free, forever, in about 1-2 years
hadlock··on GLM-5.3-Flash
set thinking to minimal and use these jinja templates: https://huggingface.co/froggeric/Qwen-Fixed-Chat-Templates

we went from 62% completion to 92% using a claude code harness

hadlock··on Starbase, LA
Most of the erosion is due to oil and gas companies dredging poorly thought out canals to move drilling equipment through the marsh, causing unslowed erosion for ~70+ years.
hadlock··on Starbase, LA
The oil and gas companies brought in huge massive drilling rigs in the 1930s-1960s. The marsh was too soft/soggy to bring in the equipment overland, so they just... dredged huge canals in the soggy marsh a couple feet deep. You can see this in the opening shot of the video with the boat in the narrow canal. Labor was cheap and they were making a ton of money. They would put the dredged mud on one side of the channel, at random, causing random tracts to flood and most of the plants died, and then the marshes slowly sank below the water. One EPA report says for every acre of dredging they did (and they did a LOT) cause 3 more additional acres of erosion. If you go zoom in on the area where they're building, there are big squares and rectangles of dredging material, surrounding what look like square lakes. The state has been trying to get the oil and gas companies to restore the land but not much luck. Part of this deal with SpaceX (allegedly) is to restore the land and close off the canals to prevent further erosion.
hadlock··on OpenAI: GPT 5.6 Sol price reduction (until at least Nov 21)
It turns out you can train a 1b model at almost 1000 tokens/s on a m5 max laptop. As a personal experiment, I've been asking Sol for synthetic training data and synthetic agentic training data (model distillation in it's purest form), plus modified opencode, codex transcripts etc for training data, and nobody's even paying me to do it. If I'm doing it has a hobby, you can bet industrial users are doing it.
hadlock··on AI Chip Architectures
Maybe it's just me, but between the extremely thin font and layout design, I find this extremely difficult to parse. Overuse and improper use of italics is confusing as well.
hadlock··on Stop Making TUIs
I guess you haven't tried ratatui then. All my cli offer a ratatui dashboard mode, and about 50% of the dashboard mode end up as a status dashboard down the road
hadlock··on Stop Making TUIs
MacOS: what is my purpose

Me: you exist so I can launch iTerm2, Chrome, and VS Code

MacOS: oh, my god

I can't ever imagine using walled garden graphics api in 2026. Particularly for work tooling

hadlock··on What Happens When the Cost of Intelligence Drops 100x
I suspect due to model distillation, their need to stay on top for IPO value is more important than absolutely crushing their competition, who will simply harvest their output and distill it for training data on a ~3 month delay. Also they are probably hitting a wall on increase in intelligence vs training time.
← PreviousPage 2 of 34Next →