HNHacker News
TopNewBestAskShowJobs

bjackman

4,994 karma · joined May 25, 2013

[ my public key: https://keybase.io/bjackman; my proof: https://keybase.io/bjackman/sigs/VqJOkkjZ3NpAA_GO1V1yRLqqvoDbwJu0bLFsLxgdj8c ]
submissionscomments
bjackman··on Grok 4.6
> With that difference in training speed, it should be impossible for Chinese models to close the gap that easily.

They have a big advantage in that they can directly distill from frontier models.

bjackman··on llama.cpp
I think ROCm is just a total second class citizen in the space TBH.

It's a shame coz it's not even really what we want, we would obviously all be better served if we could use Vulkan or something. But I guess it's inevitable that a generic framework lags behind here.

If I was AMD I'd hire a whole ecosystem team to sit next to the ROCm people and just support big users like llama.cpp to work better on their HW, e.g. giving OSS maintainers access to their board farms. Maybe they have already done that, in which case I guess I should say I'd double the size of that team.

bjackman··on Tom Stanton's supersonic trebuchet breaks sound barrier with gravity alone
It's a unit of mass used in the firearms world
bjackman··on AMD acquires Taalas to boost inference performance by etching models in silicon
I believe the fully baked-in nature is pretty important for the perf they get. With a cartridge port you now have a bus between the weights and the compute and the weights and that bus can become a bottleneck.

So I think we're looking at a spectrum here:

- Fully fixed function - i.e. Taalas

- Fixed function transformer unit (or whatever other AI architecture) with a "cartridge" for weights etc - i.e. Etched

- Flexible TPU/GPU type stuff

I think the middle of the spectrum is a bit of a dead space ATM because new models have recently been coming with significant updates to the architecture (like MoE, MTP) so by the time you have new weights you wanna load, you also want to replace the compute too. So really you probably either say "I can tolerate an old model, but I want it fast as FUCK" and go for Taalas-style, or you say "I want a near-frontier model" and you have to use flexible compute anyway.

But, caveat: this comment seems to be making me sound more knowledgeable than I actually am. Take this with a grain of salt.

bjackman··on US strikes $1.2B deal to pay German firm to halt offshore wind projects
For anyone who wants to learn more about this claim, I thought this video was cool: https://youtu.be/M9Hoy9Kn1u0?is=1O5AJPEziihH1MSJ
bjackman··on AMD acquires Taalas to boost inference performance by etching models in silicon
Dwarkesh recently pointed out [0] that these guys are almost forced to spend most of their compute on training instead of inference. This is because they need to maintain the appearance (which may also be the truth) that future models will make current models obsolete and be much more valuable.

Completely fixed-function HW can't be used for training, it's inherently a statement that "this model is Good Enough and we are now gonna start just extracting its value instead of extending it". So yeah it's an inference moat but it's not a growth moat.

Makes perfect sense for a company trying to get into the compute business, not companies who wanna be in the creating-ASI business.

Still, I guess/hope they have teams doing it in-house anyway. Just not something they'd wanna make a huge amount of noise about, it doesn't look good for To The Moon valuations.

[0] https://www.dwarkesh.com/p/why-compute-might-get-10x-more-ex...

bjackman··on AMD acquires Taalas to boost inference performance by etching models in silicon
No I think you are thinking of Etched/Sohu.

Taalas' approach (at least for their demo'd product) is to bake the whole thing in completely. IIUC the optimiser can even see the weights while generating RTL. It's like there's an "uint8_t weights[] = " in the source code.

bjackman··on Zed DeltaDB
I'm not satisfied with any of the text editors I've used (I've used all the big ones for a few years each).

Zed was a promising new direction I.e. VSCode minus the bullshit. But it's not getting the polish / productionisation it would need to actually be a better editor than the others. And now it doesn't look like it will.

bjackman··on Zed DeltaDB
I don't want someone "to make it simpler for developers to follow changes applied by an LLM in an ergonomic fashion". I want a text editor.

I want "VScode without the bullshit". That's what I thought they were building.

bjackman··on Changes at Google DeepMind: Demis Hassabis from CEO to Chair, Jeff Dean departs
this prompted me to search and find: https://addons.mozilla.org/en-US/firefox/addon/xcancel/
bjackman··on TIME Is Serving AI Bots a Different Website, with Ads Built In
ABP works for the most common of these too (e.g. popups/banners)! You have to enable it in the settings I think.

JS annoyances it won't fix are like idiotic scrolling behaviour and stuff. And for that, yeah reader mode is ideal.

bjackman··on TIME Is Serving AI Bots a Different Website, with Ads Built In
Firefox has this too, it's useful. However I only find that I use it for websites that have shit CSS that makes it hard to read on mobile. For everything else AdBlock Plus works fine.
bjackman··on Why did we wait so long for the bicycle? (2019)
You know if you tap on the text at the top, it takes you to an article that expands upon the question. In this case you'll find that article discusses the strengths and weaknesses of the materials bottleneck!
bjackman··on LLMs reward expertise
In the Linux kernel I've had a lot of luck with:

1. "find the code that does X"

2. go read that code

3. When you hit a bit you don't care about, go back to the model and ask it for the pertinent details

4. When you hit a really confusing bit, ask the model for hypotheses about what's going on. (I always phrase it as "give me some hypotheses" not "what is going on here". I dunno if this changes the output but I think it helps me stay in a mindset of uncertainty, it's important to avoid locking in any misunderstandings. Anyway I find the models do well at this task, and when they bullshit here it has a strong smell).

Before AI, parts 1 and 3 could be insanely time consuming, sometimes it felt like a infinite breadth-first-search. And part 4 was basically: either you find a human who knows the code, or you just make a mental note and hope that later on you find something that makes you go "oh, THAT'S why they <do weird thing that should 100% have a comment>!".

So yeah even though you're still reading code with your wetware the AI makes you dramatically more powerful.

This is also extremely helpful for unpicking undocumented API contracts. E.g. you can say "the x86 implementation of this API is safe to call under a spinlock, go read the other arch versions and tell me if that's true there too".

bjackman··on LLMs reward expertise
I have been pondering this and I think it's likely a gap that will get filled sooner or later.

Right now there's just so much value in building LLM tools for experts that everyone is focusing on that. But surely at some point we'll have bespoke harnesses that exist exactly to solve this kind of thing.

I think this can start with constrained problem spaces like "you are a WordPress developer, you solve problems for people with enough expertise to know they are looking for a WordPress developer" and incrementally expand from there. Maybe I'm naive but I think you can probably get pretty far with this today just by writing loads of skills and picking the right technical preferences to encode in them.

bjackman··on LLMs reward expertise
> The model outputs are much more concise than when I try and talk to GPT-5.6 Sol about mathematics. By signalling expertise, Tao shunts the model into “talking-to-mathematicians”

Maybe, but FWIW my first thought when I skimmed Tao's session was that he probably has a personal system prompt requesting this style.

E.g. even if you get it into "talking to an expert" mode I've found AI waffling through filler like "given your background in Linux kernel engineering, I'll skip the surface level and go straight to the technical meat". You do have to explicitly tell them if you don't want this.

bjackman··on Don't be a meat proxy
I think it was reasonable in the very recent past. There was a period where AI had major capabilities that had not diffused fully into the zeitgeist. So often you could see someone struggling with e.g. understanding a crash log, and say "they probably haven't thought of pasting into Claude Code with access to the codebase".

That time has now passed though.

bjackman··on Kimi-K3 on HuggingFace
It doesn't matter if you have the RAM, running a 1.5TB model for a single context stream is fundamentally inefficient.
bjackman··on Kimi-K3 on HuggingFace
I agree but worth noting that it's never gonna be very practical to run LLMs like this at home. Unless we have some sort of design breakthrough, the only "sensible" way to run them is at high batch levels on shared HW.

Like, yeah if I could spend a few grand on such a GPU I probably would coz I'm a rich nerd, but I'd acknowledge it as an extremely inefficient luxury, kinda like a sports car.

So I think you could say the real misfortune is that we don't really have the technology (be it computer tech or political/social tech) to do that shared-HW thing in way we can truly trust.

bjackman··on Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber
> This is a very fast model.

I was already impressed by how fast 3.5 Flash was. But I've never compared it to other models in its class for coding.

Why? Coz models in that class are not very useful to me. Time saved waiting for responses usually just turns into time wasted replying to low quality responses.

Google need to release a Pro model ASAP. I am skeptical of the "maybe they don't have the compute to run it" thing. Anthropic were (probably) in that situation with Mythos and they announced it anyway - that's the obvious play for investor relations as well as hype for your product.

bjackman··on NYC may require landlords and realtors to disclose the use of AI in listings
Coz you want to know what flooring, doors, cupboards, bath, sinks, railings, windows, etc they are putting in.

I went to view a flat this week, it was a building site. That's still the most important bit coz you get a feel for the size and shape of the space which is what really matters. But I'm glad to have the AI renders too.

bjackman··on NYC may require landlords and realtors to disclose the use of AI in listings
Requiring disclosure seems obvious.

Using AI for these pics is also not inherently deceptive though.

I live in an extremely overheated housing market where properties are usually sold/rented long before they actually get completed. I'm fine with landlords using AI in their renders to make claims about how the place will eventually look.

You also see people using AI to put furniture into the image (I assume they are also taking out the furniture that's actually there, belonging to the previous tenant, but doesn't fit their desired aesthetic). Again, nothing _inherently_ deceptive about this.

Main thing is just whether tenants are empowered to back out of the contract if they don't get what they were promised.

Anyone who e.g. uses AI to expand rooms/windows... Jail please.

bjackman··on The git history command
After a rebase, mybranch@{1} refers to the previous location of mybranch, so you don't need to manually track these before-rebase branches etc.

(In practice I find this syntax super annoying and usually end up typing `git reflog mybranch` and then copy-pasting the commit hash from the output).

bjackman··on An update on residential proxies and the scraper situation
Google should be able to detect this and ban those apps from the Play Store. They have the incentive too.
bjackman··on Separating signal from noise in coding evaluations
Even if nobody is "cheating" your particular definition of cheating, the benchmarks are _somewhere_ in the super-structural gradient descent. Models are benchmark-maximising machines at some level, so I think the benchmarks are inherently a bit useless.

This is not really surprising, benchmarking _people_ doesn't work. You can only get a decent measure of someone's coding abilities by personally interacting with them. Given that models are basically person simulators it would be weird if benchmarks kept being useful as the simulation got more accurate.

I think what I've just said is basically just a more roundabout way of what you said: "Goodhart's law at work". It really is a law.

bjackman··on The bottleneck might be the air in the room
Great to see there are others doing the same.

It's a extremely cool that ESPHome is able, just by existing and being good, to create this little industry of no-bullshit products. What an awesome project that is!

bjackman··on The bottleneck might be the air in the room
Yeah I was thinking airtightness might be the difference. My flat seems to be bizarrely hermetic (when you turn on the kitchen extractor fan, it struggles if you don't have a window open somewhere).

So maybe a few leaky cracks are enough that when you open a window you get a bit of a through-draft.

bjackman··on The bottleneck might be the air in the room
That's interesting coz I found the opposite, at my place to keep the level below 1k I usually have to open a window in the room I'm in, or use a fan.

I live on a noisy street so I don't usually want to do that, if I open a window at the back and keep internal doors open it will stay reasonable but significantly elevated.

So yeah I think the lesson here is you probably need to buy a sensor, different homes are gonna differ.

My home is quite small (probably 80m²) and has literally zero ventilation built in (even in the bathroom!). I live in Switzerland where it's traditional to actively ventilate your home twice a day. But that doesn't do anything for CO2. Also it's such a fucking waste of time lol. Looking forward to moving into a modern building.

bjackman··on The bottleneck might be the air in the room
IMO it's something where an intervention is often cheap enough that it's worth it even without great evidence.

But also bear in mind that regardless of "are we operating at max effectiveness", OSHA sets a legal limit of 5000ppm in a workplace, and that's about _safety_.

This article is talking about keeping levels below 1000 which is a very high standard IMO (still arguably justified given the studies mentioned). But if you are in a poorly ventilated home office you could easily hit 3000. At that point you are closer to "illegal in the US" than "earth's atmosphere".

So yeah even if you are unconvinced about micro-optimising your CO2 levels there's a very long established argument in favour of at least paying _some_ attention to it.

bjackman··on The bottleneck might be the air in the room
As a middle ground I can also recommend this unit: https://apolloautomation.com/products/air-1

Looks like it's increased in price unfortunately but I like the idea, it's basically just what you would do as a DIY project but ready built. So you can either use it like a normal commercial product, or you can just fork the ESPHome config that's on GitHub and flash it exactly like any normal ESPHome project.

← PreviousPage 2 of 32Next →