They have a big advantage in that they can directly distill from frontier models.
4,994 karma · joined May 25, 2013
They have a big advantage in that they can directly distill from frontier models.
It's a shame coz it's not even really what we want, we would obviously all be better served if we could use Vulkan or something. But I guess it's inevitable that a generic framework lags behind here.
If I was AMD I'd hire a whole ecosystem team to sit next to the ROCm people and just support big users like llama.cpp to work better on their HW, e.g. giving OSS maintainers access to their board farms. Maybe they have already done that, in which case I guess I should say I'd double the size of that team.
So I think we're looking at a spectrum here:
- Fully fixed function - i.e. Taalas
- Fixed function transformer unit (or whatever other AI architecture) with a "cartridge" for weights etc - i.e. Etched
- Flexible TPU/GPU type stuff
I think the middle of the spectrum is a bit of a dead space ATM because new models have recently been coming with significant updates to the architecture (like MoE, MTP) so by the time you have new weights you wanna load, you also want to replace the compute too. So really you probably either say "I can tolerate an old model, but I want it fast as FUCK" and go for Taalas-style, or you say "I want a near-frontier model" and you have to use flexible compute anyway.
But, caveat: this comment seems to be making me sound more knowledgeable than I actually am. Take this with a grain of salt.
Completely fixed-function HW can't be used for training, it's inherently a statement that "this model is Good Enough and we are now gonna start just extracting its value instead of extending it". So yeah it's an inference moat but it's not a growth moat.
Makes perfect sense for a company trying to get into the compute business, not companies who wanna be in the creating-ASI business.
Still, I guess/hope they have teams doing it in-house anyway. Just not something they'd wanna make a huge amount of noise about, it doesn't look good for To The Moon valuations.
[0] https://www.dwarkesh.com/p/why-compute-might-get-10x-more-ex...
Taalas' approach (at least for their demo'd product) is to bake the whole thing in completely. IIUC the optimiser can even see the weights while generating RTL. It's like there's an "uint8_t weights[] = " in the source code.
Zed was a promising new direction I.e. VSCode minus the bullshit. But it's not getting the polish / productionisation it would need to actually be a better editor than the others. And now it doesn't look like it will.
I want "VScode without the bullshit". That's what I thought they were building.
JS annoyances it won't fix are like idiotic scrolling behaviour and stuff. And for that, yeah reader mode is ideal.
1. "find the code that does X"
2. go read that code
3. When you hit a bit you don't care about, go back to the model and ask it for the pertinent details
4. When you hit a really confusing bit, ask the model for hypotheses about what's going on. (I always phrase it as "give me some hypotheses" not "what is going on here". I dunno if this changes the output but I think it helps me stay in a mindset of uncertainty, it's important to avoid locking in any misunderstandings. Anyway I find the models do well at this task, and when they bullshit here it has a strong smell).
Before AI, parts 1 and 3 could be insanely time consuming, sometimes it felt like a infinite breadth-first-search. And part 4 was basically: either you find a human who knows the code, or you just make a mental note and hope that later on you find something that makes you go "oh, THAT'S why they <do weird thing that should 100% have a comment>!".
So yeah even though you're still reading code with your wetware the AI makes you dramatically more powerful.
This is also extremely helpful for unpicking undocumented API contracts. E.g. you can say "the x86 implementation of this API is safe to call under a spinlock, go read the other arch versions and tell me if that's true there too".
Right now there's just so much value in building LLM tools for experts that everyone is focusing on that. But surely at some point we'll have bespoke harnesses that exist exactly to solve this kind of thing.
I think this can start with constrained problem spaces like "you are a WordPress developer, you solve problems for people with enough expertise to know they are looking for a WordPress developer" and incrementally expand from there. Maybe I'm naive but I think you can probably get pretty far with this today just by writing loads of skills and picking the right technical preferences to encode in them.
Maybe, but FWIW my first thought when I skimmed Tao's session was that he probably has a personal system prompt requesting this style.
E.g. even if you get it into "talking to an expert" mode I've found AI waffling through filler like "given your background in Linux kernel engineering, I'll skip the surface level and go straight to the technical meat". You do have to explicitly tell them if you don't want this.
That time has now passed though.
Like, yeah if I could spend a few grand on such a GPU I probably would coz I'm a rich nerd, but I'd acknowledge it as an extremely inefficient luxury, kinda like a sports car.
So I think you could say the real misfortune is that we don't really have the technology (be it computer tech or political/social tech) to do that shared-HW thing in way we can truly trust.
I was already impressed by how fast 3.5 Flash was. But I've never compared it to other models in its class for coding.
Why? Coz models in that class are not very useful to me. Time saved waiting for responses usually just turns into time wasted replying to low quality responses.
Google need to release a Pro model ASAP. I am skeptical of the "maybe they don't have the compute to run it" thing. Anthropic were (probably) in that situation with Mythos and they announced it anyway - that's the obvious play for investor relations as well as hype for your product.
I went to view a flat this week, it was a building site. That's still the most important bit coz you get a feel for the size and shape of the space which is what really matters. But I'm glad to have the AI renders too.
Using AI for these pics is also not inherently deceptive though.
I live in an extremely overheated housing market where properties are usually sold/rented long before they actually get completed. I'm fine with landlords using AI in their renders to make claims about how the place will eventually look.
You also see people using AI to put furniture into the image (I assume they are also taking out the furniture that's actually there, belonging to the previous tenant, but doesn't fit their desired aesthetic). Again, nothing _inherently_ deceptive about this.
Main thing is just whether tenants are empowered to back out of the contract if they don't get what they were promised.
Anyone who e.g. uses AI to expand rooms/windows... Jail please.
(In practice I find this syntax super annoying and usually end up typing `git reflog mybranch` and then copy-pasting the commit hash from the output).
This is not really surprising, benchmarking _people_ doesn't work. You can only get a decent measure of someone's coding abilities by personally interacting with them. Given that models are basically person simulators it would be weird if benchmarks kept being useful as the simulation got more accurate.
I think what I've just said is basically just a more roundabout way of what you said: "Goodhart's law at work". It really is a law.
It's a extremely cool that ESPHome is able, just by existing and being good, to create this little industry of no-bullshit products. What an awesome project that is!
So maybe a few leaky cracks are enough that when you open a window you get a bit of a through-draft.
I live on a noisy street so I don't usually want to do that, if I open a window at the back and keep internal doors open it will stay reasonable but significantly elevated.
So yeah I think the lesson here is you probably need to buy a sensor, different homes are gonna differ.
My home is quite small (probably 80m²) and has literally zero ventilation built in (even in the bathroom!). I live in Switzerland where it's traditional to actively ventilate your home twice a day. But that doesn't do anything for CO2. Also it's such a fucking waste of time lol. Looking forward to moving into a modern building.
But also bear in mind that regardless of "are we operating at max effectiveness", OSHA sets a legal limit of 5000ppm in a workplace, and that's about _safety_.
This article is talking about keeping levels below 1000 which is a very high standard IMO (still arguably justified given the studies mentioned). But if you are in a poorly ventilated home office you could easily hit 3000. At that point you are closer to "illegal in the US" than "earth's atmosphere".
So yeah even if you are unconvinced about micro-optimising your CO2 levels there's a very long established argument in favour of at least paying _some_ attention to it.
Looks like it's increased in price unfortunately but I like the idea, it's basically just what you would do as a DIY project but ready built. So you can either use it like a normal commercial product, or you can just fork the ESPHome config that's on GitHub and flash it exactly like any normal ESPHome project.