HNHacker News
TopNewBestAskShowJobs

brrrrrm

2,136 karma · joined February 9, 2021

http://twitter.com/bwasti

http://github.com/bwasti

[all posted thoughts and comments are my own]

submissionscomments
brrrrrm··on The darker side of being a doctor
I would guess “us vs them” mentalities are not good for productive societies
brrrrrm··on I asked Meta’s Muse for its filesystem and it sent me 6.8GB
I think you're confusing the expected behavior of the product offerings. Every user gets their own VM for free. would you be similarly convinced an attack has happened if AWS gave you a remote shell to the instance you rented?
brrrrrm··on Vectorized and performance-portable Quicksort
only sorts numbers? wouldn't radix be much better?
brrrrrm··on >10x More Efficient Pretraining
this is basically the only thing pre-training teams work on in labs. compute efficiency is the metric, the assumption that scaling = intelligence is considered a given.
brrrrrm··on >10x More Efficient Pretraining
they say they're looking at base models, so I think it's fairly compared as written.
brrrrrm··on The efficient frontier of LLM inference
perhaps its unfair to say this in hindsight, but it's a fairly straightforward application of little's law that's been around for some time

https://arxiv.org/html/2401.09670v2

brrrrrm··on The efficient frontier of LLM inference
this is a nice and concise writeup. what's striking to me is that these techniques really have not changed in /years/. sure, precision has become slightly lower, spec decoding acceptance has gotten slightly better and the complexity of parallelism is trickier with mixture of experts. but no new concepts in a very long time!

the absolute most impactful improvements for inference comes at architecture design time. I firmly believe everyone who cares about impacting model efficiency should look there

brrrrrm··on Qwen3.8-2.4T
this is Qwen3.8 max, right? https://qwen.ai/blog?id=qwen3.8
brrrrrm··on Muse Glimmer: 30B-parameter model optimized for always-on local agent workflows
are these all uniform quantization? or mixed and matched by layer (can't tell from the naming scheme)
brrrrrm··on Show HN: Run an 80B Qwen in 4.3 GB of RAM on a Mac, and a 35B on an iPhone
this is cool but like, are we just vibe coding NAND burners at this point? these decode times don't really tell the whole story, because prefill becomes the bottleneck.

half an hour to process 10k tokens on an M5 seems... not great

brrrrrm··on Long Range Wi-Fi – Pushing 2.4 GHz Wi-Fi to the limits (2019)
there's wifi ah, which runs on 900mhz band and has the same 30dbm limitation

I've used it with some raspberry pis to create hi-fidelity walkie talkies it's quite pleasant.

brrrrrm··on Quality non-fiction books are the antithesis of AI slop
Models are increasingly showing their ability to extrapolate into the human unexplored (math proofs being the most apparent). What gives you confidence the absurdity of life is uniquely difficult for models to source?
brrrrrm··on DSLs Enable Reliable Use of LLMs
> you are basically looking at a whole system prompt just describing the new language

whats wrong with this? You may be over-indexing on the need for large quantities of examples. These days self-play through RL is far more effective and data (not compute) efficient.

brrrrrm··on Popping the GPU Bubble
it's very much an in-domain term for folks in machine learning. heavily used when pipeline parallelism caught on in training https://alband.github.io/doc_view/pipeline.html
brrrrrm··on Ferrari Luce
it has paddle shifters - what are those for?
brrrrrm··on Running local models on an M4 with 24GB memory
what's MRT?
brrrrrm··on The best is over: The fun has been optimized out of the Internet
same can be said for a lot of things tho. e.g. nature used to be fun but then we discovered it all :’( I miss when ships literally sailed into the unknown and found surprising and novel things like hot peppers and pineapples
brrrrrm··on 1966 Ford Mustang Converted into a Tesla with Working 'Full Self-Driving'
I agree fully. Hyundai has a mockup that starts to get there (different era, but same concept) called the N vision 74[1], but I doubt we'll see it in market anytime soon. The unfortunate reality IIUC is that modern cars (electric vehicles) have certain aero restrictions (for mileage) that heavily limit design options.

[1] https://www.hyundai-n.com/en/models/rolling-lab/n-vision-74

brrrrrm··on All elementary functions from a single binary operator
meta.ai in instant mode gets it first try too (I think?)

``` 2x + y = \operatorname{eml}\Big(1,\; \operatorname{eml}\big(\operatorname{eml}(1,\; \operatorname{eml}(\operatorname{eml}(1,\; \operatorname{eml}(\operatorname{eml}(L_2 + L_x, 1), 1) \cdot \operatorname{eml}(y,1)),1)\big),1\big)\Big) ```

for me Gemini hallucinated EML to mean something else despite the paper link being provided: "elementary mathematical layers"

brrrrrm··on AI helps add 10k more photos to OldNYC
you're right, this is actually correctly placed! I was confusing the orientation. I live right around there and recognize the M&T bank in the photo on the left, so it can't be down by 9th
brrrrrm··on AI helps add 10k more photos to OldNYC
I checked 3 spots I'm familiar with and 1 is wrong

https://www.oldnyc.org/#707133f-a this is supposed to be here https://www.oldnyc.org/#702487f-a

also, if folks are interested in these old depictions of NYC, check out https://1940s.nyc/ as well!

brrrrrm··on Show HN: CineCLI – Browse and torrent movies directly from your terminal
looks cool! one bit of feedback: make your demo gif get to the point faster. either practice typing a bit quicker or speed it up 2x for the typing section
brrrrrm··on Anthropic acquires Bun
on Bun's website, the runtime section features HTTP, networking, storage -- all are very web-focused. any plans to start expanding into native ML support? (e.g. GPUs, RDMA-type networking, cluster management, NFS)
brrrrrm··on Spatial intelligence is AI’s next frontier
we've discovered some kind of differentiable computer[1] and as with all computers, people have their own interests and hobbies they use them for. but unlike computers, everyone pitches their interest or hobby as being the only one that matters.

[1] https://x.com/karpathy/status/1582807367988654081

brrrrrm··on Helion: A high-level DSL for performant and portable ML kernels
a recent wave of interest in bitwise equivalent execution had a lot of kernels this level get pumped out.

new attention mechanisms also often need new kernels to run at any reasonable rate

theres definitely a breed of frontend-only ML dev that dominates the space, but a lot novel exploration needs new kernels

brrrrrm··on John Carmack on mutable variables
one thing I've learned in my career is that escape hatches are one of the most important things in tools made for building other stuff.

dropping down into the familiar or the simple or the dumb is so innately necessary in the building process. many things meant to be "pure" tend to also be restrictive in that regard.

brrrrrm··on Poker Tournament for LLMs
> None of which are they currently capable

what makes you say this? modern LLMs (the top players in this leaderboard) are typically equipped with the ability to execute arbitrary Python and regularly do math + random generations.

I agree it's not an efficient mechanism by any means, but I think a fine-tuned LLM could play near GTO for almost all hands in a small ring setting

brrrrrm··on Reasoning LLMs are wandering solution explorers
what about "test time scaling"?
brrrrrm··on Waymo granted permit to begin testing in New York City
downshifting? these are all electric vehicles IIUC
brrrrrm··on Tour de France confronts a new threat: Are cyclists using tiny motors?
sure, numerous examples can be shown to say smart play does help. but, would you argue the net benefits of smart play are identical between a sport like basketball and racing?
Page 1 of 15Next →