848 karma · joined June 17, 2022
Here's my low-bar result, mostly from my time at RC. Is it literally published? No. Does it have novel arguments and code? Yes. winwang.blog/posts/dphm-1
If one wanted, they could put ecosystem tooling here as well (e.g. `gofmt` saving everyone time and mental health).
Also, to clarify, I'm not arguing that Erlang > Go, I don't (purposefully) use either, though I have to read Go sometimes.
Funny stuff, as if it were hell-bent on writing a paper rather than actual software.
Although, maybe benchmarking an agent on "how well can you command a swarm to annihilate the Terrans" is how it all starts going downhill...
I also think your point has at least one decent reading: that the upvotes help other practictioners update their mental model of their tools. There's probably also some value due to being an implicit "Claude Code megathread" for commenters to congregate around. News so minor that it does't even really make sense to force people to fully stay on topic, hah.
If anyone wants to write/link a much better-thought-out post, I'm all ears!
> The best way to program GPUs is face up to the reality that they are not the same machine as the CPU, write your kernels in separate files, and launch them manually
Not to mention that this is a completely sane way to use CUDA as well.
Regardless, CPUs are really good at single-thread.
For something like these compression algos, though, I imagine it would be much easier since they already have actual proofs out there.
1MW ~ 6700 NYC citizens' residential usage, apparently, lol. I don't know exactly why, but that citizen number seemed a surprisingly large (well, probably because I've played with approximately-MW lasers).
In any case, the number I'm focusing on here is the 300k cores part (x2 if you're counting in vcores). It does not seem like too much to ask for (significantly) more than that at github scale. It doesn't feel like a hardware issue, is what I'm getting at.