1,715 karma · joined September 19, 2018
Anyway, here’s Trump himself detailing the extraordinary access to White House this lobbying bought Adelsons:
https://www.reuters.com/world/us/trump-salutes-mega-donor-mi...
Side note & a hot take: that sort of abstraction never really existed for GPU and it's going to be even harder now as Nvidia et al races to put more & more specialized hardware bits inside GPUs
That is, where does it truly make a difference to dispatch non-parallel/syscalls etc from GPU to CPU instead of dispatching parallel part of a code from CPU to GPU?
From the "Announcing VectorWare" page:
> Even after opting in, the CPU is in control and orchestrates work on the GPU.
Isn't it better to let CPUs be in control and orchestrate things as GPUs have much smaller, dumber cores?
> Furthermore, if you look at the software kernels that run on the GPU they are simplistic with low cyclomatic complexity.
Again, there's a obvious reason why people don't put branch-y code on GPU.
Genuinely curious what I'm missing.
Something I recently learnt: the actual number of physical registers in modern x86 CPUs are significantly larger, even for 512-bit SIMD. Zen 5 CPUs actually have 384 vectors registers, 384*512b = 24KB!
As a HPC developer, it breaks my heart how worse academic software performance is compared to vendor libraries (from Intel or Nvidia). We need to start aiming much higher.
Re (b) I'm curious what that middle ground is. Is there any simple refactor to help GCC to get rid of this `if`? (Note, ISPC did fine here)
(c) Just to be clear, all the codes in benchmark figures (baseline and SIMD) were compiled with fast-math flags.
Regarding (a), one of the points I wanted to get across was that it didn't feel that complicated to program in the end as I had thought. Porting to AVX-512 felt mechanical (hence the success of LLMs in one-shotting the whole thing).
This is a subjective opinion, depends on programmer's experience etc- so I won't dwell on it. I just wish more CPU programmers gave it a try.
new_box.mins = _mm_min_ps(a.mins[3], b.mins[3]);
Of course some differences exist (e.g. basis vectors are fixed in FFT, unlike PCA).
AI is always good at going from 0 to 80%, it's the last 20% it struggles with. It'd be interesting to see a claude-written code making its way to a well-established library.
As a Bengali man, that's exactly how I felt when I came to USA and first visited japanese restaurants. Part of the reason we consume so much rice is that rice is kind of the main dish (not a side)- it literally takes up central and most of the space in your food plate.
https://commons.wikimedia.org/wiki/File:%E0%A6%87%E0%A6%B2%E...
For example, the cloth bending simulation is almost entirely: at __init__, call a function to add a cloth mesh to model builder obj, pass built model to initializer of a solver class; and at each timestep: call a collide model function, then call another function called solver.step. That's really it.
But I remember their entire Microsoft episodes felt like a lengthy defense of Steve Ballmer. There were too many instances of “here’s why this bad decision of Steve made sense given the circumstances” or “ here is how people underestimate the contribution of Steve on this good decision.” They were all well argued points, of course, but so numerous that I found myself wondering if the hosts does not have a relationship with Steve.
The existence of this interview does not help with that suspicion.