HNHacker News
TopNewBestAskShowJobs

alexbaden

27 karma · joined July 29, 2020

submissionscomments
alexbaden··on Intel will start making GPUs
I have a B580 too. The cool thing about it is architecturally speaking it is basically a mini version of the Ponte Vecchio (PVC) datacenter GPU. You can run most of the datacenter GPU workloads, albeit scaled down to fit the compute/memory constraints of the B580. It's a great vehicle for software development. But you can't buy PVC anymore so it's unclear what you are developing for...
alexbaden··on Intel will start making GPUs
Intel has been designing GPUs manufactured on TSMC nodes across client and datacenter for at least the past 5 years. The client chips are price competitive but not performance competitive with AMD/NVIDIA/Apple. The data center roadmap has historically been a huge mess with cancelled products left and right. But, to say "Intel will start making GPUs" seems misleading. Perhaps "Intel to try to inject sanity into its GPU roadmap" would be a better headline, though I am skeptical one hire will do anything to fix 10+ years of mismanagement.
alexbaden··on Tinygrad 0.9.0
Arguably the nvidia AI moat is PyTorch and the heavily optimized libraries behind it. The CUDA language and toolchain helped get that effort off the ground, no doubt, but PyTorch is written and optimized for CUDA first. All other backends work best with similar semantics to CUDA and have to match Cuda semantics to keep their users happy.
alexbaden··on 1.1B Taxi Rides Using OmniSciDB and a MacBook Pro
(disclaimer: I work for OmniSci)

I think this is a good point. On GPUs, SIMT is effectively automatic vectorization, so our focus has been on the memory bandwidth wall (we make use of cuda shared memory in nvidia GPU mode for aggregates like the above query). Non-random access compression on GPUs also has been a nonstarter, at least historically. With more recent GPUs and more recent versions of CUDA, perhaps this is changing. But on CPUs, we have started looking into vectorization. There is a tradeoff, though -- the vectorization LLVM passes do add time to the compilation phase, and at subsecond query speeds that time isn't always worth it.

There are also a few other tricks to get closer to roofline performance. If you sort the input data on the key you're grouping by you can see small performance improvements, mostly from better cache locality. But, part of the "magic" of OmniSciDB is that you can group on any key and get good performance without ingesting, reindexing, etc.