Full threadasdaqopqkq·aren't llms smart enough to directly write custom kernels for custom hardware from cuda code?View on HN