HNHacker News
TopNewBestAskShowJobs

shreyansh26

11 karma · joined August 18, 2017

submissionscomments

Understanding Multi-Head Latent Attention (From DeepSeek)

shreyansh26.github.io·2 pts·shreyansh26·
1

Deriving the gradient for the backward pass of Layer Normalization

shreyansh26.github.io·3 pts·shreyansh26·
0

GTC'25 Notes: CUDA Techniques to Maximize Memory Bandwidth – Part 1

shreyansh26.github.io·1 pts·shreyansh26·
0

FlashAttention in PyTorch

github.com·2 pts·shreyansh26·
1

Understanding FlashAttention

shreyansh26.github.io·2 pts·shreyansh26·
0

Ask HN: What are some good resources on Recommender Systems?

14 pts·shreyansh26·
3