HNHacker News
TopNewBestAskShowJobs

roanakb

61 karma · joined March 24, 2019

cofounder @ Trainy (yc s23)

prev ML lead @ Hive.ai

submissionscomments
roanakb··on GPU utilization can be a misleading metric
Unfortunately, SM efficiency is not accessible via nvidia-smi. The best methods to track it would be to:

1. Profile your model with Pytorch Profiler 2. Export metrics with Nvidia DCGM

roanakb··on GPU utilization can be a misleading metric
oh this looks great, thank you for bringing this up! I'll have to give it a try, but seems like the FSDP limitation on torch.compile might carry over?
roanakb··on GPU utilization can be a misleading metric
Yup, you'll see 100% utilization on a kernel over a time period if it's considered active, which includes just having a single thread executing [1]. SM occupancy is great but can be a little difficult to interpret since you're not simply trying to maximize it, unlike SM efficiency.

[1]: https://pytorch.org/blog/pytorch-profiler-1.9-released/#gpu-...

roanakb··on GPU utilization can be a misleading metric
Nice, seems like ML Productivity Goodput is a pretty well thought-out metric to understand the overall efficiency of your cluster. I'll consider adding this into our cluster management platform. Only potential drawbacks I'd guess are it being somewhat difficult to compute since it relies on metrics like MFUs, and not something we can observe layer-by-layer to understand inefficient kernels, but I'll take a deeper look. Thanks!
roanakb··on GPU utilization can be a misleading metric
Agreed, roofline plots would be quite powerful in this context. From a quick search, seems like the only way to create a roofline plot for your model would be to use Nsight [1]? Would be interested to know if there are any simpler tools, since one of the big benefits of SM efficiency is how easily the metric is accessed.

[1]: https://www.nvidia.com/en-us/on-demand/session/gtcspring21-s...

roanakb··on GPU utilization can be a misleading metric
Yup, similar to SM efficiency in that sense too. If you aren't seeing >80%, there is certainly time left on the table. But getting a high SM efficiency value doesn't guarantee you're making good use of the hardware as well. (still a better proxy than GPU util though)
roanakb··on Instrumentation checklist for running large GPU clusters
this is a good one for debugging rdma: https://docs.redhat.com/en/documentation/red_hat_enterprise_...
roanakb··on How to Use PromptGuard – Meta's Llama Moderation Model
Looks great, you guys made it really easy to integrate!
roanakb··on Monitor and Optimize your large-scale model training
Thanks! Let me know if there are any features you'd like to see added.
roanakb··on Launch HN: PeerDB (YC S23) – Fast, Native ETL/ELT for Postgres
Looks really cool! Nice work.
roanakb··on Ask HN: Who wants to be hired? (October 2020)
Location: Fremont, CA Remote: Yes Willing to relocate: No Technologies: Python, R, Pytorch, Keras, SciPy, Machine Learning, Deep Learning, Computer Vision, NLP, Javascript, MERN Stack Résumé/CV: on request Linkedin: https://www.linkedin.com/in/roanakb/ Email: roanakbaviskar at gmail dot com

I have experience with Machine Learning research and have implemented SoTA architectures. I am comfortable with Python and various ML/DL libraries.