Nvtop: Linux Task Monitor for Nvidia, AMD and Intel GPUs
github.com
github.com
If you're here because you're interested in AI performance I'd recommend instead https://docs.nvidia.com/nsight-compute/NsightComputeCli/inde... to profile individual kernels. Nsight systems for a macro view https://developer.nvidia.com/nsight-systems and the PyTorch profiler if you're not authoring kernels directly but using something PyTorch https://pytorch.org/tutorials/recipes/recipes/profiler_recip...
But if you mean the reported utilization in nvtop is misleading I completely agree (as someone who uses it daily).
I’ve been meaning to dig into the source/docs to see what’s going on. The power usage seems to be a more reliable indicator of actual hardware utilization, at least on nvidia gear.
I'd argue GB/s memory bandwidth is more worried about at the moment.
Often you can squeeze out another order of magnitude of performance by rewriting the kernel and the power draw will always stay capped at whatever the maximum is. I'd say GPU power consumption is interesting if you're CPU bound and struggling to feed the GPU enough data and/or tasks.
https://i.imgur.com/C24EV5U.png
then sudo apt install nvtop
https://i.imgur.com/SOoCdvR.png
EDIT:
Thanks, Some people were having random problems installing WSL on their systems and I found this was the easiest solution (but based on their card models, they appeared to have much older machines.
As an aside: there is no need to install Docker Desktop just to use Docker containers in WSL either, unless you want a Windows GUI to manage your containers. Just follow the official documentation for installing Docker in your Linux distro of choice, or simply run `sudo apt install docker.io` in the default WSL Ubuntu distro. Docker will work just fine with an up-to-date WSL.
> nvtop : Depends: libnvidia-compute-418 but it is not going to be installed E: Unable to correct problems, you have held broken packages.
<rant>I find broken installs a huge turnoff, especially those related to NVIDIA. With their 2.3T market cap they can't afford someone to write an universal point and click install script for ML usage? Every time I reinstall Linux I have to spend a whole day sorting NVIDIA out. Why do they have so many layers - driver, cuda, cuda toolkit, cudnn with conflicting versioning - it's a total mess. Instead of a nice install script we have a million install guides 10 pages long, all outdated.</>
Cluster admins or Ph.D. students handle these problems, allowing people to work. All this infra is already buried under Conda, Jupyter, etc. for most people already.
Sincerely,
Your friendly HPC admin.
Back in the dark days of 2015 we used to spend a day or two just getting tensorflow working on a GPU because of all the install issues, driver issues, etc. Theano was no better, but it was academic research code, we didn't expect better.
Once pytorch started gaining ground, it forced to adapt - Keras was written to hide tensorflow's awfulness. Then Google realized it's an unrecoverable situation of technical debt and they started building JAX.
With AMD, Intel, Tenstorrent, and several other AI chip specialists coming with pytorch compatibility, NVIDIA will eventually have to adapt. They still have the advantage of 15 years of CUDA code already written, but pytorch as ab abstraction layer can make the switch easier.
I don't see how Nvidia has to do anything since PyTorch works just fine on their GPUs, thanks to CUDA. If anything, they're still one of the best platforms and that's definitely not because CUDA isn't competitive.
I hate stuff that only works on certain GPUs as much as the next person, but sadly competition has only really started to catch up to CUDA very recently.
It was an example. The example was: competition from pytorch meant that tensorflow had to improve their DX to keep up.
Nor is it because of all the tensorflow models people are writing, to be honest.
Of course you can! It's a library of vectorized math operations. You don't need to do gradient descent on the graph either.
Because of copyright, NVidia gets an explicit government-enforced monopoly over the driver implementation market. Sure, 3rd-party projects like nouveau get to "compete", but NVidia is given free reign to cripple that competition, simply by refusing to share necessary hardware (and firmware) specs; and also by compelling experienced engineers (anyone who works on NVidia's driver implementation) to sign NDAs, legally enforcing the secrecy of their specs.
On top of this, NVidia gets to be anti-competitive with the driver-compatibility of its userland software, including CUDA, GSync, DLSS, etc.
When a company's market participation is vertically integrated, that participation becomes anticompetitive. The only way we can resolve this problem is be dissolving the company into multiple market-specific companies.
I've had one or two upgrade problems in the last 10 years, but otherwise the Nvidia drivers have worked great for me. My biggest complaint is they dropped support for the GPU in my Macbook, and I had to install the nouveau drivers (which I can never spell correctly).
- "gnuveau" for one masculine GPU.
- "gnuvelle" for one feminine GPU.
- "gnuveaux" for multiple masculine GPUs.
- "gnuvelles" for multiple feminine GPUs.Never had a problem with the initial installation, updates can get messy, but asides from that it's pretty much smooth sailing.
The error message is telling you that you've held back broken packages that are conflicting with dependencies nvtop is trying to install. If you sort that out, nvtop should install.
I have nvtop installed on Debian via apt, and it works just fine.
> Currently supported vendors are AMD (Linux amdgpu driver), Apple (limited M1 & M2 support), Huawei (Ascend), Intel (Linux i915 driver), NVIDIA (Linux proprietary divers), Qualcomm Adreno (Linux MSM driver).
sudo dnf copr enable atim/bottom
sudo dnf install bottom
[1] https://github.com/ClementTsang/bottom- Show a thread with all children an threads, but nothing else
- Show the whole tree but keep the selected (Shift + Space) process in a fixed screen position.
- Bubbles rows up and down into their new positions instead of having them jump around all over the place.
{UPDATE} I see: no Intel GPU support yet!
https://github.com/XuehaiPan/nvitop?tab=readme-ov-file#insta...
https://lists.freedesktop.org/archives/nouveau/2024-February...
Other than that - a cool tool!
It's one of those things which I wish existed, but I can't imagine anyone would have written. Until I do a web search.
https://github.com/koriwi/sensors2mqtt/tree/main
I have not used it yet, but that seems like how I'd want to do it.
I can see it working only with VLC.
Firefox has some support while chromium based browsers have that only formally.
In the real world you never see the video hw acceleration kicking in, neither with webrtc nor with videos.
It is a pity.
Not at my computer now, can't tell what metric to watch.
I use several video players, including mpv, vlc and ffmpeg, and they have always used without problems the hardware video decoding and encoding on all kinds of GPUs.
Only with Firefox and Chrome/Chromium in most versions the hardware acceleration is broken, even if there have been some versions where it worked fine (on NVIDIA), but at the next browser upgrade it was broken again.
This does not bother me much, because I do not like to watch video files in a browser anyway. I always download them first and I play them locally.
Sadly enough, I need the browsers to use hw-assisted en/decoding for my video communication... Which seems to be little more than a dream, ATM.
Which leads to what I said: no hw-assisted encoding/decoding actually available.
I use intel_gpu_top (from intel-gpu-tools) to monitor my GPU usage: only VLC shows usage.
Anyone has success with browsers?
Context: I'd love to have a native Go-language library to read the GPU utilization for containerized workloads.
I would have prefered a simple and brutal shell script (not bash of course) to build it on elf/linux.
It is worth trashing cmake, always, so I'll write it if I end up using that GPU monitoring tool.