I'm glad NVIDIA found a way to make GPUs programmable and got us out of the shaders tarpit, but did it have to be C++...
1,792 karma · joined March 27, 2017
I'm glad NVIDIA found a way to make GPUs programmable and got us out of the shaders tarpit, but did it have to be C++...
I can relate to the feeling when missing a dose. No tolerance for being forgetful. Lots of anxiety brought by this tether... and the permanent anorgasmia (in my case) is sad but I was dying...
Sadly I built tolerance over 2 years and had to stop, taper off (horrible experience) and try others (most that were cited in this comment section). All had painful and very annoying side effects and tapering was a pain after trying and each time feeling so disappointed and tired. Most worked on reducing ideation but the costs (inability to concentrate, brain fog, sometimes inability to work) were too high.
Turns out tolerance to venlafaxine ebbs out after some time, so after a moment I was able to restart for a year until I didn't need antidepressants anymore.
I've learned reading forums of the gamut of experiences with antidepressants and just want to add a voice to 'they can work, great even, not perfect' and between dying and trying them (with clear open eyes about the tradeoffs) I'd advise trying.
Later on used Intel Processor Trace in a similar fashion for even finer (mc-dc) coverage.
Coverage tools are very useful, if a bit hard to use...
Missing runtime checks are a pain and I remember one that struck me as hard: https://www.adacore.com/blog/the-most-obscure-arithmetic-run...
A lot of the socket featureset of io_uring seems available in AF_XDP https://docs.kernel.org/networking/af_xdp.html which shows lots of progress since I looked last.
To get an idea of what DPDK gives low-level access to there is the overview https://doc.dpdk.org/guides/nics/features.html and my "favorite annual terabit read" https://doc.dpdk.org/guides/nics/mlx5.html#mlx5-net-features for NVIDIA NICs. Broadcom has some fun stuff too. The first time you hit top RX speed (2x400G my latest) with only one busy core (yay DMA engines) is always a thrill.
I need to check about the af_xdp ecosystem around fragmentation/reassembly in UDP too, every time I needed something there DPDK had it, often with an offload path.
Some silly stuff in DPDK are very useful for testing too (in-memory devices).
Also I'm not clear on the virtualization story on af_xdp, with dpdk I got something working at full blast 400G in VMs with little (but finnicky) work.
I was remarking on the MI350P because I've had a hard time procuring "small" CDNAx systems (for e.g. development, experiments and lower-profile servers) and OAM seemed very niche (not if you're aiming for density and training/inference...).
I hope the MI350P fills a lower part of the spectrum and I can start massively porting CUDA stuff or at least work on HIP and ROCm and what I need to make most or some of our CUDA stuff run on AMD HW, then how make it run fast.
Same for gpudirect (more useful for scale-out or training).
Everyone is focused on AI but these are two interesting techs trickling down from the NVIDIA tree, relatively "easy" to use there, which maybe exist in AMD world but I somehow missed the docs and APIs and demos on how to use them...
I only regret not writing the obvious-but-buggy code that "forgot" or added some values and still passed proof...
Did you have a look at why3 and generating verification conditions from Rust or C code (as frama-c does) ?