HNHacker News
TopNewBestAskShowJobs

magic_at_nodai

57 karma · joined July 1, 2013

submissionscomments
magic_at_nodai··on Async/Await on the GPU
yes lmk how i can help. at the minimum i can get you hw and help with PRs etc. firstname at amd.com to reach me.
magic_at_nodai··on AMD RDNA 4 – AMD Radeon RX 9000 Series Graphics Cards
Im running ROCm ok on my 9070XT. You can build it from source today if you have a card.

rocminfo:

**** Agent 2 **** Name: gfx1201 Uuid: GPU-cea119534ea1127a Marketing Name: AMD Radeon Graphics Vendor Name: AMD Feature: KERNEL_DISPATCH Profile: BASE_PROFILE Float Round Mode: NEAR Max Queue Number: 128(0x80) Queue Min Size: 64(0x40) Queue Max Size: 131072(0x20000)

[32.624s](rocm-venv) a@Shark:~/github/TheRock$ ./build/dist/rocm/bin/rocm-smi

======================================== ROCm System Management Interface ======================================== ================================================== Concise Info ================================================== Device Node IDs Temp Power Partitions SCLK MCLK Fan Perf PwrCap VRAM% GPU% (DID, GUID) (Edge) (Avg) (Mem, Compute, ID) ================================================================================================================== 0 2 0x73a5, 59113 N/A N/A N/A, N/A, 0 N/A N/A 0% unknown N/A 0% 0% 1 1 0x7550, 24524 36.0°C 2.0W N/A, N/A, 0 0Mhz 96Mhz 0% auto 245.0W 4% 0% ================================================================================================================== ============================================== End of ROCm SMI Log ===============================================

magic_at_nodai··on ROCm Device Support Wishlist
ROCm on Radeon should work too and the poll above was to seek feedback on what to cards to support next.
magic_at_nodai··on ROCm Device Support Wishlist
I will provide this feedback to the docs team to clean up. I found it hard when i was making that Poll :D but I looked harder instead of trying to fix the docs. So thank you for the feedback.
magic_at_nodai··on ROCm Device Support Wishlist
Is this the repo you are referring to https://github.com/amd/go_amd_smi ? Would having a prebuilt version there help you ?
magic_at_nodai··on ROCm Device Support Wishlist
yes. We are behind on software support for all consumer cards and would love to support all cards. But are looking for guidance / feedback so we can prioritize.
magic_at_nodai··on ROCm Device Support Wishlist
I have quad w7900s under my desk that work well for workloads on my desktop that translate well to MI300x. There are some perf gaps with FAv2, and FP8 but otherwise I get a seamless experience. lmk if you have a pointer to any github issues for me to track down to make your experience better.
magic_at_nodai··on ROCm Device Support Wishlist
We do care about software and acknowledge the gaps and will work hard to make it better. Please let me know any specific issues that are an issue for you and Im happy to push for it to get resolved or come back with why it isn't.
magic_at_nodai··on ROCm Device Support Wishlist
PTX does provide a low level machine abstraction. However you still target some version of hardware ( https://arnon.dk/matching-sm-architectures-arch-and-gencode-... ). However a lot of software effort has gone into it to make it look and work seamlessly.

Though AMD doesn't have the same "virtual ISA" as PTX right now there are increasing levels of such abstraction available in compiled flows with MLIR / Linalg etc. Those are higher level and can be compiled / jitted in realtime to obviate the need for a low level virtual ISA.

magic_at_nodai··on ROCm Device Support Wishlist
hey thats me. Happy to help answer anything here and look forward to your constructive feedback to make AMD software better. We got work to do and look forward to it.
magic_at_nodai··on Ask HN: Who is hiring? (March 2024)
AMD Artificial Intelligence Group (AIG) | Remote / Global

AMD Artificial Intelligence Group (AIG) leads AMD AI strategy and drives AI roadmap across client, edge, and cloud. We build AI capabilities, including silicon, software, models, use cases, to create a vibrant AMD AI ecosystem together with everyone.

Our organization, AIG SHARK (formerly nod.ai), aims to build AI software solutions that are unified, performant, flexible, and customizable. We approach the mission with open-source and community-driven principles. We believe in the power of the community and together we can build a great stack that benefits everyone. Being a part of AMD and AIG, it also means we have a wide portfolio of hardware and customers to support for realized product excellence and business impacts.

Current Open Positions:

At any given time, we have various positions posted that you may apply to: General AI runtime: https://careers.amd.com/careers-home/jobs/36717 GPU compiler/runtime: https://careers.amd.com/careers-home/jobs/37400 GPU compiler/performance: https://careers.amd.com/careers-home/jobs/39454 GPU distribution/serving: https://careers.amd.com/careers-home/jobs/38113 CPU codegeneration (New Grad): https://careers.amd.com/careers-home/jobs/38616 NPU code generation/runtime : https://careers.amd.com/careers-home/jobs/39280 General Compiler: https://careers.amd.com/careers-home/jobs/40053 Build / Infra Ninjas

https://bit.ly/amd-sharks-hiring

magic_at_nodai··on Llama.cpp: Port of Facebook's LLaMA model in C/C++, with Apple Silicon support
We have it running as part of SHARK (which is built on IREE). https://github.com/nod-ai/SHARK/tree/main/shark/examples/sha...
magic_at_nodai··on Stable Diffusion on AMD RDNA3
Can you give SHARK a try and let us know on our discord? We can try to help. People have been using it on older AMD GPUs back to Polaris arch.
magic_at_nodai··on Stable Diffusion on AMD RDNA3
Here are a list of potential issues https://github.com/AUTOMATIC1111/stable-diffusion-webui/disc...

That said we (Nod.ai team) will add support for xformers soon so you can opt in for xformers anyway.

magic_at_nodai··on Stable Diffusion 2.0
Try SHARK on your AMD GPUs for SD. Follow the setup here: https://github.com/nod-ai/SHARK/tree/main/shark/examples/sha....

It works with Pytorch -> torch-mlir -> MLIR / IREE -> vulkan. Works on both Windows and Linux. And has a simple gradio web UI https://github.com/nod-ai/SHARK/tree/main/web but we plan to enable better UI integrations very soon.

Join us on discord https://discord.gg/RUqY2h2s9u if you have any trouble. Appreciate any / all feedback.

magic_at_nodai··on PyTorch on Apple M1 MAX GPUs with SHARK – faster than TensorFlow-Metal
unlikely since the interface from ANE is not public and it may change between hardware versions.
magic_at_nodai··on PyTorch on Apple M1 MAX GPUs with SHARK – faster than TensorFlow-Metal
I updated the blog with the reference. Basically it crashes to compile the model with https://github.com/NodLabs/shark-samples/blob/main/examples/.... The coremltools converter is very version specific (like all vendor conversion kits) and still on a version of TF I couldn't get on conda. Also it doesn't allow for training and only FP16 for inference with ANE. All our tests were with FP32.

//part of nod.ai/shark team.

magic_at_nodai··on PyTorch on Apple M1 MAX GPUs with SHARK – faster than TensorFlow-Metal
Yeah the ANE and AMX on cpu are wrapped behind Accelerate Framework and CoreML. So you will have to use CoreML (which wasn't able to compile the latest TF BERT). ANE is also inference only. So if you want training you will have to use the GPU with Apple's Tensorflow-Metal.

//part of nod.ai / SHARK team.

magic_at_nodai··on PyTorch on Apple M1 MAX GPUs with SHARK – faster than TensorFlow-Metal
Thanks to: LLVM/MLIR --> For the awesome compiler infrastructure

IREE --> For the awesome backend to MLIR

SHARK/nod.ai --> For adapting IREE for use on various hardware and fine tuning for target hardware.

//part of nod.ai / SHARK team

magic_at_nodai··on PyTorch on Apple M1 MAX GPUs with SHARK – faster than TensorFlow-Metal
This is not part of regular pytorch install.

If you can build torch-mlir and SHARK from src you can use it. So hopefully soon we can make pip installable packages but for now the interfaces are in constant development so you will have to build from source.

//part of nod.ai / SHARK team.