HNHacker News
TopNewBestAskShowJobs

lmxyy

105 karma · joined March 21, 2023

submissionscomments
lmxyy··on Radial Attention: O(nlogn) Attention for Long Video Generation with 2-4× Speedup
Introduce Radial Attention — a static sparse attention mechanism with O(nlogn) complexity for long video generation! Here are some key features: * Plug-and-play: works with pretrained models like Wan, HunyuanVideo, Mochi * Speeds up both training&inference by 2–4×, without quality loss * Compatible with pre-trained LoRAs. When applied to 8-step FusionX LoRA, Radial Attention further delivers a 1.6× speedup
lmxyy··on SVDQuant+NVFP4: 4× Smaller, 3× Faster FLUX with 16-bit Quality on Blackwell GPUs
Thanks for pointing this out. I've fixed the prompt. Both FLUX and PixArt use T5 for the text encoder, which has limited capability. Our quantization method can preserve the image quality and contents of the original 16-bit ones well.
lmxyy··on SVDQuant+NVFP4: 4× Smaller, 3× Faster FLUX with 16-bit Quality on Blackwell GPUs
It has already released as in https://github.com/mit-han-lab/nunchaku?tab=readme-ov-file#c....
lmxyy··on SVDQuant+NVFP4: 4× Smaller, 3× Faster FLUX with 16-bit Quality on Blackwell GPUs
I think so. There are already some techniques called rotation, which have similar effects. But it will incur additional overheads in diffusion models.
lmxyy··on SVDQuant+NVFP4: 4× Smaller, 3× Faster FLUX with 16-bit Quality on Blackwell GPUs
FLUX-schnell is only 800ms on RTX 5090.
lmxyy··on SVDQuant+NVFP4: 4× Smaller, 3× Faster FLUX with 16-bit Quality on Blackwell GPUs
SVDQuant supports NVFP4 on NVIDIA Blackwell GPUs with 3× speedup over BF16 and better image quality than INT4. Try our interactive demo below or at https://svdquant.mit.edu/! Our code is all available at https://github.com/mit-han-lab/nunchaku!
lmxyy··on RTX 5090 Workstation Configuration Journey
With the arrival of the RTX 5090, we built a high-performance workstation to maximize its AI computing potential. In this blog post, we share our experience—from overcoming setup challenges to testing its performance.