LLMs up to 4x Faster With Latest NVIDIA Drivers on Windows
blogs.nvidia.com
blogs.nvidia.com
1: https://nvidia.custhelp.com/app/answers/detail/a_id/3188/~/w...
Thanks for catching my bullshit, I stand corrected (and should have been wiser).
Maybe someone wrote a script that automatically scrapes the website and downloads the newest driver automatically?
Has a scheduled update-check if you want that too.
https://github.com/ggerganov/llama.cpp/issues/34
If you meant eGPU support, IIRC that is beta for everyone right now.
I mean I guess they go a lot wider on a GPU, but there's no reason an APU couldn't do that too.
Haven’t found any good setup instructions for Linux or my Google skills are failing me.
Prompt: A cat wearing pajamas Model: deliberate_v2 Size: 768x512 Batch Size: 4 Clip Skip: 2 Samler: DPM++ 2S a Karras Steps: 70 Seed: 2024828515
Generation times avg over 5 runs
================================
without TensorRT: 19.2 seconds
with TensorRT: 12.3 seconds
Benchmarks
Batch Size 1 / 2 / 4
=======================================
without TensorRT: 31.11 / 36.23 / 42.34
with TensorRT: 55.32 / 58.06 / 62.27
It's faster, but clearly it's not 4x faster. I suppose they cherrypicked benchmarks against generation techniques not using xformers or SDP Attention. Also, it appears to be limited to a max batch size of 4.