Energy-Efficient Llama 2 Inference on FPGAs via High Level Synthesis
arxiv.org
arxiv.org
A great example of the unexpected things that happen when put great code into the commons.
I bet Andrej never expected anything like this when he released it.
So as far as I can understand, the biggest "bottleneck"/limiting factor with using FPGAs for LLMs is the available memory -- with current large models exceeding 40 GiB in parameter size, GPUs and TPUs with DRAM look like the only way to go forward for the months to come ... Thoughts?
Up to 38 TFLOPs of FP16 performance, up to 116 Gbps transceiver rates, and up to 3.9M logical elements (LE).
Memory bandwidth of over 1TBps using the NoC, using in-package HBM2E (up to 32GB capacity) and hardened DDR5/LPDDR5 memory controller (supporting 5,600 Mbps).
Maybe doesn't come into the power envelope (Agilex 5 is looking good on that front) of a low-power device, but it's an awesome chip. About £8k for a DevKit though.
An interesting twist is that this DRAM might not need to be a central pool where bandwidth must be shared globally -- e.g. the Tensortorrent strategy seems to be aiming for using smaller chips that each have their own memory. Splitting up memory should yield very high aggregate bandwidth even with slower DRAM, which is great as long as they can figure out the cross-chip data flow to avoid networking bottlenecks
Thus the GPU is storing a 110M model in gigabytes of external RAM, and paying the power penalty associated with the excess capacity, while the chosen model 110M fits neatly within the FPGA's on chip RAM, and the design can trim all that overhead accordingly.
A more fair comparison would either run a larger model that had both systems hitting external RAM, or they would compare power/performance against some sort of inference ASIC that had all the RAM on chip (maybe a cerebrus, but scaled according to the portion of the wafer actually used for the model).
That being said, it's neat that they open sourced their work, and it's worth looking down this path more.
A more apt comparison would have been with a phone made in the past 5 years, even without an AI accelerator chip I'm sure you could manage 20-30+ t/s from a 110m model but this depends entirely on the memory bandwidth of the phone.
The problem is FPGAs are heterogeneous and highly optimized to reduce latency, instead of for efficiency, so there are strong limits to the approach.
On the plus side, the LLMs are a lot of layers, so you could take one layer per FPGA, and just use the high speed links in the chips to feed the results to the next FPGA.
You'd have maybe even a millisecond of latency, at a clock rate of 100 Mhz, but possibly a million tokens per second of serial / parallel streams of execution.
The part of this ecosystem that is as non-democratized as it can be is training. It's currently impossible to train decent enough model with resources that are available to one person.
And from a hardware cost perspective the AWS f1.2xlarge instances they used are $1.65/hr on-demand, vs say $1.29/hr for an A100 from Lambda Labs. A very interesting line of thinking to use FPGAs, but I'm not sure if this is really describing a viable competitor to GPUs even for inference-only scenarios.
AWS instance prices are more of a supply/demand/availability thing, it would be more interesting to compare from a total cost of ownership / perf-power-area prespective.
Is there something I am missing making FPGA potentially more viable, besides not feeding into NVIDIA’s greed?
The reason FPGAs work as an in-between product is because they are also a commodity product. All you need is a team to program the device. You can buy a million of them off the shelf right now, and there's a hundred other applications they're good for so you're not restricted when the time comes that your ASIC is no longer needed for whatever task it's explicitly designed for.
anyway if you're tempted by this, i strongly advise you to ponder this:
> Run the Hardware build, should take around ~12 hours.
> am i missing something? since when does vitis connect to vivado and do the p&r too?
I haven't done much HLS, but isn't that the normal case? Translating the HLS into HDL and then do the pnr with vivado?
> Translating the HLS into HDL and then do the pnr with vivado?
As far as I know, vitis does not give you tcl scripts (or whatever) for vivado - you have to do that yourself.
[link to example HLS makefile](https://github.com/Xilinx/Vitis_Accel_Examples/blob/f61637e9...)
Memory bandwidth for FPGAs seems worse, so for serving models don't GPUs still win out?
the Xilinx Virtex UltraScale+ VU9P FPGA prototyping boards seem to be 9000 USD. Anything in the 1000$ range ?
fair warning: one does not "play around" with an FPGA. they are the antithesis of user friendly.
As for cheaper FPGAs, the paper notes that the bottleneck is the size of on-chip memory. So I doubt it will be easy to find cheaper model to reduce costs.
Another hidden fee would be Vivado and Vitis (tooling) licenses, which you need for most upper-end FPGAs.