Its using vanilla llama-2 from Meta with no fine tuning. The point here is the speed and responsiveness of the underlying HW and SW.
> This is not about the model, it’s about the relative speed improvement from the hardware, with this model as a demo.
To compare apples to apples look at the tokens per second of other systems running Llama 2 70B 4096. We're by far the fastest!