HNHacker News
TopNewBestAskShowJobs

oscarrovira

4 karma · joined February 12, 2017

Exploring.
submissionscomments
oscarrovira··on Deploy Gemma 7B with TensorRT-LLM and achieve > 500 tok/s
There’s a percentage of users that do care about token generation speed as they chain multiple API calls. The performance is all thanks to TensorRT-LLM, Mystic takes care of the engineering of getting a scalable endpoint out of it, i.e, not having to manage your infra.