I think even easier in fact - what's happening behind the scenes w/ an LLM is far more opaque
53 karma · joined May 1, 2017
We do some on-the-fly optimizations as well (like compiling into CUDA graphs or fusing together calls) which ends up resulting (for some inference engines) faster token throughput too.
In fact one of our customer's use cases is exactly what you describe, allowing users to "hibernate" container workspaces.
Pretty wild first experience doing it for me - I'm used to the hassle of sideloading a bootloader and then flashing.
Thanks :)
It's not unheard of for them to employ people who's primary role is grant writing to try and get (for example) SBIR funding
Are all your roles subject to ITAR restrictions?