You’re kind of screwed on the latency optimization side bc your bottleneck is a language model though. Why wouldn’t you eventually just use a specialized model that isn’t constrained by running a language model under the hood?
They even named the model after "Jevons Paradoxon" - They anticipate that their model lowers the cost of adopting this kind of AI significantly, unlocking a lot of use cases.