Your impact is measured by raw performance masked by time in market. Pure Rust or whatever will have higher raw performance but lower overall (my guess) impact because it misses out on a wider market. Python API will have in general a very slightly less raw performance but much wider time in market.
edit: you've removed a line while I was resonding like: "maybe go is a sweet-spot"
I think if you are looking to replace python as a runtime you'd maybe be better off arguing for safety, python is much more easily corruptible so if maybe you don't trust your machine(s) you are training on and don't want to be mislead by someone hacking your initializations then arguably running a compiled program with certain guarantees is safer.
D'oh! Yea that changes things. I would be considering UI integration from the inference side.
“Guardrails” are often just calls to one or more (usually smaller) classification/moderation models.
also business rules, no?
Right, neither of those are CPU intensive in a different way than LLM inference itself (the latter is LLM inference itself.)
> also business rules, no?
Business rules can vary quite a bit in content and complexity, but either tend to be simple enough that they won’t impose much additional load, or complex enough that you are probably going to want to simply use an existing rules engine (many of which, regardless of their implementation language, have Python bindings) which are going to behave the same way no matter what language you call them from.