Q: From my (very, very limited!) understanding, I’m under the impression that part of the “inference engine inertia” is that model- or at least architecture-specific code is required for most, if not each new open-weight model coming out.
Assuming I got that right, do you plan on supporting everything vLLM/llama.cpp can do, such that Magnitude becomes a drop-in replacement for as many (economically/pareto-viable) models as possible, or do you want to focus on the best possible support for only a select few models/classes of models?