Full threadeximius·This is "evaluating" LLMs in the sense of benchmarking how good they are, not improving LLM inference in speed or quality, yes?View on HN