Thank you!
I’m incredibly skeptical that OpenAI is spinning up custom ASICs for improved inference performance, but they never thought of optimizing KV cache until a tiny Chinese lab did it? Give me a break.
I’m incredibly skeptical that OpenAI is spinning up custom ASICs for improved inference performance, but they never thought of optimizing KV cache until a tiny Chinese lab did it? Give me a break.
timeline suggests not.