Happy to see Chinese OSS models keep getting better and cheaper. It also comes with a 50% API price drop for an already cheap model, now at:
$0.28/M Input ($0.028/M cache hit) > $0.42/M Output
$0.28/M Input ($0.028/M cache hit) > $0.42/M Output
(The inference costs are cheaper for them now as context grows because of the Sparse attention mechanism)
Output: $1.68 per million tokens.