It's a bit interesting how the open models are able to keep pace with the closed models except whole maintaining a steady following time.
If using Anthropic models for distillation doesn't make them better, why do it?
This entire story makes certain American AI shops look cartoonishly evil and Chinese ones relatively sane. Not only they want to grab without giving anything back, they also want to sabotage everyone else's AI research and do plenty of terrible things like media manipulation on the global scale and getting in bed with the government. This can't possibly end well, for the Americans in the first place.
The interesting question is: will Anthropic release a Fable like model with an architecture similar to Kimi, and get the inference cost gains? They should surely beat Kimi because they can internally distill as much as they want.
I don’t feel sorry for the model companies
Yeah.
LLMs
It's hard to feel sorry about the breach of their ToS, which ultimately is all Anthropic can argue, when they are constantly being sued by countless IP owners