I want and hope this to succeed. But the tea leaves don't look good at the moment:
- model sizes that the industry was at 2-3 gens ago (llama 3.1 era) - Conspicuous lack of benchmark results in announcements - not on openrouter, no ggufs as yet
- model sizes that the industry was at 2-3 gens ago (llama 3.1 era) - Conspicuous lack of benchmark results in announcements - not on openrouter, no ggufs as yet
quantizations: available now in MLX https://github.com/ml-explore/mlx-lm (gguf coming soon, not trivial due to new architecture)
model sizes: still many good dense models today lie in the range between our small and large chosen sizes
Note that we have a specific focus on multilinguality (over 1000 languages supported), not only on english