I should try it, I guess - I always default to Flash 3.8.
I should try it, I guess - I always default to Flash 3.8.
Specifically, 3.8 Flash suffers from far more hallucinations than 3.1 Pro and frequently ignores the context between consecutive messages. These two issues are fatal to the chat experience.
Because of this experience, I believe that benchmarks are quite limited to evaluate a model's real-world performance, and most current benchmarks focus on coding performance rather than chatting. At the moment, I have almost no reason to use 3.8 Flash.
Also, 3.8 Flash for coding... I've done several experiments with the same prompt and same codebase against other models, and it's never turned out well for Flash against GPT Sol or Astra. No reason to use it for coding, either, except for quick scripts and prototypes, where it's nice because it's so fast.