Does Anthropic do something like this as well, or is there another reason Claude Sonnet 3.5 is so much better at coding than GPT-4o?
Does Anthropic do something like this as well, or is there another reason Claude Sonnet 3.5 is so much better at coding than GPT-4o?
"Which data specifically? Gerstenhaber wouldn’t disclose, but he implied that Claude 3.5 Sonnet draws much of its strength from these training sets."[0]
[0]https://techcrunch.com/2024/06/20/anthropic-claims-its-lates...
I remember when I first started using activation maps when building image classification models and it was like what on earth was I doing before this... just blindly trusting the loss.
How do you discover biases and issues with training data without interpretability?
It's impossible to say because these models are proprietary.
This reminds me of the passage found in the description of the fuckitpy module:
"This module is like violence: if it doesn't work, you just need more of it."