> however they do not show that this approach produces similar accuracy as large LLMs.
I think they have demonstrated their case pretty well, unless there is some serious degradation of the scaling - 7b is pretty big.
I think they have demonstrated their case pretty well, unless there is some serious degradation of the scaling - 7b is pretty big.
[0] https://twitter.com/gordic_aleksa/status/1682479676910870529