I'll try it again now that it's out of preview and has been updated with more post-training. It presumably can't be worse, so maybe it's better enough to compete with a 31b model.
I haven't tested it yet so I cannot comment on the quality (nor the comparison with 10x smaller (!) models)
Of course the bigger model embeds more knowledge, but when neither model has the knowledge necessary to perform the task, hy3 makes idiotic decisions all the time whereas gemma 31b has a decent hit rate.
hy3 feels like someone who's read a lot of books and says the right words but has nothing of substance between their ears, gemma feels like a reasonably intelligent person who doesn't understand the domain, the latter is muuuch easier to work with than the former.
I've only used the Hy3 preview, so I don't want to judge too harshly, yet. But, I wasn't very impressed with it a couple of months ago.