According to my understanding the blog post, FreeWilly2 performs near or above ChatGPT4 for most test-cases. Is this true?
Am I misunderstanding this? Is this not a big deal?
Am I misunderstanding this? Is this not a big deal?
Versions being worked on now will do much better.
GPT 4 is far better and will likely not be beaten by any current open models and approaches but maybe an ensemble of them.
All I see is "compares favorably with GPT-3.5 for some tasks".