[1] GPT-4 Technical Report, https://cdn.openai.com/papers/gpt-4.pdf
[1] GPT-4 Technical Report, https://cdn.openai.com/papers/gpt-4.pdf
They can exclaim the model says 40% less "xbox live gamer words" which people outside the company couldn't validate.
tl:dr OpenAi is now a business
Worth watch Yannic talk about the problem and other cool ML topics too. https://www.youtube.com/watch?v=2zW33LfffPc
Here we'd need to see more about its design and safety, else you may be getting recipes for veggie dishes when what you really wanted was fried chicken.
I’m no LLM expert, but I don’t think you can eyeball the arch and say “that’s going to confuse veggies for fried chicken”.
"GPT-4 and professional benchmarks: the wrong answer to the wrong question OpenAI may have tested on the training data. Besides, human benchmarks are meaningless for bots."
Having multiple tests would be a stronger test say with example prompts: "whats going on in this picture", "what would a person think seeing this image" etc..
gpt4 is cool as a numbers box but this is not reasoning logic and without papers hasn't been proven either.