Let’s ask in good faith. Can you suggest something that it can’t do? Functional things. I’ll reply in good faith and consider it.
Then you’re either going to say it can or you’re going to say that requires more than 10000 tokens.
This isn’t an interesting conversation and I don’t think you are presenting this challenge in good faith for the reason I gave above.
There are several models with greater than 1200 elo
at the beginning 2022 it was useless because the output was garbage (hallucinations and fake data).
nowadays its still useless, but for different reasons. it just regurgitates things already known and published and is unable to come up with novel hypotheses and mechanisms and how to test them. which makes sense, for how i understand LLMs operate.