It depends on your word list, are you using the same word list as wordle?
The assumption you need to make for my analysis to be correct is that the letter patterns in the 2,500 possible answers is statistically similar to the distribution of letter patterns in the original 12,000. There are probably some differences between the distributions, and I'd love to rerun my code with the actual word list Wordle uses, but in the absence of that list, I think that my code does about as good as possible.
[1] uses for the answers; I assume it allows all 12,000 for guesses. [2] NYTimes does not specify which source they used