Thanks for pointing this out. He has his own set of candidate submissions to the NIST post quantum competition, so throwing shade at other candidates makes a bit more sense now.
I've done some analysis on one of his candidates and it was by far the slowest on cores which do not feature wide vector operations. My opinion is that you shouldn't just optimise for Intel (as many submissions do) for all the obvious reasons.
Edit: A quote from the slides by Nigel Smart which Bernstein criticises - "[We] Would caution NIST against putting too much emphasis on academic measures of performance of algorithms for this reason"
- I couldn't agree more.