While the tests are about the language, noobermin's comment that I replied to was not.
Here's my summary of the tests:
2.4: generator comprehensions; the obvious way if that's what you need
2.5: if-else expressions; a bit too easily used when if/else statement is more appropriate, but there are times where it's a good fit.
2.7: set notation: {0} is the obvious improvement over set([0])
3.0: ... as a token - I don't really understand why Python changed here. I only use it in indexing. It's not obvious when I would use it at all.
3.1: multiple context managers; the obvious improvement over a multiple indented managers
3.3: yield from list; the obvious improvement over for x in lst: yield lst
3.5: @ operator added for matrix multiplier. Solves a special pain point. Obvious only for that case.
3.6: underscores in numbers. Generally better for larger numbers. Mostly obvious.
3.7: async - I've not done async in Python yet, so can't judge
3.8: x:=expr; the debate that broke van Rossum. Not going there.
3.9: allow more complicated @annotation expressions; more useful than the alternative if that's what you need
3.10: match; no experience with it. I suspect it's more obvious
3.11: more async that I can't judge.