I think the 'pelican test' is becoming useless. It's been around long enough that now I'm sure good examples are in the training data, and hell they might even do some hand tuning to make it do a decent job since they know people will ask about it.
But either way, with no real way to visualize the result of the text it starts with - it will always be stabbing in the dark. It can't understand conceptually what any of it should look like and then refine the SVG to improve it gradually. It just throws darts at a wall and hopes it comes out alright.