We saw substantial performance improvements from the increased register count of AMD64 when I lead the porting effort of a video editing/compositing suite back in the early days.
Because our engine took advantage of memory-mapped unnamed file handles to cache frames, we didn't really need the extra address space of larger pointers. I was able to exhaust system memory by manually managing whether or not buffers were mapped into our address space.
Though it had been made for something else, I lucked into being able to use that same engine for interprocess legacy plug-in support. AMD64 support also wasn't necessary to get more physical memory support out of the chips of the era. Operating systems supported 36 and 40 bit physical address spaces, and applications could use higher (would-be negative) addresses by indicating support for it. Those wanting to really cut loose could turn to manual mapping. (See: PAE, LAA)
In the end, we shipped AMD64 support because there were substantial performance wins. And, yes, I profiled every possible software/hardware combination to determine where gains and losses were coming from.
I'm not saying that things were universally faster/better. RIP-relative addressing, though generally quicker than absolute 64-bit, is a headache, and the loss of 80-bit floating point intermediates broke some rare code (which would have broken in strict mode, anyway). I'd love to see some evidence of slow-down now, but, until then, I'm unconvinced. I had a dev bring me "evidence" of huge performance differences between ARMv7 and ARMv8 a couple of years ago (and there are some if you dig). I asked him if he had compiled with or without thumb in the linked modules.
He let me know of his new results a few days later, a little red-faced.
I'm not saying that your detection of slowness is the same thing, but many aspects of performance are measurable (and measured). If it's slower, we can probably measure just how much.