Caching is often not such a bad idea regardless of GC.
Cache effects might play a role here as well. Deallocating memory immediately after use allows you to allocate it again when it is still hot in the cache. So eg if you're iterating a structure and doing lot of interleaved allocations/deallocations they are likely to not cause cache misses (assuming the allocator is not dumb and uses some kind of LRU strategy).
With GC and deferred deallocation, you're moving a lot more data between the main memory and the caches, because memory for sure gets pushed out of cache by the time it is reclaimed. And additionally, GC has to touch quite a big number of objects when tracing (and move unneeded stuff to cache, and push needed stuff out of cache). Memory bandwidth is a scarce resource these days.
If you want to see these effects in extreme, try running a GC based program in an environment that is low on memory but has swap enabled. Hitting the first full GC is basically performance game-over, regardless concurrent or not.
On the other hand, manually managed apps can often deal with big chunks of their heap swapped out without terrible consequences.