I've never done workload optimisation myself, but have looked into it just-in-case. It's a really interesting field and unfortunately, you need to know about both your compiler and target CPU/GPU to get the big gains. I believe most modern CPUs have registers for cache use. You have to do some guesswork to get IPC numbers, since core frequency can vary.
Increasing instructions-per-clock is all about minimising program branching (essentially 'if' statements). Because a CPU core can execute instructions faster than main memory can fetch em. It's a fun game to look at an 'if' statement and figure out how you could instead make it an arithmetic operation :).
Maximising cache hits is all about how you structure and access data. For example, if a cache entry is n bytes long, you want to ensure your struct is smaller than n bytes. Having very consistent access patterns can also help (e.g. arrays-of-structs vs structs-of-arrays).
This field is super deep, it's very fun to learn about!