It’s a backtracking search program looking for integer solutions to a specific problem. I’ve tuned a series of bloom filters to fill my L2 cache so that I rarely have to touch main memory (this alone took my IPC from 0.1-0.3 to 3.0 per thread). Without SMT it’s 4.6 IPC/core.
I think it’s only able to exceed 4.0/thread with SMT off because of a uops cache? From what I’ve read the Zen5 front end only had a 4-wide instruction decode per thread.