Can you share any more details about the workload? Always interesting to hear of something like that which isn't a useless microbenchmark.
Does it use vector? What can you hit with SMT disabled?
Does it use vector? What can you hit with SMT disabled?
I think it’s only able to exceed 4.0/thread with SMT off because of a uops cache? From what I’ve read the Zen5 front end only had a 4-wide instruction decode per thread.