The biggest per core "efficiency" improvement will come from RDMA, since IO tends to be the largest bottleneck in Redis (and now Valkey). However, if you are running with deep pipelines (sending multiple commands as a batch without waiting for responses from each commands) than the benefits of RDMA are pretty limited since the bottleneck is on command processing. Multithreading doesn't help much either.
One of the strategies that is being used in the multi-threading is using CPU memory prefetching to pull memory closing to the CPU so that we aren't stalling on fetching data from main memory while executing commands. We still want to try to apply these techniques without multi-threading to improve the efficiency of single or double core installations.