Thanks!
> I always disable C-states deeper than C1E
AWS doesn't let you mess with c-states for instances smaller than a c5.9xlarge[1]. I did actually test it out on a 9xlarge just for kicks, but it didn't make a difference. Once this test starts, all CPUs are 99+% Busy for the duration of the test. I think it would factor in more if there were lots of CPUs, and some were idle during the test.
> Try receive flow steering for a possible boost
I think the stuff I do in the "perfect locality" section[2] (particularly SO_ATTACH_REUSEPORT_CBPF) achieves what receive flow steering would be trying to do, but more efficiently.
> Would also be interesting to discuss the impacts of turning off the xmit queue discipline
Yea, noqueue would definitely be a no-go on a constrained network, but when running the (t)wrk benchmark in the cluster placement group I didn't see any evidence of packet drops or retransmits. Drop only happened with the iperf test.
1. https://docs.aws.amazon.com/AWSEC2/latest/UserGuide/processo...
2. https://talawah.io/blog/extreme-http-performance-tuning-one-...