After analyzed the data, I found branch stall cycles and data access stall cycles were causing huge number of delays in the critical code path.
I used the following tricks to get rid of the stall cycles.
1) Use branch_likely to force gcc to make sure there is no branch at all in the critical path of executions. (save 30+ cpu cycle per branch, there are a lot of branch stall cycles if one just simplely follow the gcc generated "optimized" code. MIPS CPU 200Mhz)
2) Use prefetch ahead of data structure access to get rid of un-cache data delay. (save ~50+cpu cycle per data stall, also, there are lot of them in the critical path.)
3) Use inline functions, etc to get rid of call stalls in critical path.
The system got ~100x increase on the overall system thru-put with those techniques with just pure C optimization from standard -O2 build.
I think it might be possible to create a build system that can automatically collect the profiling data (branch stall cycles and data stall cycles) and use the branch likely and prefetch instructions to auto-optimized the critical path code.
Specifying which code path / function call sequences are the real critical path probably still require programmer's touch.
As result of using data prefetch code in proper place, I don't used cache locking nor doing any kind of CPU affinity trick to generated the optimized obj code without any stall cycles for critical code path.