If you want to go fast in this space then you need to care about data layout and how the system is structured end-to-end. Calling a function per object is going to hit a wall regardless of how you schedule.
Hundreds/Thousands of updates is not small but it's also not massively impressive either. I've done ~2,800 node scene graphs on underpowered ARM chips back in '09 at 60FPS including rendering. You have to use NEON, and be aware of your caches. No scheduling magic is going to change that unless you're just deferring work which sounds like what may be happening there.
FWIW I've also done this in Java via FlatBuffers(which uses ByteBuffer internally) to keep data coherency when driving animations frames so it doesn't require dropping down to C/C++/Rust(although C#'s value types do make it easier).