It is going to be difficult to achieve a hard 350 us minimum latency requirement on a general-purpose OS with a virtual memory system, regardless of the allocator you are using.
Hard realtime linux systems do that as a basic given, and all real-time code I've ever seen was C, C++ or SmallTalk, and effectively bypasses the kernel entirely.