Java and C++ concurrent maps scale reads quite well, both in practice and in research. I was comparing only in-memory read-only map lookups, not those that persist like FASTER and MassTree. I should note though that decade or older hardware will have lower multi-threaded throughput by using many slow cores (Azul, Sparc), so those raw numbers are not comparable to today's x86 machines.
In my experience when adding atomic writes to mostly uncontended fields during this lookup, the performance degrades to 33% of the raw map. Therefore 25M/s did not make sense and I believe there is a limiting factor that could be removed to increase throughput.
[1] https://preshing.com/20160201/new-concurrent-hash-maps-for-c...
[2] http://people.csail.mit.edu/shanir/publications/LazySkipList...
[3] https://arxiv.org/pdf/1809.04339.pdf
[4] https://dl.acm.org/citation.cfm?id=3210408&dl=ACM&coll=DL
[5] https://www.usenix.org/legacy/event/atc11/tech/final_files/T...
[6] https://web.stanford.edu/class/ee380/Abstracts/070221_LockFr...