The problem with these benchmarks is that the Chinese models tend to be incredible on paper, and absolutely terrible in practice :/
You probably refer to GLM-4.7
I work on mid-sized projects currently (200k to 1kk lines of code).
Isn't that a million?
When I worked in banking, the codebases were often larger than a million.