Good overview.
At the other extreme, some recent works [1,2] show why it’s sometime better to scale down instead of up, especially for some humanlike capabilities like generalization:
[1] https://direct.mit.edu/tacl/article/doi/10.1162/tacl_a_00489...