The drive so far has been to scale up to push performance, not to reduce cost. Although everyone offers cheaper versions of models they don't base their marketing on these. Now that the momentum of scaling up on training is slowing down they are pushing inference time computes and efficiency comes to the forefront.