I may be in agreement, and I was an idiot to misunderstand your comment and reply based on it.
I especially agree with the last sentence that models largely learn features in the dataset, but I don't understand why you would describe it as
> There's nobody that's like "today we're working on fingernails" or "today we're making hair physics work better"
If there were a business case for that, I would characterize curating a dataset of fingernails and fine-tuning or augmentind a model based on that as "today we're working on fingernails".
And the same too with eye reflections. So with the right dataset you can get eye reflections right, albeit in a limited domain. (E.g. deepfakes in a similar setting as the training data). In fact you can look at the community that sprung up around SD 1.5 (?) that fine tunes SD with relevant datasets to improve its abilities in exactly a "today we're going to improve its ability to produce these faces" kind of fashion.
Where did I misunderstand your comment? I seem to arrive at the completely opposite response from the same fact.
I also noticed that you say
> the improvements are really just giving the models techniques for more accurately replicating features from reasoning data.
You seem to refer to aspects of a model unrelated to dataset quality. But fine-tuning on a curated dataset may be sufficient and necessary for improving eye reflections and fingernails.