The model F=GMm/r^2, for example, has a causal and ontological semantics: F is a force, M a mass etc. these are pieces of reality. And this formula (though actual a little suspicious in many ways, GR fixes this) nevertheless says there is a force between masses that has certain properties etc.
Now you can say that astrologers who recorded positions of the stars in books helped 'create' this model in the sense that this data was inspiration to newton. But he didnt derive the model from this data: there are an infinite number of (causal) models consistent with the data (statistical models).
Rather newton played around with creating geometries, just like the vase-makers in plato's cave. Newton built various ways the world might be first, projected data out of them, and compared that to 'the statistical data of his day' (ie., astrology).
There's nothing in the data to tell Newton he was right. Indeed, vast amounts of it told him it was wrong: such a law does not describe the known solar system at his time, very far away from it.
Nevertheless 'modelling shadows' isnt science; and his job was science. So one has to compare actual explanatory models, and his was the best.
What you're describing above is hypothesis testing which occurs long after theory building. Broader theories create causal models, causal models create sets of predictions, we call some subset a hypothesis and by hypothesis testing we can select, in an often psuedoscientific way, between causal models.
This technique occurs long after the invention of science, arises out of explanatorily bankrupt areas, as a way of 'giving researchers something to do'. It's wholly pointless without theory-building, it is just averaging shadows.
The science we think of when using the term 'Science' owes very very little to the modern practice of hypothesis testing. Comparing hypotheses is an intellectual part of assessing explanations -- identifiable formal statistical methods entered in the early 20th C.
For almost all of scientific history 'data' functions much more like reductio-ad-absurdum premises in philosophical arguments than as sets of numbers from which to derive distributions.
That latter system, in most cases, fails. It provies a wholly illusory sense that data can decide matters; and applies in cases requiring extreme non-physical assumptions (eg., of the normalcy of the underlying data, or of a fast rate of convergence of the central limit theorem).
Much real-world phenomena studied by stats cannot really be studied by data analysis at all; and the whole method of 20th C. statistical hypothesis testing is the opening sales pitch to entire fields of pseudoscience.