[1] https://en.wikipedia.org/wiki/Pisces%E2%80%93Cetus_Superclus...
[2] https://en.wikipedia.org/wiki/Galaxy_filament
[3] http://spiff.rit.edu/classes/ladder/lectures/wars/wars.html
Even if there's a sampling thing going on, why should this particular sampling method have an asymmetry?
Maybe there's no reason, or maybe there is a reason.
Just make the usual assumption: The data consists of random variables that are independent and have the same distribution. Or "independent and identically distributed", i.e., the i.i.d. case.
Then can apply the theorems called the law of large numbers, either the weak version or the strong version.
Then in the application of the law of large numbers, the number of samples is crucial, but, possibly surprisingly, the number of cases in all the data does not get involved, is not used in the law of large numbers, and is irrelevant. That is, to be more explicit, with this approach, we don't care how many galaxies there are, 10^12, 10^100, 10^1000, whatever.
Right: With different approaches, might care a lot about (a) the number of galaxies there were in total, (b) the number of galaxies used in the analysis, i.e., if in (b) we had enough samples to characterize (a), to have a good approximation of the results we would get if we analyzed all of (a).
For more, might claim that (A) the distribution we use in the law of large number is just that of ALL of the galaxies, a large number but not infinite (commonly a distribution is for infinitely many) and (B) drew the sample to be analyzed by rolling dice or some other mechanism that achieved independence so that, net, we have i.i.d. as wanted for the law of large numbers. So, here the distribution is not of some imaginary collection of infinitely many galaxies but just of the large but finitely many galaxies there are in THIS universe or THIS observable universe, or ..., this universe once we get tired of collecting data, start to run out of grant money, want to get home with the kids, want to watch the soccer matches, ...!
For one thing, these galaxies were observed and cataloged - something that already puts them on unequal footing with most of the universe' galaxies. Furthermore, of the catalogued galaxies, how were these particular galaxies selected for inclusion in the study? Humans are very bad at choosing things randomly - even when assisted by computers.
For such things will want to learn about the careful definition of independence and, for more, about conditional independence and the Radon-Nikodym theorem. There is a nice proof of this result by von Neumann in, e.g., the W. Rudin Real and Complex Analysis.
> But these are _not_ independent datum
You have not shown that.
> ... unequal footing
Your "unequal footing" is not part of probability theory.
Pick a galaxy X. Note what you know about it.
Now someone tells you it is in the catalog and also lets you see the other galaxies in the catalog.
Now what more do you know about X than you did before????
Can read more about independence in
Jacques Neveu, {\it Mathematical Foundations of the Calculus of Probability,\/}
After all, there are galactic forces at work here (gravity, initial conditions etc). How do we know that everything that’s observable to us isn’t biased? Ie the galaxies that are asymmetric are more easily observed for some secondary reason (closer to us, brighter to us, etc)? Even if the galaxies in our observable universe are independent variables, how do we know that our observable universe is an independent variable of all observable universes? Eg what if the structure of the Big Bang sent different asymmetries in every direction and thus each observable universe will always have some asymmetry from the perspective of any single observer but overall there’s no asymmetry.
Or heck. The ages of the galaxies are presumably hard to account for individually in this kind of combinatorial analysis. How do we know that the asymmetry isn’t because the galaxies are all in different stages of development / the light from them took variable amounts of time to reach us?
Simple. I never claimed that the random variables were independent. Instead I just observed that if one made an ASSUMPTION of independence (and the rest of i.i.d.) then could use the theorem the law of large numbers and get a result without being concerned with what fraction the sample size was of the whole population.
I didn't even try to argue that there was independence or even approximate independence.
The independence assumption is really common in applications of probability and statistics. Maybe then people comfort themselves by believing that independence holds approximately and close enough. As I recall, there has been at least one attempt to investigate approximate independence, but that work is likely not well known.
Questioning independence in this discussion of cosmology is appropriate.
Again, in practice, an independence assumption is common, and with all of "i.i.d." could apply the strong law of large numbers and get results about a large population from a comparably small sample thus responding to the stated concern I was responding to:
> ... a frighteningly small sample to draw conclusions from?
But that's still an assumption, which means that if all your samples come from the local cluster, you cannot extrapolate from your measurements to the entire universe -- because you haven't ruled out that the asymmetry is influenced/controlled by the layout of the local cluster.
Gee, I seem to remember writing:
> Just make the usual assumption: The data consists of random variables that are independent and have the same distribution. Or "independent and identically distributed", i.e., the i.i.d. case.
So you used the word assumption. I used he word assumption. Gee ....
Quite broadly in practical statistics and applied probability, "usual assumption" is correct, that is, that the assumption is usual, that is, usually made, is correct, not that the assumption itself is correct.
E.g., the strong law of large numbers is commonly applied, and it has an independence assumption. Yup, the weak law of large numbers needs only an uncorrelated assumption which is weaker and implied by independence.
Where did I claim that the independence assumption is always correct?
What are you arguing about?????
I used to be a professor. I didn't like it, thought it was not very productive for anyone and financially irresponsible for me.
But I studied probability from some of the best profs and references and wrote a dissertation on stochastic optimal control. If I were to get serious here, I'd have to teach a course, starting with sigma algebras, careful definitions of random variables, various cases of convergence of random variables, ..., conditioning, conditional independence, the Radon-Nikodym theorem, the Markov assumption, martingales, etc. Instead, there are books by Loeve, Breiman, Neveu, ....
I am doing a startup and don't want to be a professor.
I merely pointed out that your answer did not address that part of the question.
> ... a frighteningly small sample to draw conclusions from?
That's all I responded to.
No way did I try to answer all possible probability and statistical questions in astronomy and astrophysics back to the Big Bang, hope for a Nobel Prize, apply for a chaired professorship, .... Again, I just responded to
> ... a frighteningly small sample to draw conclusions from?
In really simple terms, with common assumptions, the accuracy of a sample mean is JUST from the size of the sample and has nothing to do with what fraction the sample is of some population the sample was drawn from.
This point seems to be worth making. E.g., once while working myself and my wife through our Ph.D. degrees, I was asked to estimate the survivability of the US SSBN fleet under a special scenario. From some WWII era Koopman work on encounter rates, and more, I found a continuous time, discrete state space Markov process and generated sample paths via Monte Carlo techniques. Actually used the random number generator that passed the Fourier test and was published by a group from Oak Ridge and based on the recurrence
X(n+1) = X(n)*5^15 + 1 mod 2^47
So I wrote some code, generated sample paths, and found event by event, essentially hour by hour, the expected number of SSBNs surviving. Right, the scenario included lots of weapon types so had lots of combinations of what was surviving.
It happened, nearly uniquely, that my work right away got a little review from a well known probabilist. He asked:
"How can your scenario generation fathom the enormous state space?"
that is, the number of combinations.
This question is related to
> ... a frighteningly small sample to draw conclusions from?
Soooo, even a well qualified probabilist can ask such a question.
I answered: Pick a point in time, t. There the number of SSBNs surviving is a random variable. It is bounded. The sample paths are independent. So, each sample path gives an independent random variable for time t. So we have an i.i.d. case. So the strong law of large numbers applies. So, run off 500 sample paths, add them up, divide by 500, and get a good estimate of the number of SSBNs surviving within a gnat's ass nearly all the time. The probabilist was socially shocked and offended by my mention of a "gnat's ass".
I went on, intuitively, Monte Carlo "puts the effort where the action is". The probabilist's response was "That's a good way to look at it."
In a sense, the most common sampling situation is a finite sample from an infinite population so that the sample is 0% of the whole population -- and we are back to
> ... a frighteningly small sample to draw conclusions from?
0% "small". So, actually 0% small is common.
So, it is common for people to consider what fraction the sample is of the whole population being sampled. This being the case, in my post here I continued and outlined some ways regard the population as finite. That can take us into sampling from finite populations which can be conceptually tricky -- I want to avoid that stuff, and for something the size of the universe, maybe not for a deck of 52 cards, that is reasonable.
For a simple observation, put 50 independent random variables in a box and they are still independent.
That's all folks!
But you can also consider a qualitative analysis. A standard assumption is that the Big Bang should be largely isotopic (else why not, which is the main point of the article). So where you look shouldn’t matter. Even if orientation were uniform in average (impossible to tell), non-uniform regions tell an interesting story.
Now in reality we know the universe isn’t smoothly uniform: parts of it have curdled into galaxies, stars systems etc. Also there’s the puzzling imbalance between matter and antimatter. Why? This article points out another asymmetry, which could provide a clue.
The probability that you've picked the exact 0.0001% of galaxies that had a specific pattern is vanishingly small.
The Copernican principle is the default, unless we have some reason to doubt it in a particular case.
Sample size in a survey will just give you more or less error. You can calculate exactly the expected error of such sample size.