Does this correlate with the racial proportion data?
Does this correlate with the racial proportion data?
It is wrong to get overly exact here though. We are talking about a overall score variance of just 10pts in the face of an achievement gap of nearly 30pts, with an absolutely massive 20% decline in white population share over the 26 year period.
I'm in a different field but an arbitrary 27 years (1998-2024) of data and the autocorrelation both would get flagged in a review for me. Not to give statistics homework but you should test and correct the autocorrelation issue you're describing in the model and if the data goes farther back I would go farther back, too (with how you're describing the errors going off pattern then back on it sounds like the inference here would be at least somewhat unstable depending on year chosen).
Edit: these are the assumptions and basics of how to correct for violations, in particular you're describing a violation of assumption 2 but you should test for all of them - https://www.statology.org/linear-regression-assumptions/
There's a channel on YouTube called Statquest that teaches statistics in a pretty accessible way if you're interested in analyzing this sort of data.
To make my point in the original post, I held the group scores fixed and extrapolated combined scores based on changing population sizes. This is straightforward. Simple even.
\(\widehat S=\sum_g p_gS_g,\)
Why did I call this "linear"? (Not a "linear model", your term, because it is not). The model function actually tracks the population function. Over the 25-year period, however, this trend is in fact linear because the White-to-Hispanic population trend is linear on that timescale. Over a longer timescale it would not be.
I don't mean to disparage amateur statistics since often there are interesting things that experts miss that amateurs have insight into but this sort of data with few data points (the data you have doesn't sound independent) and a lot of confounding factors is not trivial to accurately model. Not that everyone needs to learn statistics but this sort of data (and also this topic) probably deserves a bit more of a rigorous approach.
As a very simple example, picking a time range for the model because the model doesn't work when you go farther back adds a lot of potential bias into the model and that along with the non-independence issue (and without looking at the data myself I don't know if there are other issues) are going to lead to overconfidence in conclusions.