If its only a statistical fluctuation, the p-value due to sequential sampling follows a markov process (random walk) between zero and one (there is no convergence). So starting from a significant result, the p-value would be just as likely to decrease and become "more significant" as increase and "lose significance". The significance cutoff is irrelevant to this behavior, although the random walk will spend less time in these extreme regions to begin with.
Here is an example of the random walk in R. It was my first attempt when going up to 100k samples, after n = 45 the p-value was ~5e-4 (~3.3 sigma), later at n = 36441 it went even more extreme in the other direction, corresponding to a p-value of ~1.9e-6 (~4.65 sigma). Here is the plot and code: https://image.ibb.co/hOHO08/p_markov_chain.png
set.seed(1234)
n = 1e5
a = rnorm(n)
b = rnorm(n)
p = sapply(10:length(a), function(i) t.test(a[1:i], b[1:i])$p.value)
plot(p, type = "l", panel.first = grid())
> min(p)
[1] 0.0005394751
> 1 -max(p)
[1] 1.908121e-06
Perhaps I'm wrong and if you run enough of these and subset to only look after a certain threshold has been breached (ie 5e-2 or 3e-7) there will be different behavior, I doubt it though. However, that is all unimportant since it is based on a false premise:>"_if_ what you're seeing is a statistical fluctuating"
You aren't you are seeing that, so everything that follows is irrelevant. The background model is imperfect (eg nobody believes in a standard model with no higgs boson, everyone knows it is wrong somehow). In this (realistic) case convergence to statistical significance is guaranteed with enough data. However, in the rare (non-existent?) case where the background model is thought to be perfect, it is a different story.
>"Confidence in rejecting the null isn't "somehow translated into support for the researchers' favorite theory", it's translated into support for the existence of a signal, which makes sense since the null is the background-only hypothesis."
This amounts to claiming that in most peoples minds (including the authors of the paper) the Higgs boson, gravitational waves, etc haven't been "detected with 5 sigma confidence"... How much would you want to bet on this? The null model could be rejected due to a loose cable (ftl neutrinos), but the p-value doesn't distinguish between these explanations. Perhaps there was "a loose cable" in the LHC experiments but nobody looked very hard since the deviation agreed with preconceived notions? The p-value doesn't help you here.
>"After we establish rudimentary support for a model by finding one or more such signals, we can then embark on the process of doing precision measurement to really test that model (as is being done for the SM)."
Why not just skip the first step? Instead collect data to check whatever models that exist. When none of them can explain it, modify or come up with new models until it fits. Then derive new predictions from those and collect new data to check them, repeat.
EDIT: I ran it again and got this: https://image.ibb.co/g9rGL8/p_markov_chain2.png
> min(p)
[1] 0.005265033
> 1-max(p)
[1] 1.453814e-05