Neil Ferguson's Imperial model could be the most devastating software mistake
telegraph.co.uk
telegraph.co.uk
I know more and more of us are getting sucked into the 'tinfoil hat zone' with increasing rapidity thanks to the internet - and you could competently argue I'm going there with this - but that sounds like it was done to maybe downplay fears and the fallout we've seen over the past few weeks around this code and its outsized impact on our current position.
It was a 50k LoC monstrosity at one point - in a single file.
Neil Ferguson was put in the spotlight as the "source of truth" by the UK media - I get that its more nuanced and I hope the UK government wouldn't have based its decisions to lockdown on this, and they instead took a wide range of data points - but the media pounced on Neils model being the thing that scared Johnson shitless into rolling back the idea of herd immunity.
So then people wanted to see the model, and saw it turned out to be horrific in its quality, its lack of testing, and its potentiality to contain mistakes (which it did - given Neil rolled back his predictions after a while) I think that probably spooked some high profile people into action.
The whole thing is weird and Neils reputation is in tatters in the media over the way this thing has played out (as well as for some of his extra-curricular activities).
Can you please source this? Sorry to have to ask, but last time someone said this it turned out they were comparing predictions in dissimilar scenarios from two different countries.
Yeah, I read different stories in different newspapers. Such as the rest of the EU, er, politely insisting that the UK was doing something really stupid and needed to change course immediately, not giving Johnson much of a choice. I strongly doubt that the entire EU had based their models on those of a single UK research team.
I would also be strongly surprised if this was the only study used by, well, most of the world to determine their policies. Even if one study isn't as well-{architected, tested} as a computer engineer would hope (hint: they never are), the fact that pretty much all models/studies lead to the same overall conclusion is quite encouraging.
https://github.com/mrc-ide/covid-sim/issues/116
The discussion isn't terribly clear to me. Is this non-determinism in which random number gets used where, like if you seed an RNG once and then use it from multiple threads that get scheduled in varying order? That wouldn't be good, but the model would still be performing as intended, just with no way to make it deterministic for testing. Or is it some other kind of non-determinism?
I'm not sure anyone knows, which indeed isn't totally comforting. I'm not sure the "bug uncertainty" is any bigger than all the existing known uncertainty about the assumptions that go in to his (or any such) model, though.
Per my other comment, it looks like Imperial tried to do that here, but it's not quite working. They seem to think the model is fine and just the "deterministic from seed" mode is broken, but hard to be sure until they fix it.
Based on the GitHub thread you linked to earlier, there are two issues:
1) multithreading causes problems. Quoting Carmack at https://twitter.com/ID_AA_Carmack/status/1258192134752145412 , "issues with random seeds and parallel floating point reductions are extremely widespread, and a great deal of science gets done with non-deterministic simulations."
Think of it this way. IEEE 754 arithmetic is not commutative. Not even addition. (A+B)+C might not be equal to A+(B+C).
Now what happens if you split A+B+C+D+E+F+G+H+I across three threads. The first compute A+B+C, the second D+E+F, the third G+H+I. The threads finish in the order 1, 3, 2, causing the reduce step to be (A+B+C)+(G+H+I)+(D+E+F).
That can very well be different from the single-threaded, left-right addition of those numbers.
Hence the reply "We are aware of some small non-determinisms when using multiple threads to set up the network of people and places. (Look for the omp critical pragmas in the code). This has historically been considered acceptable because of the general stochastic nature of the model. We do expect the code to be deterministic when using a single-thread."
"omp" here is "OpenMP", a low-level set of compiler directives for multi-threaded programming.
2) The followup reports an reproducibility issue with single-threaded code, which wouldn't be subject to the above problem.
Thing is, the reproducible uses the same RNG seed, first when generating a save file, and the second time when loading a save file; both continue from that point. The reporter showed a small difference in the output between the two.
This means the RNG in the first case may have been used many times during the creation of the initial data, which didn't happen in the second time. The two generators have different internal states at this point, so should not be expected to produce identical results.
Which is why the reply "if the RNG is used in the generation of the network, but not when a cached network is loaded, this might explain the discrepancy. The random sequence in the former case would be advanced relative to the latter."
I agree that the non-determinism may be pseudo-random numbers used in a differing order, whether from differing seeds or from differing scheduling of multiple threads. Ideal practice would be to correct for these--a thread-safe RNG, an unsafe RNG per thread, seed saved in output files for restart, etc. That they didn't do that doesn't mean their model is wrong, just that it's harder to show that it's right. Note that even in your quote above, they hedged with "might".
Lots of equations are quite sensitive. You don't need PDEs. The original Lorenz equations are only three ordinary differential equations - https://en.wikipedia.org/wiki/Lorenz_system#Overview , and that's enough to give chaos.
> Ideal practice would be to correct for these--a thread-safe RNG, an unsafe RNG per thread, seed saved in output files for restart, etc.
If you want to reproduce a simulation (assuming it doesn't have issues like IEEE math ordering), then you need to repeat the entire simulation process.
You can't expect that steps A + save checkpoint + B is the same as load checkpoint + B.
If you want people to be able to reproduce from a checkpoint, then you need to do step A + checkpoint, then exit, the run checkpoint + B. That makes "A + checkpoint" as one reproducible and "checkpoint + B" another.
Perhaps that's what they did? I don't know. I only read that single issue, plus HN commentary. And that issue doesn't appear to highlight anything novel or unexpected for this sort of simulation.
> they hedged with "might"
Yes, and that's the mentality drilled into researchers. Because perhaps that code is wrong.
But the internal RNG state difference seems much more likely to me.
And I'm not sure I'd be totally reassured that the simulation is deterministic when run single-threaded in a single start, if they run it multi-threaded for real? How could you be sure the non-determinism multi-threaded is all from the RNG issue, and not from real bugs that show up only multi-threaded, like race conditions accessing shared memory?
With some effort, it should be possible to make the multi-threaded simulation deterministic too. If you want to run the simulation in multiple starts, then you can make that deterministic by saving the final seed in the intermediate output file and restarting from that. I've seen stuff like that in other non-deterministic code, and I'd find it (a) an unreasonable degree of rigor to expect from random academic code, and (b) a quite reasonable degree of rigor to expect in code whose output helps make decisions determining the safety and freedom of billions of people. But (a) and (b) unexpectedly became one and the same, so here we are.
This issue parallels the foundational story of chaos theory. Quoting https://en.wikipedia.org/wiki/Chaos_theory :
> [Lorenz] wanted to see a sequence of data again, and to save time he started the simulation in the middle of its course. He did this by entering a printout of the data that corresponded to conditions in the middle of the original simulation. To his surprise, the weather the machine began to predict was completely different from the previous calculation. Lorenz tracked this down to the computer printout. The computer worked with 6-digit precision, but the printout rounded variables off to a 3-digit number, so a value like 0.506127 printed as 0.506. This difference is tiny, and the consensus at the time would have been that it should have no practical effect. However, Lorenz discovered that small changes in initial conditions produced large changes in long-term outcome.
At the very least it shows that it's easy for intuition to be wrong.
Do you accept that "issues with random seeds and parallel floating point reductions are extremely widespread, and a great deal of science gets done with non-deterministic simulations", re-quoting Carmack?
That holds with my own experience.
You ask "How could you be sure the non-determinism multi-threaded is all from the RNG issue, and not from real bugs that show up only multi-threaded, like race conditions accessing shared memory?"
You cannot. Rather, how can you be sure that any program is bug-free?
The question is, can you be confident enough to use it to make policy statements? Eg, people use imperfect models all the time when making laws.
Note that this sensitivity isn't usually just a numerical implementation detail. For example, it's one of the reasons why the weather is so hard to predict, because the physical effects that create the weather are extremely sensitive to the real physical initial conditions, which we can measure only with some uncertainty, typically much worse than our numerical precision. So when such behavior is present, modelers are usually aware that it is, since it has consequences well beyond floating point reductions. I haven't seen that discussed in any of Ferguson's papers, so I strongly suspect that floating point issues may prevent bit-exact results but do not explain the variation seen.
But in any case, I think we agree that's probably not the biggest difference, and the biggest difference is probably random numbers in the wrong order? I typically attempt to determine whether a program is bug-free by testing it. That's harder when a program is non-deterministic, so people often include the ability to derive all pseudo-randomness from a fixed seed. It looks like Imperial College tried to do that here, but it didn't work. I agree that's not a flaw in the model, just something that makes it harder to prove that no flaws exists. Would you at least agree that in a perfect world, with unlimited resources, that's something you'd fix?
We know that transmissions rates are not fixed. While R0 models the average, sometimes 1 person infects 100 other people, and sometimes no one is infected.
These 100+ events are rare, but not extraordinarily rare.
If it happens early in a simulation, then your simulated model will result in a higher growth than one where that event occurs later.
So sometimes you'll have a large outburst early, and sometimes it will be delayed a week or two.
> My superficial understanding of Ferguson's model is that it's an extended SIR model, with different classes of patient with different behaviors, geographical grouping, etc. From my intuition, it doesn't seem like that kind of DE would be too chaotic
The paper "Impact of non-pharmaceutical interventions (NPIs) to reduce COVID- 19 mortality and healthcare demand" appears to be the one which has influenced many. It cites https://www.nature.com/articles/nature04795 as the model used, which cites https://www.nature.com/articles/nature04017 which gives a summary:
"We constructed a spatially explicit simulation of the 85 million people residing in Thailand and in a 100-km wide zone of contiguous neighbouring countries. The model explicitly incorporates households, schools and workplaces, as these are known to be the primary contexts of influenza transmission .. Random contacts in the community associated with day-to-day movement and travel were also modelled."
I expect that sort of agent-based modeling, by nature, to be chaotic, along the lines I outlined above.
You write: "So when such behavior is present, modelers are usually aware that it is, since it has consequences well beyond floating point reductions. I haven't seen that discussed in any of Ferguson's papers"
How many papers of his have you read? The paper I quoted from, https://www.nature.com/articles/nature04017 , includes:
"Multiple assumptions are inevitably made when undertaking preparedness modelling for a future emergent infection. Sensitivity analyses are therefore critical for assessing the robustness of policy conclusions. Here, critical assumptions not already discussed include (1) the ratio of within-place to community transmission, (2) the expected generation time, Tg, of a new pandemic strain (largely determined by the duration of viral shedding and therefore infectiousness), (3) the level of heterogeneity in individual infectiousness (for example, ‘superspreaders’20), (4) antiviral efficacy/take-up, and (5) the sensitivity and specificity of case detection during the control programme."
You rhetorically asked "Would you at least agree that in a perfect world, with unlimited resources, that's something you'd fix?"
That's a bogus argument. If I had that sort of unlimited resources there are many more things I would rather use it on than making this simulation be the way you want it to be.
For one, the noise and induced chaos are certainly far less than which result from the uncertain input values to the program.
And I'm not sure what you intend by those quotes? Sensitivity analysis is a normal thing you do with almost every model. Those sensitivities are usually small enough that numerical error is negligible--your 100:1 example above is such a small share of the dynamic range of typical data types that it wouldn't usually get any special attention. So "sensitivity" usually means the author is talking about the behavior of the ideal model, correctly neglecting numerical effects. In the rare cases where that numerical error isn't negligible, the author would typically remark on that explicitly, and on the additional steps they took to build an accurate model despite that, and on why that extraordinary sensitivity doesn't invalidate their model in the real world (given that real-world measurement uncertainties are usually much bigger than an LSb). That doesn't seem to be the case here.
Chaos is real, and it's cool; but in most (not all, most) practical work, it doesn't turn out to be that important. I first mentioned it only to show I was aware of cases where a single-LSb error could get amplified huge, and didn't think this was one of those rare cases. I agree that floating point issues may result in non-bit-exact results, but no one other than you is suggesting they explain anything close to the variation observed. I think we even agree that variation is probably mostly random numbers in the wrong order? So I'm not sure why we're continuing to discuss chaos.
In my first post on this page, I wrote:
> I'm not sure the "bug uncertainty" is any bigger than all the existing known uncertainty about the assumptions that go in to his (or any such) model, though.
So I think I agree with you? I'm just saying that from my experience, it's normal good practice in stochastic code to add a "deterministic from seed" mode. It's certainly possible to develop a good model without that, in the same way that it's possible to write good software without tests or source control--it's just harder.
shrug Once upon a time, when I did simulation work, chaotic effects were everywhere.
To see if perhaps my views were unusual, I did a search and found https://www.ncbi.nlm.nih.gov/books/NBK305903/ - "Assessing the Use of Agent-Based Models for Tobacco Regulation", "Appendix B - Agent-Based Models for Policy Analysis"
"An agent-based model (ABM) is a computational simulation model of a many-agent system that captures the behaviors of the system’s autonomous agents and their interactions with each other. An ABM is a computational instantiation of a complex adaptive system (CAS). A CAS is a dynamic model that represents individual agents and their collective behavior"
That describes the Imperial College model.
"ABMs are nonlinear models; highly nonlinear is the usual term. Deterministic nonlinear models have three characteristics that make them difficult to use statistically:
" * Sensitive dependence on initial conditions: A map exhibits sensitive dependence on initial conditions at x0 if there is some distance d > 0 such that no matter how close to x0 another initial point x′0 is chosen, xt and x′t will eventually be at least distance d apart.
" * Complicated limit dynamics: The limit behavior of typical linear dynamic systems is simple: Either the system converges to a steady state from any initial conditions, or they blow up, diverging to infinity. Nonlinear dynamical systems have much more complicated dynamics, including stable limit cycles of various periodicities, strange attractors, and chaotic behavior.
" * Sensitive dependence on parameters: Small changes in parameter values can lead to abrupt changes in the qualitative character of system dynamics.
This is in alignment with my understanding of the issues. These are well-known, and I have difficulty understanding your background enough to not have faced them before. Perhaps since your intuition is based on experiences which are distant from this sort of modeling, you can't use your experience as a guideline?
You wrote:
> it's normal good practice in stochastic code to add a "deterministic from seed" mode
There's nothing in the bug report that you linked to which suggests that the result is not "deterministic from seed" on single-threaded code, so long as you repeat the re-run the code exactly.
The reported non-determinism was not a result of re-running the code exactly, but of running it in a slightly different state.
With this sort of model, "Small changes in parameter values can lead to abrupt changes in the qualitative character of system dynamics." Which is reasonable because in real-life we also expect that small changes in how things turn out may lead to 'abrupt changes in the qualitative character of system dynamics'.
The stuff you're quoting above is pretty standard description of dynamical systems in general. It's true, but I don't see how it applies to Ferguson's model with respect to numerical error; and if it does, then I don't see how his model is useful. I could believe his model is chaotic in the sense that the outputs vary dramatically over the expected uncertainties in the inputs, but the numerical inaccuracy is many orders of magnitude below that.
Finally, I keep saying their "deterministic from seed" mode was broken multi-threaded or after restarts, and you keep saying it was fine single-threaded in a single start. I think both are true, so I'm not sure what we're arguing about?
Yes, with the modification of the statement to "I would not be surprised if it changes by many percentage points."
The bug report you linked to says only "This mostly manifests as a time- offset, but there are numerical changes beyond that." It shows a different by day 80 of 100,000 deaths, or 20% overall. It does not show if by day (say) 120 those numbers have gotten further together or closer.
If they both end at about the same numbers, separated only by a week, then it does not change the overall conclusions in the paper at https://www.imperial.ac.uk/media/imperial-college/medicine/s... . It does shift some of the dates, and I am annoyed with the lack of error bars or the like in their plots.
You asked "how could that model possibly be useful, given that none of the inputs are known to anywhere near that?" You check it for sensitivity across a range of values and see how the overall conclusions change.
Quoting that paper I linked to, "Such policies are robust to uncertainty in both the reproduction number, R0(Table 4) and in the severity of the virus ... However, there are very large uncertainties around the transmission of this virus, the likely effectiveness of different policies and the extent to which the population spontaneously adopts risk reducing behaviours. This means it is difficult to be definitive about the likely initial duration of measures which will be required, except that it will be several months."
When all the lines in a spaghetti map show the hurricane bearing down on you, you prepare for bad weather, even though the hurricane track is almost certainly not going to follow any of those paths.
You write "but the numerical inaccuracy is many orders of magnitude below that."
What I know about chaotic systems is that I cannot make that conclusion without much more study about the relevant phase space. If the end result is an attractor, then even intermediate chaotic behavior is not much of an issue, depending of course on the specific details of the policy guidance.
You write "I think both are true, so I'm not sure what we're arguing about?"
My comment started because of your comments at https://news.ycombinator.com/item?id=23213711 where you wrote that discussion on the GitHub comments wasn't "terribly clear" to you. I gave my interpretation. This specific thread started because you wrote "They seem to think the model is fine and just the "deterministic from seed" mode is broken", while I don't see in that bug report where they state the '"deterministic from seed" mode is broken'.
Is the GitHub commentary clear to you now, and if not, what is unclear?
https://github.com/mrc-ide/covid-sim/pull/121
Linked directly from the issue, too. It would indeed seem to have all been random numbers in the wrong order.
There is no way to independently verify the claims that were made on March 16th. Indeed, the claims originally made were not peer reviewed.
Code was finally put into github, but that is not the code that existed originally.
Important policy decisions require transparency and accountability. Not making source code and data available when it is so easy to do so defeats those goals.
Related commentary about another Telegraph article along the same lines (6 comments) at https://news.ycombinator.com/item?id=23208113 .
This has had a lot of attention on HN recently: https://news.ycombinator.com/item?id=23139508 , https://news.ycombinator.com/item?id=23099499, https://news.ycombinator.com/item?id=23101077 as examples with comments.
the last time this came up it was firmly in the former camp, and was pretty small beer as a result.
For the U.S., the Imperial College team predicted 2.6M deaths if nobody did anything, 474k deaths if full suppression measures were put in place at 1.6 deaths / 100k population / week, and 84k deaths in the U.S. if full suppression measures were put in place at 0.2 deaths / 100k population / week. (from https://www.imperial.ac.uk/mrc-global-infectious-disease-ana...). A range of 84k to 474k deaths seems pretty reasonable given the U.S. is at something like 90k deaths today.
Man, I should have chosen another career area.