Scientists begin building highly accurate digital twin of our planet
ethz.ch
ethz.ch
In particular, this is a notoriously poor approach to modeling complex large-scale interactions between humans and their environments. There was a study I was involved in last year to determine why one of the epidemiological models for COVID was so badly off target. The root cause was modeling human behavior in the same way you would model weather, which is quite inappropriate but the implementors of the epidemiological model did not have the expertise to know better.
The selection of GPUs is also not appropriate for the nominal objectives of the program. When modeled correctly, i.e. not as weather, these kinds of things aren't the kind of workload GPUs are good at. They tend to look more like very high dimensionality sparse graphs -- latency-hiding is more important than computational throughput. CPUs are actually quite good at this.
This looks more like a program more designed to produce press releases than useful results.
https://www.wsws.org/en/articles/2020/03/18/covi-m18.html
https://www.nature.com/articles/d41586-020-01003-6
https://www.imperial.ac.uk/news/196496/coronavirus-pandemic-...
The model authors have since argued that the data was correct, but people responded to the pandemic by changing the way we live. That's OP point: that feedback cycles and corrections exist and they make modeling dynamic systems very difficult.
Even a cursory glance at the actual data, even the data available at the time, shows they were completely and utterly wrong.
"Lots of people could die if you keep behaving as you currently are"
"Okay, lets behave differently"
And then less people die.
Trying to frame it as "they modelled it wrong" is nonsense. What even is the point of predictions like this if not to change behaviour - predicting outcomes based on everyone taking precautions and not telling people what might happen if they don't would be dangerous and irresponsible.
And even their best case scenarios overshot the mark -- and by a lot.
This isn't to criticize modeling -- it's only to point out how hard it is to get right.
> The models tended to overshoot the number of deaths by huge amounts. For example, the Imperial College of London estimated 40m deaths in 2020 instead of the 2m that occurred.
The very article you cited pointed out that the 40m figure was based on a "left unchecked" scenario. It was not an attempt to predict the actual number of deaths that would occur. Claiming that this is indicative of overshooting because the actual number of deaths is 2m is completely wrong.
It's not that. It's that when the system you model responds to the existence of your model, it becomes anti-inductive. It's no longer like weather, but is now like the stock market[0]. Your model suddenly can't predict the system anymore[2], it can at best determine its operating envelope by estimating the degree to which the system can react to the existence of the model.
--
[0] - I use the term anti-inductive per LessWrong nomenclature[1], but I've also been reading "Sapiens" by Yuval Noah Harari, and there he uses terms "first order chaotic" for systems like weather, and "second order chaotic" for systems like the stock market.
[1] - Introduced in https://www.lesswrong.com/posts/h24JGbmweNpWZfBkM/markets-ar....
[2] - I think it becomes uncomputable in Turing sense, but I'm not smart enough to reduce this to the Halting Problem.
https://en.wikipedia.org/wiki/Lucas_critique
and also https://en.wikipedia.org/wiki/Campbell%27s_law
https://en.wikipedia.org/wiki/Goodhart%27s_law
Edit: I think this is not the first time the good people from lesswrong dug up some well known idea and gave it a new name. Good thing, too, giving this important concept more attention. Too often we forget how many people have dealt with the problem of modeling complex systems in the past. And while we can not read everything, it's often a good idea to have at least a glance at where they failed!
Too often I read/review some new "revolutionary" paper based on the idea that hey, we can model this process (involving people) like XYZ from physics, where this stuff works great! Surely, this is better than the plebian approaches in the literature! And then, to the shock of all involved, it doesn't work great....
Also: https://xkcd.com/793/
Were they actually so naïve that their model did not allow for the possibility that human beings change their behavior in fear of death by pandemic?
Of course there are many ways that things went wrong, and not everyone made the same mistakes.
This is the same reason as to why Econophysics, despite all its promise and grandeur, has yet to yield any useful results.
Interdependent human beings are a lot more complex entities than particles or flows.
Isn't it interesting how often one reads about science in the press that seems more geared towards generating press releases than useful results?
This hasn’t always been the case so it’s fair to consider what circumstances might foster a better situation, where research can be directed to areas most promising to add progress and value to society at large.
Consider palæontology as an entire field; there is no financial benefit nor practical application to be had for it, yet it seems to find ways to attract funding all the same, most likely because it does have a habit of generating spectacular news articles which sponsors would probably enjoy the publicity of.
But there is truly no practical benefit for mankind to be had in trying to answer to what extent various dinosaur species were endothermic and feathered.
Except we’re an insatiably curious species and it sates our curiosity.
The image of dinosaurs that became canonically entrenched in popular culture is almost certainly completely wrong, but the truth is of little consequence, exactly because it is not used for anything that might depend on it's veracity.
It really does not matter whether it be accurate or completely false, for this purpose.
> Healthcare is recognized as an industry being disrupted by the digital twin technology.[45][34] The concept of digital twin in the healthcare industry was originally proposed and first used in product or equipment prognostics.[34] With a digital twin, lives can be improved in terms of medical health, sports and education by taking a more data-driven approach to healthcare.
The down side is everything explodes exponentially - setup time, mesh count, solve time; and we usually get worse results than more focused simulations because we can't squeeze enough detail in across the board.
It generally starts because some manager hears that we've created 8 different specialized models of something due to different areas of interest, and has the bright idea of "lets just create a single super-accurate model we can use for everything". I've been fighting against them my entire career, although 10 years ago it was "virtual mockups"
The next buzzword in the pipeline seems to be "virtual lab" which I can't figure out either. I've been simulating laboratory tests for over a decade and no one can explain to me why that isn't exactly what we're already doing.
None of this is to say that this team isn't doing great work, but somewhere along the way it got wrapped up in some marketing nonsense.
Edit: Restructured my reply to better address OPs question.
"Light fields" is one that always annoyed me. People who are apparently unaware of centuries of knowledge and methods in electromagnetism, developing "new" ways to solve problems crudely. That's great if they can make some cool new imaging system, but is it research deserving of long-term high-risk funding? It's just something that anyone skilled in optics can work out if they thought to build it.
Either way I wouldn't think too much about it. Tech is full of these things. I have been working with AI and neural networks for years before it was called Machine Learning. Now I'm forced to use the term ML to sound relevant even though it is the same thing.
So you can have sensor and measurement data from the real thing be streamed to the model in (ideally) real-time, you can make decisions off of the state of the model, and have those decisions be sent back out into the real world to make a change happen.
The specific wording of digital twins originated from a report discussing innovations in manufacturing, but I find that railway systems and operations make for some of the best examples to explain the concept, because they manage a diverse set of physical assets over which they have partial direct control, and apply conceptual processes on top of them.
Here's three assorted writings [1][2][3] that explain how railways would benefit from this.
[1] https://www.anylogic.com/digital-twin-of-rail-network-for-tr...
[2] https://www.railwayage.com/analytics/how-digital-twins-suppo...
[3] https://www.railwayage.com/analytics/realizing-the-potential...
I work as a Computational Researcher at Stanford Med. My work is quite literally translating 3D scans of the eyes (read MRI) into "digital twins" (read FEA Models).
I think that there is a subtlety in differentiating a digital twin from a model/simulation in intent. Our intent is to quite literally figure out how to use the digital twin specifically, NOT the scan that it is based on, as a way to replace more invasive diagnostics.
Of course, in the process, we figure out more about diagnosing medical problems as a function of just the scans themselves too.
Maybe we were in the business of building better business builders, I am still not sure.
I work with digital twins in chemical manufacturing, and there the term is directly coupled with Model Predictive Control. The basic idea is that you build a model of the system (e.g. a chemical plant) you want to control, use that model to optimize controller behavior, apply the results to real controllers in the real system, and then sample the system to reground the model. Rinse, repeat. Such a model is called the "digital twin" of the real system - the idea is that it exists next to the system and is continuously updated to match the real world.
For GE's digital twins in the jet engines, they will build a high fidelity representation of the each individual engine based on as built parts, and then they will simulate every flight based on accelerometers, force sensors, humidity sensors, temperature and pressure sensors which they have placed in the engine. This is different from a general model or simulation which will build a model from CAD and then have a series of expected flight simulations and use that to predict life of the engine.
As an example, my startup, Bractlet (bractlet.com), uses detailed, physics-based energy simulation (aka "digital energy twin") technology as a tool to optimize HVAC design and controls in large commercial buildings. I'm sure efficacy varies widely by domain, but it's worked extremely well for us so far; we typically help our customers save about 20-30% on their energy expenditures annually, and the digital twin is the bedrock of our approach.
- How do model random variables with high variance like occupancy or equipment use? Are you using sensor/monitoring equipment or just trying to model it as best as you can?
- Did you guys use the initial energy model (built by by the design/consultant team) or do you guys build your model from scratch?
- What energy simulation engine do you guys use? I'm just curious if there's any big advantage between the engines for digital twin applications.
- There's almost never a model. We build it from scratch in every case.
- We have a highly customized fork of EnergyPlus. They are all difficult to use in our experience, but EnergyPlus allows us to get the detail required for our needs.
I think I've seen something like it on the latest season of Westworld. Jokes aside this reminds me a little bit of the brain project. I'm really not sure if attempting to build full fidelity models of hugely complex systems is the way forward to understand this systems. It seems increasingly that scientists are trying to replace theoretical models of the say, the mind or the city with purely data driven approaches that don't necessarily produce any insight, or even accurate forecasts. Sometimes it feels like with increasing computing power we've gone back to the problems of empiricism of the mid 20th century.
I foresee a Big Science project with little to no benefit for meteorology, climatology, geology, etc and a lot of public money going into building a really, REALLY big computer.
Which isn't necessarily a bad thing to buy per se, but if the intent really was to understand Earth, they'd be better off putting the money into fundamental research in the Earth sciences.
It may seem like a logical desire to integrate the existing massive data feeds about various aspects of the planet into a progression of coherent states of it. With the hope to gain insights for the dynamics and trends.
The first challenge is to figure out if such feeds are indeed integratable and at all could lead to any sort of coherency.
The next challenge is to understand what should be considered a "state". It is a complex system, yet one needs to choose the parameters.
It is nice to have such a global view of the observations. But I'm sure the goal stretches further and means to produce projections and forecasts. This may have some political impact, but practically may just be as good as a speculation, given the vast scale.
I think they are mesmerizing to look at. https://youtube.com/shorts/sdlx8yxdLPo
But I want to be charitable and so I wonder if there really is a game-changing idea in here that the press release does a poor job communicating.
Could this wind up physically measuring a ton of stuff that hasn't been measured at a decent resolution before, and so produce genuine meaningful improvements?
Or is there really some kind of new viable "supercomputer" architecture unique to climate modeling that will pay off massive dividends?
The AI part worries me the most, since it can be a notorious black box where critical biases and errors get amplified without even being detectable. But are there actually techniques here to drastically speed up the production of expensive calculations, that are cheap to verify as correct?
Or is there genuinely a divide between earth scientists and computer scientists where neither side is benefiting from advances in the other, and there's a huge genuine opportunity here for a massively productive paradigm shift?
I'd really like to hope there's something of value here, and perhaps someone here who has worked with climate modeling and knows actual details about this project has some insight.
The best we will ever be able to create is approximate models that have varying tradeoffs. The term "digital twin" obscures the fact that there are tradeoffs involved in the first place. It also causes harm to decision-makers, who ought to be able to choose which tradeoffs to make.
https://www.euronews.com/living/2020/07/09/living-in-a-bubbl...
Highly sceptical the digital version of this will be playing with a full deck of data resulting in skewed results. still think it's a good idea but definitely not the accursed 'settled science', more an experiment.
Reading the article, it seems that the purpose of the twin isn't simulation, it's mapping. That data can then be used as the input to simulation of your choice.
"Small differences in initial conditions, such as those due to errors in measurements or due to rounding errors in numerical computation, can yield widely diverging outcomes for such dynamical systems, rendering long-term prediction of their behavior impossible in general. This can happen even though these systems are deterministic, meaning that their future behavior follows a unique evolution and is fully determined by their initial conditions, with no random elements involved. In other words, the deterministic nature of these systems does not make them predictable. This behavior is known as deterministic chaos, or simply chaos. The theory was summarized by Edward Lorenz as: Chaos: When the present determines the future, but the approximate present does not approximately determine the future."
Ed Lorenz was a weather modeler; he invented chaos theory to explain the long-term failure of his models. But weather models and climate models are nevertheless useful today if we're aware of their limitations.
If you need clarification, isn't this similar to Deep Thought?
Due to the complexity of the real world we can't even predict weather for 2 days in advance and you they think they can predict the world...
The sad part is that they will then use this model and its result to create policies not knowing that you can never rely on such a "copy" because even the smallest difference will create an insane amount of divergence even 1 day down the line of future simulation let alone years...