Particle physicists turn to AI to cope with CERN’s collision deluge
nature.com
nature.com
The hardness is due to possible ambiguities of what collection of springs caused the observed intersection points, imperfections in the helical shape of the springs and the cylindrical shape of the detectors, limited resolution of the measured intersection points, and missed (detector efficiency »only« 99 %) or added (detector noise) intersection points.
Is the news that they want to use ML in the trigger selection now?
Also, which one of the four experiments is doing this?
EDIT: Ah, it's CMS.
Not quite (and that wouldn't be news). It's using ML for track reconstruction. Not even LHCb does this.
It's a trade-off.
It's common practice to check data/simulation agreements between variables before using them for training. Reweighting procedures are used to improve agreement.
It depends on the experiment, LHCb for example does not use simulated background.
In any case, the more complex your model gets (number of variables) the exponentially more simulated Monte Carlo events you need to fill that multidimensional space.
You're absolutely right about how the amount of data needed for statistical significance grows with the size of the parameter space you're searching.
That depends on the analysis. If you're looking at a partially-reconstructed decay (common in semileptonics) then you can't rely on the regular trick of choosing a sideband sample to work as your combinatorial background. Also, it's very common to model specific misidentified or partially-reconstructed backgrounds using simulation.
> In any case, the more complex your model gets (number of variables) the exponentially more simulated Monte Carlo events you need to fill that multidimensional space.
I understand this argument if you're trying to model the efficiency in nD space (through splines, histograms, moments etc), but that's usually done when you're fitting to n variables, e.g. in an amplitude fit. If you just want the efficiency of a cut on the score from an MVA algorithm, I don't think it matters. What definitely matters is that the behaviour on simulation reproduces that of real signal as faithfully as possible.
You might have a point there.
But surprisingly the value they get out of it is certainly only 1 team's work.
“incomparably more difficult”
It immediately looked like a really interesting challenge to me, but after reading a bit about the state of the art it seems like three months is a pretty short time to come up with a meaningful result even if you could work on it fulltime. Many people already invested a lot of time in that problem and existing solutions are quite sophisticated and good. The material actually mentions that they expect that you will have to take into account things like adjacent detectors overlapping by a few pixels or how particles may light up several pixels if they hit the detector at very shallow angles and cross several pixels as they pass through the detector.
The first thing someone probably considers is something like a Hough transformation and it turns out the creators of the challenge mention that in the material and submitted a solution based on this as a bench mark which achieves a score of about 20 %. If I read the related documents correctly, a meaningful result will require a score of at least about 90 % and the state of the art would probably be somewhere around 95 % to 98 %. The current leader is at 26.48 %, admittedly the challenge is only 5 days old. I am really curious where the scores will be at the end.
I'm not sure I like this approach, however. AI is not pixie dust you can sprinkle over your hard problems and even if you are 100% sure your AI-based system matches perfectly your current systems, you simply can't guarantee it will match the cases you never tested (the never-seen-before data from never-done-before experiments). You'll still need to do the well understood process once you flagged the interesting data (and possibly threw away the sets mislabeled as uninteresting).
A factor of 10, which is what's mentioned in the article, is what you expect to get with Moore's law in three or four years. With current off-the-shelf advances, I'd expect more than a 10-fold improvement in the next three years at the leading edge HPC world. Maybe CERN's problem is one that specialised compute units could solve better.
https://www.kaggle.com/pranav84/beginner-s-guide-to-cern-s-p...
Does anyone know how long this currently takes?
I have no idea what amount of time state of the art algorithms will use, but I could well imagine that it is essentially a tunable parameter that is chosen to get the best accuracy given the amount of data you have to process and the time available to complete the task.
That seems also to be - but I did not look at the code - more or less what the creators of the challenge implemented and submitted as a benchmark implementation, admittedly with the expected poor performance score of only about 20 %.
Not withstanding that, I am not yet convinced that a Hough transformation combined with something similar to a quad tree could not work. More specifically I am thinking of delaying the creation of votes. Roughly the first point just becomes a node in the tree corresponding to the bounding box of its entire possible parameter space. Only when we encounter a second point whose possible parameter space overlaps with that of the first one we split up the two volumes into one volume for which both points vote and a few volumes for which only one of the points vote.
This obviously requires that the possible parameter spaces do not have terrible shapes that are hard to bound and I also could see nearly perfectly overlapping volumes cause issues due to the generation of many small volumes for the imperfection in the overlap. There are probably more issues and possibly even show stoppers, but without picking up a pencil and really thinking about it, I not really tell whether or not it could work out. But, as said, I am also unable to see immediately why this could never work.
Sounds like the LHC’s physicists have an idea.
If CERN had an unlimited budget, I suspect they'd do it however they did it before.
When I last worked there in 2015 a typical pile up situation was having about 50 collisions per detector reading. It is no simple problem to simultaneously reconstruct 50 collisions from the same set of overlapping detector measurements.
Towards the end of last year, they had to start levelling the instantaneous luminosity to 75% of what they could achieve,† primarily to reduce the load on the grid.
† Edit: the maximum peak luminosity is still 200% of the design value, so the performance is beyond initial expectations.
The CERN grid currently[1] has 1 EB of storage and 750k of CPUs, and that pushes out 2 million jobs per day. From own experience of one of the 162 sites, you have something like 6 GB of ram per CPU core, and often jobs need more than that so that in practice you are memory bound.
[1]: https://indico.cern.ch/event/466934/contributions/2524828/at...