CERN pushes storage limits as it probes secrets of universe
news.idg.no
news.idg.no
Designing the trigger is extremely complex, since the various detectors within each experiment have greatly varying response time. Only data from fast detectors is available for the low level trigger.
In addition, the speed of light is a real barrier for the lower levels of the trigger; by the time the debris from a collisions reach the outer reaches of the experiment (this is usually where the muon chamber is), there have already been an additional collision at the center.
(The speed of light is about 1 ft/nanosecond, the radius of the muon chamber in CMS is about 25 ft, and the time between bunch crossings is about 25 nanoseconds.)
The design of the trigger is a very important and often contentious process. A bad trigger will throw out important physics events, and trade-offs can favor one physics search (e.g. the Higgs) over another (e.g. supersymmetry).
First they say that they "generate around 1 petabyte of data per second"
Then they say "ATLAS produces up to 320M bytes per second, followed by CMS with 220M Bps. The data from ALICE amounts to 100M Bps and LHCb produces 50M Bps." only that sums up to 690M Bps ... definitely not 1 petabyte per second. (That is, assuming that 1M Bps means 1 million bytes per second, or just under 1 megabytes per second.)
And then, later on, they talk about a different mode in which "more data is produced by the four experiments, about 1.25G Bps in total." which is still not 1 petabyte per second.
What's going on?
Further preliminary analysis is performed on the retained data, broadly categorizing the energy and other characteristics of the collision. That allows individual physics groups around the world to download only the data that is likely to pertain to their specific research, e.g. the Higgs boson, multiple dimensions, etc.
There was some talk of transferring data via Bittorrent or perhaps a custom protocol involving fountain codes. That never got off the ground. Instead, the Russians were working on a custom peer-to-peer system with a monolithic centralized set of indices, a system which is hopefully working better than it used to.
P.S. - Here's a hummingbird-speed video of building our prototype fileserver node for local physics analysis of ATLAS data [before I learned about electric screwdrivers]: http://www.youtube.com/watch?v=8y6MpPNqxmw
There's also many different groups that need LHC data to perform many different analyses, so the more data the better.
So all in all, you would get a few dozen weeks of real operations a year, which would include some stable beam (if you aren't into beam gas studies or cosmics).
Strictly speaking, (down)time can be "irrelevant", in the sense, that with higher luminosity you get more data (LHC can catchup on Tevatron easy). You can have 5 nine uptime, with 1 bunch circulating, or lots of bunches of particles (the beam is not continuous, it comes in trains of particles). So one thing is, that you also go for as many bunches as possible...
But the periodical reports are on cdsweb, because of the public funding agencies ie. for the machine itself, setting an upper boundary: "Downtime statistics over the 2010 run" -- Chamonix 2011 Workshop on LHC Performance, Chamonix, France, 24 - 28 Jan 2011, pp.70-74. Then the DAQ of your experiment of choice comes on top of this...