I've been told by a physicist friend of mine that one of the big problems is the amount of data produced by LHC sensor arrays. While the LHC is running, 600,000,000 particle collisions happen every second. Every collision produces particles that decay in complex ways into more particles, and so on.
I forgot the numbers, but they have severals layers of filtering between sensors and long-term storage. First there are FPGA-based real-time filters, next to the sensors, which throw away most of the data as "noise." Then there are local CPUs which throw away most of the remaining data, again classified as "noise" or uninteresting. Finally, what remains (30,000 TB/year) is stored long-term to be later analyzed by physicists.
All levels of filtering and analysis, from the FPGA to the physicist's algorithms, make use of the Standard Model itself and the rest of known physics to figure out what's "interesting" and what's "noise."
One big problem is thus: how can we find new things if we are only looking for what we already know? Hence the need for machine learning and automatic pattern discovery.