80 karma · joined February 14, 2016
We run most things on Alma Linux, but there's also RedHat in the mix.
A good intersection of CERN and Gentoo is EESSI. They use CVMFS, which is a FUSE filesystem used for software distribution from CERN, and Gentoo as part of their software stack. See links below for more info:
https://github.com/xrootd/xrootd/blob/master/cmake/XRootDVer...
and also the genversion.sh script at the top of the repo.
I use these plus #cmakedefine and git tags to manage the project version without having to do it via commits.
EESSI (https://eessi.io) has taken this model further by using CVMFS, Gentoo Prefix (https://prefix.gentoo.org), and EasyBuild to create full HPC environments for various architectures.
CVMFS also has a docker driver to allow only used parts of a container image to be fetched on demand, which is very good for cases in which only a small part of a fat image is used in a job. EESSI has some documentation about it here: https://www.eessi.io/docs/tutorial/containers/
If you come to CERN, you can also see the NeXT machine that was the first web server.
https://github.com/RandyGaul/cute_headers/blob/755849fc2819d...
See an optimized quaternion multiplication implementation in SSE by me here:
https://stackoverflow.com/questions/18542894/how-to-multiply...
https://github.com/RandyGaul/cute_headers/blob/755849fc2819d...
I fixed this exact problem in a highly used library in high energy physics:
https://gitlab.cern.ch/CLHEP/CLHEP/-/commit/5f20daf0cae91179...
Many believe the C++ compiler will magically optimize the switch away, but in some cases, like the example above for CLHEP, it doesn't happen, so you end up with bad performance.
- https://indico.fnal.gov/event/23628/contributions/240607/
- https://indico.cern.ch/event/1338689/contributions/6077632/
Faster Math Functions:
https://basesandframes.files.wordpress.com/2016/05/fast-math...
https://basesandframes.files.wordpress.com/2016/05/fast-math...
Even faster math functions GDC 2020:
https://gdcvault.com/play/1026734/Math-in-Game-Development-S...
- Gentoo Prefix https://wiki.gentoo.org/wiki/Project:Prefix
- European Environment for Scientific Software Installations (EESSI) https://www.eessi-hpc.com/
- REANA (Reusable Analyses) https://reanahub.io/
> In general, managed tools will give you stronger governance and access controls compared to open source solutions. For businesses dealing with sensitive data that requires a robust security model, commercial solutions may be worth investing in, as they can provide an added layer of reassurance and a stronger audit trail.
There are definitely open source solutions capable of managing vast amounts of data securely. The storage group at CERN develops EOS (a distributed filesystem based on the XRootD framework), and CERNBox, which puts a nice web interface on top. See https://github.com/xrootd/xrootd and https://github.com/cern-eos/eos for more information. See also https://techweekstorage.web.cern.ch, a recent event we had along with CS3 at CERN.
https://news.ycombinator.com/item?id=39657703
See the CI workflow using my script here: https://github.com/xrootd/xrootd/blob/master/.github/workflo...
Also, with some advice from Henry I wrote a nice new setup.py for the Python bindings of XRootD integrated with the CMake build. I didn't want a dependency in scikit-build in the end, but his advice helped me a lot.
The idea to classify cycles into front-end bound, backend bound, bad speculation or memory bound is brilliant. Once you know which one your program suffers from, it's easy to know what can be done to improve things.
1. https://cmake.org/cmake/help/latest/prop_test/WORKING_DIRECT...
2. https://cmake.org/cmake/help/latest/command/configure_file.h...
3. https://cmake.org/cmake/help/latest/prop_test/FIXTURES_SETUP...
I have examples for Julia sets and the Mandelbrot set, including an implementation with AVX2 intrinsics.
These days with std::simd more widely available there's less of a reason to use VecCore, but the examples may still be educational enough. I chose Julia sets and the Mandelbrot since they are perfect examples of simple problems that compilers fail to vectorize on their own.
In high-energy physics, ROOT is /the/ toolkit for data analysis, and I guess jsROOT (https://root.cern.ch/js/) could also be used to load data to be shown in Framework dashboards. I thought the idea of Framework as a blogging engine with powerful data visualization built-in could be very interesting. Think, for example, about physicists pulling open data (https://opendata.cern.ch) and writing about their analysis or someone pulling data from https://ourworldindata.org/ in their own visualizations to support their case while writing about a particular subject, etc.
EOS was designed and developed with the unique needs of the LHC experiments in mind. The advantages it has are features used by the experiments, like support for remote access via the XRootD protocol, which is used for data analysis with ROOT (only the parts of files needed by an analysis are downloaded); rich support for client authentication methods (Kerberos, X509, OIDC, etc); and support to also FUSE mount everything to give a convenient POSIX-like view of the data. EOS needs to sustain ingestion of data at high rates from experiments (10s of GB/s each) for several months at a time during data taking without any downtime, while at the same time having tens of thousands of clients connected reading data as well. It's also integrated with the CERN Tape Archive (CTA) and File Tranfer Service (FTS), used for long term archival and data management across sites, respectively.
In the cases where block/object storage is needed, like storage for VMs, S3 storage for various uses, etc, then ceph is better suited. It has lower latency and EOS does not offer block-level access. In addition to providing storage services for OpenStack/Openshift, ceph is used to provide storage to back AFS and CVMFS, for example. CVMFS is another interesting piece of CERN's infrastructure, it's a read-only, HTTP-based FUSE filesystem used to distribute the software used by the experiments to grid sites around the world. Dan van der Ster, mentioned in the article above, has a good overview of ceph usage at CERN here: https://youtu.be/2I_U2p-trwI?si=Tsq4h8NIu4vSZQwt
If you are interested in EOS, we have the EOS workshop coming up in March: https://indico.cern.ch/event/1353101/
Our EOS clusters have a lot more nodes, however, and use mostly HDDs. CERN also uses ceph extensively.
We also have meetings dedicated to performance, some of which are not public, but this series from ROOT is: https://indico.cern.ch/category/14122/ If you search above, you will see many discussions about performance. The CI for ROOT also has a set of benchmarks to catch regressions, and Geant4 has two systems to track performance, a CI job checking every merge request, which I've set up myself (not publicly accessible), and a more complex system to track performance run by FNAL: https://g4cpt.fnal.gov/
These are just some examples from the projects I've worked on. There are also efforts to port stuff to GPUs and HPCs, and many other projects like event generators that are also undergoing performance work for HL-LHC. If you Google you can probably find a lot more stuff than what I already mentioned. Cheers,