ROOT: analyzing petabytes of data scientifically
root.cern
root.cern
I left about 5 years ago, and ROOT was in a process of change. They already ripped out the old CINT interpreter and moved to a clang-based codebase, and now you can run your analyses in Jupyter as far as I know (in C++ or Python). I heard the code quality has improved a lot, too.
TTree is succeeded by RNTuple, which is basically CERN's take on Apache Arrow, they're incredibly similar
I haven't used ROOT so I don't know how well it would work to write bindings for it in Haskell; it can be hard to provide a good interface to an implementation that was designed for a totally different style of use. Possible, just difficult.
EDIT: there's one at https://hackage.haskell.org/package/HROOT
In other words, why do you think it’s not a good fit?
It's much more likely the physics community would adopt Julia, or maybe Rust, and even that has been pretty slow.
(nothing I said above should be construed as taking a position about the suitability of any specific language or lack thereof for doing scientific computing. I have opinions, but I am attempting to explain the reason factually with a minimum of bias)
Using something like Haskell for ROOT is ridiculous for a lot of obvious reasons. A simple and dismissive "no" invites the cautious reader to discover them on their own rather than waste engaging in a protracted debate. Maybe it's better to reject the idea out of hand and spend our time elsewhere.
You’re free to not do any of that, of course, but be prepared to defend the fact that you’d prefer not engaging in discussion and instead just shallowly dismiss something.
When the cost of complexity of interacting with an API is paid by the LLM, optimizing this particular part of software design (also one of the hardest to get right) will be less fashionable.
Also I really like their 404 page [2]. And no it is not about room 404 :)
Past HN discussion on Julia for particle physics: https://news.ycombinator.com/item?id=38512793
But this particular problem (per row computation) have different options to tackle now in hep-python ecosystem. One approach is to leverage array programming with NumPy to vectorize operations as much as possible. By operating on entire arrays rather than looping over individual elements, significant speedups can often be achieved.
Another possibility is to use a library like Awkward Array, which is designed to work with nested, variable-sized data structures. Awkward Array integrates well with uproot and provides a powerful and flexible framework for performing fast computations on i.e jagged arrays.
For the record, vector-style programming is great when it works, I mean Julia even has a dedicated syntax for broadcasting. I'm saying when the irreducible complexity arrives, you don't want to NOT be able to just write a for-loop
Just a recent example, a double-for loop looks like this in Awkward array: https://github.com/Moelf/UnROOT_RDataFrame_MiniBenchmark/blo... -- the result looks "neat" as in a piece of art.
Later, even when Pytorch added support for 3.12, nothing changed (so far) in Taichi.
how is this a lame excuse
>but it fails on a bunch of PyTorch-related tests. We then figured out that PyTorch does not have Python 3.12 support
they have a dep that was blocking them from upgrading. you would have them do what? push pytorch to upgrade?
>Later, even when Pytorch added support for 3.12, nothing changed (so far) in Taichi.
my friend that "Later" is feb/march of this year ie 2-3 months ago. exactly how fast would you like for this open source project to service your needs? not to mention there is a PR up for the bump.
I stand by my original comment.
Another example: Gravitational waves were found with GStreamer at LIGO: https://lscsoft.docs.ligo.org/gstlal/
That being said, I don't know whether it's actually a good idea for someone external to actually use it. My experience may be a little outdated, but it's quite clunky and dated. The big advantage of using it for CERN or particle physics stuff is that it's basically a standard, so it's easy to collaborate internally.
The other one, gstreamer, is a beautifully designed platform with an architecture so nice it can be easily abstracted and reused in completely different scenarios, even ones that probably never occurred to the authors.
Say WHAT now?!
Also, for the longest time, the I/O format wasn't very well documented, with only 1 implementation.
Now, thanks to groot [1], uproot (that was developed building on the work from groot) and others (freehep, openscientist, ...), it's to read/write ROOT data w/o bringing the whole TWorld. Interoperability. For data, I'd say it's very much paramount in my book to have some hope to be able to read back that unique data in 20, 30, ... years down the line.
[1] https://go-hep.org/x/hep/groot (I am the main dev behind go-hep)
- https://github.com/go-hep/hep/blob/main/groot/rhist/efficien...
I understand this has not been a priority so far.
It kinda works if you open a magic file with a specific on-disk representation which bypasses this, but that’s not a solution at all.
My plots looked a lot nicer ;)
This was way before the Python ecosystem gained traction. And R ML packages were also just starting.
If you don’t use RDataFrame, or it’s just histogram plotting, be very careful with pyroot.
Trying to bring C++ into the Clojure world and Clojure/interactive programming into the C++ world.
Oh gosh. The nightmares. - What obviously shows that you can build extraordinary stuff in horrible environments.
Without Cling, this sort of thing wouldn't be feasible in C++. Not in the way which Clojure dialects work. The runtime is a library and the generated code is just using that library.
There are also a couple of C++ live environments in the game industry.
https://con.cern is not yet used, so...
It was a huge mistake, borne out of greed and recklessness.
They also screwed up some old URL/email parsers/sniffers hardcoding TLDs. Largely the fault of bad assumptions to begin with.
Other than the above, I don’t see much of a problem. Whatever problems people like to point out about gLTDs already existed with numerous sketchy ccTLDs, like .io. Guess what, the latest hotness .ai is also one of those.
It has its rough edges, but you do get a lot of good synergy out of this setup for sure.
comments here have already mentioned couple horror stories of people accidentally/by inexperience doing a lot of work above the framework - if you can save that by not being slow, why not?
Modern ROOT of course replaces CINT with Cling and STL containers are well supported.