I was reading the piece, waiting for the punchline of what kind of unholy beast of a workstation she was using, and wasn’t disappointed.
Thing is, physicists, heck, scientists in general, are not programmers. Fortran 77 and python are pretty much the only shows in town - and your usual data crunching script will be huge, procedural, in log time, and will eat mountains of memory while making disks thrash as hard as humanly possible.
For instance, I helped out a postdoc who I shared a lab with with his ephemeris calculator - it’d take in a series of fits images, and it’d tell you the ephemeris of whatever object you chose - by editing the source, and putting in the x/y of the object in the first frame.
Thing was, it took all night to do this for a single object from half a dozen frames. Most of the time was spent opening and closing each file to read each pixel, and then stuffing those pixels into a gigantic array, and writing that array to text files, and then re-reading those files, and then doing matrix multiplication and all sorts of amazingly baroque stuff that must have seemed like a good idea at the time. He was running it on a monster (for the time) of a workstation, with 64GB of ram and several TB of storage.
I banged together an app in C++ for him with a basic tcl/tk gui, and what had taken him a day of setup and a night of processing and an ungodly machine instead took about as long as it took for him to click on an object and click “go” - on my creaking laptop with 128mb of ram and a 1.4gb hard drive.
This was far from singular - after this, I found myself being “the guy” to talk to about your slow scripts - and that was basically every script in the department.
So no, not better, considerably worse, and I can’t help but think that having a more cross-disciplinary approach to science (embed tech people!) would yield benefits across the board.
I've definitely seen file I/O in the middle of FORTRAN loops. Entirely correct from a functional perspective, and total disaster from an actual runtime perspective.
To be fair, how many blog posts have been written over the years on programmers doing ridiculous things with SQL, or network IO, or file IO, etc. etc.? If trained programmers struggle with these things, I'm willing to give non-programmers some slack.
For example once $LARGE_BIOTECH retired its inhouse supercomputer (A TruCluster composed of multiple GS1280s, the epitome of classical UNIX power) and replaced it with Linux and NFS. The principal engineer noticed his pipeline never finished on Linux but how no idea where to start.
Within five minutes I noticed that GNU/Linux sort uses $TMPDIR for large sorts (multipass). They had pointed TMPDIR at an NFS mount and were doing a multipass sort over NFS. On TruCluster, tmp pointed to an ultrafast cluster diskl.
The bottom line is: In this kind of science, you are used to "big data", to massively parallel computing, to computation and statistics on various levels of abstraction. But the engineering skills outcome greatly vary, because from the physics perspective, computers are just a tool for getting the job done.
For instance, big data in astrophysics is quite different from big data in accounting. Complexity in astrophysics programming is also very different from complexity in the banking buisness. People tend to get arrogant due to their years-long experience, but in the end all they have is just years-long experience in that particular domain, let it be buisness or physics.
That is to say, if you have to ask, be skeptical of any terse answer.
I don't care who has the best paint brush technique. What is interesting is the art and how to creatively solve an interesting problem. There is no shortage of people who can masterfully paint absolutely uncreative shit.
This article though didn't really get deep enough to learn much from.
I'd argue this is simply an unnecessary comment altogether.
Clickbait titles revolve around sounds and language though, so I don't know if there's a way to combat what's profitable and not strictly speaking misleading.
I thought it was the oldest joke in the book that at the start of the academic year, physicists pronounce to all and sundry they're superior to all non-physicists. They then they spend the rest of the academic year fighting over what the most superior branch of physics is.
As they say, the best jokes have an element of truth to them...
She actually explains that in the article:
"one object might be anywhere between roughly 50 gigabytes to maybe a terabyte. By the time I'm done reducing that data and imaging it, it'll probably have roughly tripled to quadrupled in size. I tend to have larger surveys, so by the time I'm done I might have several tens of terabytes of data."
> their math/programming/modeling skills are not better than the average compsci,physics phd.
No, but the data they handle regularly is just much, much larger than in many other fields.
Do they? I don't know and I am really curious: Do you have any source for this?
> It is stored forever in publicly accessible places (a large amount of it, for free).
Same for telescope arrays' data, but the same problem: How to get the data from its source into your number crunchers? When I know the pipeline, I can simply calculate the answer, but when I need to explore/play with data, I will inevitably run into IO roadblocks.
> How to get the data from its source into your number crunchers?
You do not simply "get" this data. It's simply too large, conceptually infinite. You query it for the parts that you are interested in, and then you extract only those parts.
As an example of free access, see the access hub for the Copernicus program: https://scihub.copernicus.eu/ There is an interface for downloading particular images, and an api to query the archive of images with space/time constraints to obtain a manageable list of URLs to download. You have sentinel-2 for optical images, sentinel-1 for radar images, and sentinel-5p for hyperspectral.
If you want higher resolution images, they are also publicly accessible, but typically not free. There's plenty of commercial satellite imaging providers nowadays, and lots of companies work by extracting information from this huge corpus of data.
Once I visited the astro folks at Caltech and they were light years ahead of everybody else in every tchnical dimension. But that's CAltech for you.
That is, not so much.
At the PhD and research level the data types and questions are so specific that the analyses often are bespoke (in academia, the more bespoke the analysis, the more you can sell its novelty). PhD-level data analyses have very little relevance outside of its own field.
There is nothing profound to be had here, even if they have the prestigious title of "PhD astrophysicist".