In my PhD I was analyzing petabytes of structured data from particle detectors that required relatively complex algorithms just to reconstruct, before any physics analysis is ever started, and the analyses were just as complex. My experiment’s total code base was at least 100kloc of C++ and Fortran if not more, and used a farm with over 20k cores for massive distributed processing, single jobs regularly taking hundreds or thousands of cores and terabytes of memory.
I’m currently working on distributed GPU accelerated hydrodynamic simulations where a single input is in the ~1-10Gb range of structured mesh data and the calculation requires ~200+ Gb of GPU memory.
I really don’t see how excel would be a viable option for basically anything I have worked on in my scientific career, beyond my undergraduate toy analyses.
https://genomebiology.biomedcentral.com/articles/10.1186/s13...
https://www.theverge.com/2020/8/6/21355674/human-genes-renam...
So did Excel not cut it, or was it a preference to use Python? Back then Excel supported up to 65,535 rows, and would typically crash with over 8,000 rows. I worked in that era in Excel. One model once was split across three spreadsheets. Only one spreadsheet could be open at a time, they would take 20 minutes each to load, and there was about a 50% chance they would crash while loading.
So what do you do if you need over 65,000 instances of labeled data? For a neural network it's nice to have a million instances, yes a million.
R and Python have in them what is called a dataframe. It's a spreadsheet, but in another programming language. We tend to load our data into those, which is just like Excel, but without the hardware limitations. So in many ways, today it's just like it once was, but we get to choose which programming language to use while working in a spreadsheet, and let's be fair, the Excel programming language isn't exactly great.
We record (electrical) neural activity from behaving animals. The raw data is about 1 Gb/minute, and we collect hours of it per day, 5-7 days/week. The lab next door has microscopes that produce image sets that are in the Gb to Tb range (big swathes of human brains, imaged at micron resolution). An RNAseq experiment is ~20 Gb range; other genomic things are similar, maybe a bit smaller.
All of this would be intractable in Excel, even over a long weekend :-) On top of that, Excel often doesn't do much of what you would need: there's basically no support for signal processing or image analysis or genomic work.
Examples that immediately come to mind include : numpy, tensorflow, etc ...
Even heavy-duty I/O can be made to crank with python if you do it properly.