> People forget how fast modern hardware is and overestimate how much useful data they have (or need).
Where your model breaks down is when you have folks throughout your engineering org who have data needs but don't have folks to spin up bespoke pipelines like this and spend time optimizing them. You have data, you want to write something SQL-ish or have some nice APIs, and you want to get results dumped into a predictable place. When you have mixed workloads, mixed data sources, and those jobs are being tweaked and changed frequently, the actual underlying compute cost is hardly the issue. Getting the data, chewing on it [fast enough] without having to spend much time optimizing, getting it to its destination, and making that happen regularly and reliably is where the value is.