But no one cares about those kinds of productivity gains. Just the ones that will completely replace us.
But no one cares about those kinds of productivity gains. Just the ones that will completely replace us.
My comments are more in the context of OLAP queries and other non-normalised data often queried via SQL.
I train non-LLM transformer models on (older and rarer) datasets, and automating the ingestion of sprawling datasets with hundreds of columns, often in a variety of local languages and different naming conventions adopted over decades, with quite a few duplicated columns…. The LLMs perform badly, it’s nigh impossible to test (for me as a user in prod) and it’s nearly impossible for the LLM companies to test (in training) to RLVR and RLHF this.
> I train non-LLM transformer models on (older and rarer) datasets, and automating the ingestion of sprawling datasets with hundreds of columns, often in a variety of local languages and different naming conventions adopted over decades
All of this sounds like basic data processing
Laid off your DBAs I see.
I do enjoy giving the frontier models wacky projects that I can't even find examples of how to do online but I don't expect any results or need them and some have done really well with it while others fall on their face (models)
[0]: Like https://www.oreilly.com/library/view/sql-queries-for/9780134...
Unfortunately I am very good at forgetting things I resented having to learn, and SQL is definitively one of them.
But if you ever need to query unknown data, then probably you should learn SQL a bit deeper.
I think you may be describing the experience of 6-12 months ago.
An eight-join query is going to be nigh on unmaintainable should the requirements change, leading to a change-break-change-break spiral as your preferred coding agent tries to fix its previous fixes.
Maybe the wise way to use AI would be to sort out the schema.
A highly normalized DB can easily end up with 8 joins required for some function. That's really not out of the question. "Sorting out" the schema then would be... denormalization, which is a thing, but you need to know why you're doing it. And I think 8 joins isn't enough of a reason.
I'd rather get it from the LLM and review