5 karma · joined May 25, 2026
Consider this prompt:
Make sure that, whenever a job runs, we can tell where it started from, who triggered it, what settings were passed in, and what changed later. In the future, if something breaks, we want to be able to trace it back and understand what happened.
That is fine.
But an expert can say:
Jobs should persist provenance metadata."
That only works if the model is trained, specifically the way you want, to understand that second sentence. If not, any model could work with the first sentence, but not with the second.
You've crated a need for expert training at the model level (which is insanely expensive to create and maintain) rather than accessible natural language discussions that work on any model and understood by anyone. Denser isn't "better" because the words have more power.
You can test this with any platform, simply type "improve this response" after any result and see what you get back (no thinking credits, nothing fancy just three words and a few seconds)
Interestingly they did not share what prompts they used to run any of this – highly suspect.
Maybe your job isn’t at risk of being replaced, but your entire organization is.
If you move those things to software and utilize tools that are cheap at scale (databases, web search etc.) the hardware arms race ends and the price becomes sustainable. With the right tools preparing dynamic context for a conversation, models are used for their reasoning and not for their knowledge. And waiting even a minute or two for a model to prepare a response, evaluate it, and iterate to improve quality makes a huge difference.
The problem is that the entire market right now is set up for resolving things with hardware instead of software. Bigger models, faster chips, data center sprawl etc. Clearly that is not sustainable. But if you use AI inference for reasoning only, and traditional software like databases, web searches, etc. for the things they do well (at a fraction of the cost) the economics flip.
Trying to force expensive training updates and ever larger models that know who the first baseman was for the winning 1939 world series time is useless and where the cost-to-value ratio is broken. Pull what you need from a search api, and send the results to an LLM to analyze. No fine tuning, no super big chips and even small models can produce good results.