New features which have some random amount of error do not make me happier. They make me more afraid.
I have been grilled on single commas being missing in documents - for good reason! It was a financial model.
For me - LLMs are
- English majors - Decompression systems
They cannot conform to my expectation of code
- their output is inherently randomly distributed - they are being pigeon holed into doing logic / Thinking.
Any use case that - exposes an LLM to infinite interactions (open to users at internet scales)
Simply represents a unsolvable edge case issue to me.
Production solutions need to figure out - a way to verify the output is correct (fast, scaled, cheap!), - figure out use cases that accept uncertainty gracefully.
From what I am seeing, there is a bit of both happening.
1) Get the LLM output close enough to 99% correct, 2) send it to a human as a “co-pilot” to verify it is correct.
It is fatiguing to track every new variation on getting somewhat closer to some figure between 89 to 99.9999%.
I love the tech. I love what it actually achieves. I use it regularly to make my life better. This is not an attack on the tech.
My life would be MUCH easier, if someone could tell me their bot or solution is going to screw up massively 89% of the time, and that expert users can identify the 11% cases 100% of the time.
Also that my experts arent going to suddenly see their workload expand by infinity% if this tool goes live.