> Have you experienced this kind of problem? In what tasks does it show up most for you?
I have experienced this type of problem. A colleague asked an LLM to convert a list of items in a text to a table. The model managed to skip 3 out of 7 items from the list somehow.
> Would solving it be valuable enough to pay for? Do you see this as something LLM providers will solve themselves soon, or is there room for an external solution?
The solution I have found so far is to prompt the model to write and execute code to make responses more reproducible. In that way most of the variability ends up in the code, but the code outputs tend to be more consistent, at least in my experience.
That said, I do feel like current providers will start to or are already working on this.