While coding assistants seem to do well in a range of situations, I continue to believe that for coding specifically, merely training on next-token-prediction is leaving too much on the table. Yes, source code is represented as text, but computer programs are an area where there's available information which is _so much richer_. We can know not only the text of the program but the type of every expression, which variables are in scope at any point, what is the signature of a method we're trying to call, etc. These assistants should be able to make predictions about program _traces_, not just program source text. A step further would be to guess potential loop invariants, pre/post conditions, etc, confirm which are upheld by existing tests, and down-weight recommending changes which introduce violations to those inferred conditions.
ChatGPT and tab-completion assistants have both given me things that are not even valid programs (e.g. will not compile, use a variable that isn't actually in scope, etc). ChatGPT even told me that an example it generated wasn't compiling for me b/c I wasn't using a new enough version of the language, and then referenced a language version which does not yet exist. All of this is possible in part b/c these tools are engaging only at the level of text, and are structurally isolated from the rich information available inside an interpreter or debugger. "Tackling" unreliability should start with reframing tasks in a way which lets tools better see the causes of their failures.