I'm always wondering if typos and grammar mistakes impact significantly the quality of the response. After all LLMs are next-token predictors, and I suspect that in the training set (internet), bad writing is correlated to low-quality content?
I was wondering about this one too.To make things worse I also use speech to text and that introduces its own inaccurate transcriptions and typos. Are there any reliable research around this ?
I think any of the current interfaces using thinking tokens and other contexts, make the actual input one types such a small part of the total input tokens that it probably doesn’t have much influence.