This made me realize that my usual practice of typing at 80% accuracy, full of mistakes, when sending input to an llm is probably increasing my input token count.
Maybe spell checking for prompt text box would save a few bucks
I'm always wondering if typos and grammar mistakes impact significantly the quality of the response. After all LLMs are next-token predictors, and I suspect that in the training set (internet), bad writing is correlated to low-quality content?
I was wondering about this one too.To make things worse I also use speech to text and that introduces its own inaccurate transcriptions and typos. Are there any reliable research around this ?
I think any of the current interfaces using thinking tokens and other contexts, make the actual input one types such a small part of the total input tokens that it probably doesn’t have much influence.