Why AI Spending Isn't Slowing Down
wsj.com
wsj.com
Reasoning/chain of thought seem like diminishing returns to me and I worry they are a bit of a dead end/local optimum. Reasoning models call the language models tens of times so it will be tens of times less efficient than the underlying language model but the quality is not tens of times better. It also feels finicky to me. The bump from Chat GPT-3 to Chat GPT-4 was a reasonable positive shift across the board. The reasoning models produce answers with a different vibe, maybe better overall but worse at some tasks better at others. I can use O1 at no additional cost so I do use it fairly often but I often consciously opt for 4o either because I prefer the results of the quality boost from O1 isn't worth the wait.
We could have accomplished a lot of good with those resources.
Speech to text using whisper is almost perfect.
I once worked with someone who wasn’t fluent in English, but could read it really well. He had whisper running during meetings because he could read at the speed it translated, but couldn’t keep up with our casual speech.
Honestly, we didn’t even notice he was using speech to text to keep up with us for a few weeks. We only noticed because of screen sharing.
I imagine that would be a big social boon for deaf people as well.
All that to say, it’s not like there was _no_ good done with all that money.
Any day now you'll barely notice he isn't even there anymore; which seems to be the end goal.
It's not like this mad rush into la-la-land doesn't have negative consequences for society.
This isn't true. On benchmarks whisper is not SOTA. It is said to be noise resistant but it doesn't compare well with Conformer based architectures ever on Librispeech mixed. Definitely not perfect, and it doesn't work for medical transcription.
Although all of this is from a production lens. For personal use, honestly nothing is as easy to use as Whisper (even works on a laptop).