Splitting TPUs into dedicated training vs inference chips feels like an admission that the bottleneck has shifted from FLOPs to memory bandwidth + latency. Are future gains to come more from memory/system design than raw compute scaling? What’s that saying about Scaling laws?
I first compute the embeddings to the dictionaries and store these. At query-time, when the user inputs the description, its embeddings are again computed by a lighter model. Then comparison happens between this and the stored ones
The comparison to automobiles changing streets is thrown around a lot. But I feel AI is fundamentally different. It is not a technological change like the internet which brought us huge amounts of opportunities in so many different directions. AI’s goal is to automate (in other words, replace) us.