The main signals on Search have been AI/ML for >5 years now.
At least for Google, AI isn't just a parlor trick.
The main signals on Search have been AI/ML for >5 years now.
At least for Google, AI isn't just a parlor trick.
Say more? "Passing the entire web through ML-model inference to generate 'signals' — IQL expressions? — at the same global-concurrent throughput as previous CPU-based indexing" sounds like something that should have required Google to get so many TPUs fabbed that it would have eaten the entire world's chip-fab capacity for multiple years. It's hard to imagine what they could have done to get around that.
Did Google come up with some kind of model-hardcoded ASICs that could do this single task orders-of-magnitude more efficiently than more "flexible" GPU/TPU-based approaches would? Maybe, for example, they used some process to convert the weights of an arbitrary Transformer into VLSI for a search memory / memristor network / etc — essentially giving them an inferencing DSP?
Or maybe they figured out a mostly-lossless method (that either doesn't use ML itself — or maybe uses ML hyper-optimized into some CPU-viable model that fits in L2 cache) to pre-process + normalize webpage documents into chunks that could be effectively content-deduplicated? (And then the ensuing inference passes — being context-free — would only ever need to be done once for a given chunk, significantly reducing the inter-page costs, and massively reducing the same-page over-time costs, of inference.)
Or something else I can't even think of. (I'm guessing it's this one.) Either way, crazy stuff.