but it takes a huge amount of training/test data to do it and takes a lot of computational work after you've got that data. What TREC revealed is that most of the things that would obviously improve search relevance don't, and it took 5 whole years of 20 teams working on it before a useful discovery was made.
There was an article about TikTok that revealed just how wrongheaded the viewpoint of the current web is. As much as Google fetishizes data, the data collected by sites like YouTube is useless because they offer you too many things to click on so you click around like a chicken with its head cut off. There are maybe 5 things that you like out of 20 that they show you, but which one you click on is random so the signal is mostly noise and worst of all they can't come to the conclusion that you didn't like any of the other 19 things they showed you.
TikTok shows you exactly one thing at a time so the opinion that they capture is meaningful.
It reminds me of one of the first "learning to rank" papers where I talked the management at the CU Library to let the Thorsten Joachims group run our search engine and we realized just how poor of a signal you get from search engine usage and how challenging it is to feed it back to improve your results.
Really marking up judgements for all the results that turn up and using "pooling" to add new results to the original set when your search engine is essential to make really better search. Otherwise you can be just another one of those guys who posts to Medium about the "semantic search" engine he built but can't tell you if the results are any better than any other search engine.