I wonder what will happen to actual content then. Currently YouTube is showing info about most watched section of the clips. It saves so much time! Now imagine that happening to everything above.
I wonder what will happen to actual content then. Currently YouTube is showing info about most watched section of the clips. It saves so much time! Now imagine that happening to everything above.
One basic but effective demonstration I’ve seen was summarising a 30 minute talk on YouTube into dot points.[1]
I watched the video and read the summary afterwards and was almost completely satisfied with the summary.
At scale, the flexible compression and expansion and navigation of information is potentially huge … like Google Maps for the internet.
[1]: https://gist.github.com/simonw/9932c6f10e241cfa6b19a4e08b283...
With that being said, one of the main challenges ahead will be multimodal learning. We're sort-of there combining text with visual data, but there are many other modalities out there as well.
I guess the problem is taking a new document (like a search term) in the higher dimensional embedding and reducing it to three dimensions for searching in that reduced space and expecting that to also maintain the same nearest neighbor ordering.
EDIT — answered my own question: there is indeed an OpenAI Embeddings API :
https://beta.openai.com/docs/guides/embeddings
That plus a vector similarity engine (FAISS for example) is the key to these types of apps (Thanks to Simon Willison’s blog, which he pointed to elsewhere in this thread)
... and our tool "sees" the above text as ...
For the problem of data-supported AI search, the content really matters. fragen.co.uk's edge is semantically chunking the content, in other words we are splitting the content up into key facts ready for recall. Splitting the content up into key facts ready for recall makes data-supported AI search solveable.
(hope it's visible how an LLM like GPT able to use/quote the above can perform seriously better at those bothersome it/what/where questions and follow ups)
Personally I think question answering is still very gimmicky, in particular because I can't clearly understand why the answer is what it is. Ctrl-f is completely explainable, I know why it works and why it fails, and is a much more useful tool to engage with books. The main problem seems that the full text of most books is not available unencumbered to process as one sees fit
Demo we have created for our website: https://speech-kws.ozonetel.com/ozosearch
“AI, what does this book say about integrity?”
“(answer), based on paragraphs on page 143 and page 210”
“Show me those paragraphs inline, starting with the paragraph where the author transitions in to the topic”
Using chat GPT to solve this problem seems pretty awesome, it would be super handy to query my (larger than it has any right to be) library of reference books using not-exactly-competent requests.
This and translating technical material into 5th grade reading level explanations are probably the best two uses of chat GPT I've seen yet.