https://cdn.some.pics/snekoil/66594bca93e19.jpg
No way llama3-7B Q4 would be able to under the same on a home GPU
https://cdn.some.pics/snekoil/66594cb8cddcd.jpg
No way GPT4 can perform a web search to determine the source is the onion and flag it as satirical or summarize the page and flag it satirical
https://cdn.some.pics/snekoil/66594e1a3de29.jpg
No way Claude can flag that it’s satire.
https://cdn.some.pics/snekoil/665955fb22618.jpg
No way WizardLM30B would suggest it’s satire because eating rocks is not scientifically backed.
Yes, it’s not reliable, but it shows that there’s enough information encoded in even small edge models without access to the largest data warehouse of web content and human interaction and hundreds of millions of dollars in content deals to put a “likely fake on it”.
Yes, it may be unsolved at SERP scale, not chat scale, but Google chose to have this fight on SERP and launch so they have to take the hits.
Yes replacing context with Gemini god quotes to monopolize traffic onto google is bad, but given the above results augmenting results rather than appropriating them seems to have some potential as seen with Perplexity.
Yes LLMs are flawed, but as some competitors show they can improve on the very low bar that is Google search results and not just by removing the paid misinformation / ads. Google showed they can make it worse.
Can you solve the problems you have posed with some reliability? If yes, there’s a good chance it’ll be automated at some point.
Claiming something is unsolvable rarely pans out unless it violates some fundamental law of physics.