Show HN: I've built a locally running Perplexity clone
github.com
github.com
It's basically a LLMs with access to a search engine and the ability to query a vector db.
The top n results from each search query (initialized by the LLM) will be scraped, split into little chunks and saved to the vector db. The LLM can then query this vector db to get the relevant chunks. This obviously isn't as comprehensive as having a 128k context LLM just summarize everything, but at least on local hardware it's a lot faster and way more resource friendly. The demo on GitHub runs on a normal consumer GPU (amd rx 6700xt) with 12gb vRAM.
I'm wondering if there's a search API that would make the backend seamless for something like this.
We can escape all ads and dark UI patterns by delegating this task to AI agents. We could have it collect our feeds, filter, rank and summarize them to our preferences, not theirs. I think every web browser, operating system and mobile device will come equipped with its own LLM agent.
The development of AI screen agents will probably get a big boost from training on millions of screen capture videos with commentary on YouTube. They will become a major point of competition on features. Not just browser, but also OS, device and even the chips inside are going to be tailored for AI agents running locally.
I personally use AI for text style changes, as a summarizer of ideas and as rubber duck, something to bounce ideas off of. It's good to get ideas flowing and sometimes can help you realize things you missed, or frame something better than you could.
There are projects like flareresolverr which might be interesting
EDIT: I just noticed that most of the code is Go. Still going to play with it!
is it possible to make it only use a subset of the web? (Only sites that I trust and think are relevant to producing an accurate answer), and are there ways to possibly make it work offline on pre installed websites? (wikipedia, some other wikis and possibly news sites that are archived locally), and how about other forms of documents? (books and research papers as pdfs)
How do you parse and efficiently store large, unstructured information for arbitrary, unstructured queries?
Also, i don't own an Nvidia Card or Windows / MacOS
I'm wondering because most news websites now have a lot of polluting elements like popups, would they also go into the database ?
So I think there may be some useless data in the vector, but that may not be a issue since it is coming from multiple sources (for simple question at least)
This all would be due to optimisations within model inference code and techniques, hardware and packaging of software like the above.
Don't see billion dollar valuations for lots of AI startups out there to materialise into anything.
Why? It's much more efficient to have centralized special purpose hardware to run enormous models and then ship the comparatively small result over the internet.
By analogy, you don't have a search engine running on your phone right?
Will not happen any time soon. Consumer hardware can't even run GPT-4 locally, and won't be able for a looong time. Each GPT-4 instance runs on 8 A100. The cost of such system is ~$81K. Not even in the ballpark of what most consumers can afford.
But in a few years we might be able to have LLMs running on our phones that work just as well if not better. Of couse as you mention the LLMs running on large servers might still be much more powerfull, but the local ones might be powerfull enough.
Privacy, security, latency, offline availability, access to local data and services running on the device, just to name a few.
I think that the models will evolve and grow as more powerful compute/hardware comes out.
You may be able to run scaled down n versions of what state of the art now, but by then the giant models will have grown in size and in required compute.
The 6 year old models will be retro computingish.
Somewhat like how you can play 6 year old games on a new powerful PC but by then the new huge games will no longer play well on your Old mach
That's why I think these private companies will have the best AIs for many decades.
Visualized in a chart with star-history: https://star-history.com/#nilsherzig/LLocalSearch
> Useful for searching through added files and websites. Search for keywords in the text not whole questions, avoid relative words like "yesterday" think about what could be in the text. > The input to this tool will be run against a vector db. The top results will be returned as json.
Presumably each clarification is an attempt to fix a bug experienced by the developer, except the fix is in English not in Go.
also love your last commit
>fix: copilot is stupid and i should not blindly trust it
>https://github.com/nilsherzig/LLocalSearch/commit/9f45e24f15...
Everything wrong with code gen in a nutshell
Good to see more of these non API key products being built (connected to local llms)
I can only encourage other makers to post their projects on HN and put them out into the world.
I too use local 7b open-hermes and it's really good.
Q5 is minimum.
https://huggingface.co/NousResearch/Hermes-2-Pro-Mistral-7B-...
If you have the time, could you explain what you mean by "Q5 is minimum"? Did you determine that by trying the different models and finding this one is best, or did someone else do that evaluation, or is that just generally accepted knowledge? Sorry, I find this whole ecosystem quite confusing still, but I'm very new and that's not your problem.
If you're RAM constrained, you'll also have to make trade-offs about the context length. e.g. you could have 8 GB RAM and a Q5 quant with shorter context, vs Q3 with longer, etc.
I ran this originally on a M1 with 32GB, I run this on an Air M2 with 16GB (and mac mini M2 32GB), no problem.
I use llama.cpp with a SwiftUI interface (my own), all native, no scripts python/js/web.
7b is obviously less capable but the instant response makes it worth exploring. It's very useful as a Google search replacement that is instantly more valuable, for general questions, than dealing with the hellscape of blog spam ruling Google atm.
Note, for my complex code queries at $dayjob where time is of the essence, I still use GPT4 plus, which is still unmatched imho, without running special hardware at least.
[0] https://huggingface.co/TheBloke/OpenHermes-2.5-Mistral-7B-GG...
You built this in your spare time?
The following things jump out to me:
- How much a hype cycle invites insane amounts of money - How trash the entire VC world is during a hype cycle - What an amazing thing ingenuity and passion are
Great job!
“Needs tool usage” and “found the answer” blocks in your infra, how are these decisions made?
Looking at the demo, it takes a little time to return results, from the search, vector storage and vector db retrieval, which step takes the most time?
Die LLM makes these decisions on its own. If it writes a message which contains a tool call (Action: Web search Action Input: weight of a llama) the matching function will be executed and the response returned to the LLM. It's basically chatting with the tool.
You can toggle the log viewer on the top right, to get more detail on what it's doing and what is taking time. Timing depends on multiple things: - the size of the top n articles (generating embeddings for them takes some time) - the amount of matching vector DB responses (reading them takes some time)
You mean the? The German is bleeding through haha
They do train their own models now, but for about a year they just forwarded calls to models like gpt3.5T. You still have the option to use models not trained by perplexity.
C.ai and Pi comes to mind
I must say, though, that they are doing a commendable job integrating sources like YouTube and Reddit. These platforms benefit from special preprocessing and indeed add value.
> According to the sources provided, Chrome on iOS is not powered by Safari. Google's Chrome uses the Blink engine, while Safari uses the WebKit engine.
I find it amusing how when people show off their LLM projects their examples are always of it failing, and providing a bad answer.
Besides i think the following sentences arent wrong? Its just a 7b model give it some slack haha
What would be the best self hosted option to build sort of a textual AI assistant into your app? Preferably something that I can train myself over time with domain knowledge.
I'd start with "librechat" and mistral, so far that's one of the best chat interfaces and has good support for self hosting. For the actual model runner, ollama seems to be the way to go.
I believe it's built on "langchain", so you can switch to that when it makes sense to. When you've tested all your queries and setup with librechat, know that librechat is a wrapper around "langchain".
I'd start by testing the workflow in librechat, and if librechat's API doesn't do what you want, well I've always found fastAPI pleasant to work with.
---
Less for your use case, and more in-general. I've been assessing a lot of LLM interfaces lately, and the weird porn community has some really powerful and flexible interfaces. With sillytavern you can set up multiple agents, have one agent program, another agent critique, and a third asses it for security concerns. This kind of feedback can help catch a lot of LLM mistakes. You can also go back and edit the LLM's response, which can really help. If you go back and edit an LLM message to fix code or change variable names, it will tend to stick with those decisions. But those interfaces are still very much optimized for "Role playing".
Recommend keeping an eye on https://www.reddit.com/r/LocalLLaMA/
Just taking a guess, but I wouldn't expect more than a couple tokens (more or less like syllables) per second. Which is probably to slow, since it has to read a couple thousand per search result.
It's hard to provide minimum requirements, since there are so many edge cases.
Btw. This is the first time I hear about Perplexity, which after 10 minutes of experimentation, looks like a worse clone of Phind.
If you could tell me more about your goals, I can probably provide a more narrow answer :)
Extremely impressive, cannot wait to actually implement this on my M2Pro (mac).
I see this but what search engine lets you get results in json for free?
I will try my best to catch up with everyone <3