Creating a search engine for fun and because Google sucks
vincents.dev
vincents.dev
I'm convinced that's a lot of what passes for software development in the 2020s: Make a splash screen and display some branding on top of older, robust, actually-complex projects.
Perhaps they would like startpage.com?
If that was the case then their description is completely off, as what they are doing is not creating a search engine but filtering results from other search engines.
I mean, just because you pipe your output to grep it does not mean you created an entirely new app.
Please provide evidence for such a claim.
https://www.startpage.com/en/privacy-policy?t=device
"Our search result pages may include a small number of clearly labeled "sponsored links," which generate revenue and cover our operational costs. Those links are retrieved from platforms such as Google AdSense. In order to enable the prevention of click fraud, some non-identifying system information is shared, but because we never share personal information or information that could uniquely identify you, the ads we display are not connected to any individual user."
Still, there are ads.
Alternatives while probably fine in America, suck in the Nordics. I think people forget just how much search traffic happens in this category.
[1] - https://docs.searxng.org [2] - https://github.com/ItzCrazyKns/Perplexica
It's a metasearch engine that can query multiple search providers at once, including google, so you're not missing out on the good results you expect. Pick an instance at https://searx.space/ and tell your friends!
I imagine I could do it on consumer hardware for less than $10-20k.
Perhaps common crawl has done much of the heavy lifting already and I just have an indexing task.
I'm on a single ~$17k server now, probably utilizing about 20% of it's capacity.
High level language with the ability to make downcalls to a low level language is probably a sane compromise.
You can register with cloudflare as a search engine crawler and they largely will let your traffic through. If you get blocked by individual sites, you can usually just email them explaining the situation and get unblocked.
I'm having very little issues with this.
The sales pitch is simple, you download the crawler, point it at your own blog index it and build an index of pages that your blog links to. Then index the pages those pages link to. etc You simply crank up the depth whenever you like it.
If you are a half decent blogger you have articles that link to most of the important websites that fit the subject of your blog.
You put a search box/page on your website that connects to your desktop client, your visitors can search with options:
- articles on this blog,
- related pages you've linked to,
- 1-5 depth to broaden the topical search (but less related articles)
- search other instances
It scales so well because searching your own blog is the most important, linked pages is pretty nice to have, deeper crawls are still useful but much less important and searching other instances, the anti climax if you like, is great but the least important.
The most crappy hardware can do 50 000 per day, if you run it slowly in the background [say] 100 pages on average per day is still 36 500 every year.
More usual is to be excited about the new found tool and run it for a few hours the first day. You are initially shocked how useful it is. The next day you crawl a few more pages until you get bored with it. You look again after a while and do one more good crawl. Few years later and you have an oddly large index.
You might want to run it automatically when your rss updates.
If you use it once in a while it is easy to ban some instances full of spam.
YaCy checks all results returned by other nodes by fetching the html and looking for the keywords on the page. This worked well. A very stale index may reflect poorly on the node but it may also be full of material that is important to you.
You would get crusty results at times but this is a feature not a bug. There is no man behind the curtain who is the big decider what you may and may not look at.
If your client is not running the search box/page on your blog only does p2p but it is likely able to still search your domain. What is a lot of posts for a blog is not a lot for a crawler. You can glue all kinds of products onto this. Besides a db YaCy keeps the full text of all pages crawled but only the text. If users want a feature that cant be done for free you can sell it to them. If someone has a website that is hard to index they can customize their crawler themselves or pay to have it done.
If you want to throw money at it and have a blog search engine you can limit the results by things that have an rss or atom feed.
I visited the page on mobile and I have to scroll horizontally back and forth before I can read?
That is not fun at all.
So, until a better one comes along, Google's search is fine for me.
If those two don't apply Google is often pretty much useless, just giving you low quality/ AI spam blog posts, useless ad-ridden product comparison sites and marketing pages.
There's a reason why adding "reddit" to Google searches has almost become a meme now.
Why do people always jump on 'umm Google bad' trends? It's all well and good to say it, but as someone who hasn't ever experienced this (apparently common) problem before, I would like to see to believe.
I still miss the ability to use "".
Don't get me wrong, I still find anything I search for, on page 2-N
I switched to duckduckgo about 5 years ago and look back occasionally.
DDG simply works better for me. I find what I search for usually on page 1.
Just to put things into context: I don't care about google, meaning I have no beef with them. I used google search for many years until I switched the default engine, simply because the search results for how I search got worse over time.
It'd be so easy to filter I wonder why they/Microsoft don't bother. Oh, wait...