SearchHut
searchhut.org
searchhut.org
I wonder if simply parsing all of Wikipedia (dumps are available, and so are parsers capable of handling them) and building a list of all domains used in external links would do the trick. Wikipedia already has substantial quality control mechanisms, and the resulting list should be essentially free of blog spam and other low-quality content. Wikipedia also maintains "official website" links for many topics, which can be obtained by parsing the infobox from the article.
Full disclosure: I'm the CEO.
what is stopping companies/users from abusing/gaming on this system with bots?
Users vote on their own preferences... not on other users' preferences. so i dont see how it will be gamed unless we start incorporating user preferences into global ones - which we would only do if we thought it worked better and isn't being gamed.
[1]: http://elempresario.mx/sites/default/files/scith/specific-he...
I've been considering building my own search engine for a while for my niche topic which has <50 websites and blogs on the web.
I can't tell how useful this will be but it'd be fun to give it a go.
I mean the edit history is public and there's plenty of people that actually pay attention to edits and the like so they would be found out soon enough, but still.
I'm sure this is an ongoing discussion when e.g. political figures' pages are protected as well - who becomes the gatekeeper, and what is their political angle?
Considering that Wikipedia has become a pillar of most scientific work (imagine writing a math or computer science paper without Wikipedia – utterly unthinkable), it's safe to say that knowledgeable people have collectively decided that its quality control mechanisms are "good enough", or at least better than those of any other resource of comparable depth and breadth.
And that puts Wikipedia's link pool leaps and bounds ahead of whatever dark magic current search engines are using, which mostly seems to be "funnel the entire web through some (easily gamed) heuristic".
Thanks for all the work Drew, I hope you guys manage to come to a conclusion that you are satisfied with!
[1]: https://paste.sr.ht/~sircmpwn/048293268d4ed4254659c3cd6abe67...
For what it's worth, my search engine got prematurely "announced" for me on HN as well, while likewise hilariously janky. I don't think the launch is the end of the world. (I guess I had the benefit of serving comically bizarre results when the queries failed, so it got some love for that)
The bigger struggle is, because a search box is so ambiguous, people tend to have very high expectations for what it can do. Many people just assume it's exactly like Google. It's something a lot of indie search engine developers struggle with. Even if your tool is (or potentially could be) very useful, how can you make people understand how to use it when it looks like something else? Design wise it's a real tough nut to crack, a capital H Hard UX problem.
Startups and side projects are messy and sometimes things don't go as planned. Contracts get canceled, DoS takes down your homepage when you launch losing all those free leads, people leak new features and your sixth deployment erases most of the production database.
There are a lot of great ideas that start out as bad as the first release of thefacebook, AirBnB, Twitch and Youtube. Still, they iterate on these wonky, almost-working sites and end up making something great.
YC pushes this idea constantly; put something in front of people and iterate. Drew was following that advice and I applaud him. https://www.youtube.com/c/ycombinator/videos
Still, most projects aren't a search engine. I see people put high expectations on how things will go and often it's just really hard to realize some of those hopes.
Sometimes you just have to take what you can get and iterate. Don't give up.
So, I started rolling out new features of our search engine on Twitter, Slack, etc instead of here.
[edit: nevermind, found it easy enough. Seems to be Go and PostgreSQL with the RUM extension]
I misinterpreted a don't-share-yet announcement to mean the announcement post only and not the entire software and announcement. I don't mean this as an excuse; that's just the context.
So this is out before Drew et. al. intended it be hence some 404s and so forth as commented by Drew in this very thread here: https://news.ycombinator.com/item?id=32105407
I let my excitement get the better of me this time and I hope people revisit SearchHut in a week or so after these quirks are resolved.
You (or Drew or someone else) can resubmit it later with a bogus query string to skip HNs dupe checker.
I went ahead and polished up the announcement for an early release:
https://sourcehut.org/blog/2022-07-15-searchhut/
Let me know if you have any questions!
e.g. "Any websites engaging in SEO spam are rejected from the index" - how is determined whether something is SEO spam or not? More clarification of whats allowed/not allowed would be nice!
https://searchhut.org/docs/docs/webadmins/requirements/
And there's some advice for web masters on ways to improve your site's ranking without running afoul of this rule:
https://searchhut.org/docs/docs/webadmins/recommendations/
But ultimately, it's subjective, and a judgement call will be made. If it's minor you might get a warning, if it's blatant then you'll just get de-listed.
e.g opencrawl, internet-archive, archiveteam
It strikes me the resources to crawl, update, and manage/index data is a common problem.
If you happen to be cloud hosting this, and if you do not have a global rate limit, implement one ASAP!
Several independent search engines have been hit hard by a botnet soon after they got attention, both mine and wiby.me, and I think a few others. I've had 10-12 QPS, sustained load for weeks after weeks from a rotating set of mostly eastern european IPs.
It's fine if this is on your own infrastructure, but on the cloud, you'll be racking up bills like crazy from something like that :-/
I don't recall if it supported SourceForge and GitHub (2008) but it certainly included gzipped tarballs which were popular and prevalent at the time.
I tried a couple test queries:
> lambda decay to function pointer c++
I get some FSF pages and the wikipedia for Helium?
> std function
I get... tons of Rust docs?
> std function c++
All rust docs? The wikipedia page for C++??
Interesting idea, but this seems like it would be the primary failure mode for an idea like this: as soon as you are researching outside of the curator's specializations, it doesn't have what you're looking for. Yet these results would both be fixed simply by adding cppreference.com to the index. Let's try and give it a real challenge:
> How to define systemverilog interface
And as I might expect, I get wikipedia pages. For "Verilog", for "System on a Chip" and for "Mixin".
1st google result:
> An Interface is a way to encapsulate signals into a block...
Working as expected
> Notice! This product is experimental and incomplete. User beware!
But in reality, if you know where is the answer that you are looking for, why would you use that search engine.
I use DDG and if I want to search on the Scala docs, I use "!scala whatever I'm searching" instead of just search with DDG.
There will (soon) be a form to request new domains are added to the index, so if there are any sites you want indexed which are outside of my personal expertise, then you'll be able to request them.
It is also meant to be very simple to run in the case you want to index your own category of sites. For instance cooking content is specifically not indexed but if _you_ wanted to you can spin up an instance and index cooking sites yourself.
I watched him crank out a prototype of this in Go in about three hours.
It really feels like the old days at sr.ht. It’s fun again.
How can I watch?
One meta-thought, I think projects like this are surfacing something interesting: The underlying technology to make a pretty good search engine is no longer especially difficult for programmers or for server. This is potentially a very good thing, as it means the end of the Google era.
I can imagine a future that is almost a blast from the past, where there are a lot of different search engines, those engines are curated differently, and while none of them index the entire Internet, that’s what makes them valuable and better than Google (which I think cannot defeat spam).
I’m trying to think of a historical parallel, Where are some service used to be very difficult to provide and therefore could only effectively be done by a single natural monopoly, but technology progressed and opened up the playing field, breaking the monopoly. Television has some similarities. Perhaps radio vs podcasting. What others?
I don't mean to be conspiratorial, I'm sure there are good intentions behind this, the consequence however is effectively locking in Google as the default gateway for the Internet.
But it passes tests that are very important for me:
1) It's fully accessible by Tor. No CAPTCHAs or "We don't serve your kind in here" messages.
2) It works in a text browser without JavaScript and renders in a sensible way without style requirements.
10/10 for accessibility. Something Google and other search engines could learn from.
I can't tell whether this is a neglected Google product that they were going to refresh but lost interest in, or something that is undergoing a breath of fresh air.
As you say, I was able to add a list of domains and get some pretty decent results from it. The UI makes me feel like Google are not interested in making it a truly successful product, though.
edit: Looks like this OSS project was launched and cancelled in a single day.
Touché :-)
However the idea of a federated search is some measure of protection against that. If it ever happens.
- search queries are performed directly from the clients computer so can't protect their privacy (since Custom Search JSON API have a daily limit of 10k queries)
- forced to use javascript, and the way it's implemented makes it difficult if not impossible to do even basic things like the loading animation cards
- ads are loaded from an iframe so you can't do any styling (except extremely limited options that they make available in their settings, but no matter what then it will be very ugly if you want to have a light/dark theme)
But there are of course many benefits as well, such as it being 'free' (Bing is ridiculously expensive IMO, and feels impossible to join their ad network to offset the costs.. which might explain why you see countless Bing proxies shut down after a few months) and search results are no doubt better than the ones you'd get from Bing.
Which raises the question: does archive.org offer their Wayback Machine index for download anywhere? Technically, why should anyone go through the trouble of crawling the web if archive.org has been doing it for years, and likely has one of the best indexes around? I've seen some 3rd-party downloaders for specific sites, but I'd like the full thing. Yes, I realize it's probably petabytes of data, but maybe it could be trimmed down to just the most recent crawls.
If there was a way of having that index locally, it would make a very powerful search engine with a tool like yours.
> List of accidents and incidents involving commercial aircraft
I think it's similar to how Google's search works internally, though I doubt the separation is based on a list of domain (as in DNS) names. IIRC they have a set of search modules, and what they return (and how fast they return it) all gets mixed in to the search results according to some weighting. Right below the ads.
If you look at a search system that way, it's easy enough to add modules that do things like search only wikipedia, and display those results in a separate box (like DDG), or parse out currency conversion requests, and display those up top based on some API (like Google). etc
It is possible for a site's results to be of different quality: maybe one article about MySQL is not so informative, and an article about Python on the same site is a reference.
The search engine operated by the author is unlikely to acknowledge that.
=> 404
https://paste.sr.ht/~sircmpwn/0cab5e3137c2c2077b5aabf9e2fc8d...
It was intended to be larger prior to launch. Here's some other domains I want to index:
https://paste.sr.ht/~sircmpwn/84d052f14a9a282698b5e5f7a9d9d9...
Cool.
Postgres as backend is maybe not the best choice and there are already many sites that index specific pages and take suggestions. The hard part is getting relevant results when having a large index.
Still thank you for a new web search.
:).
Seems like it would be an interesting experiment to see what the results would be, indexing only the content / meta tags of “index.html”.
!url <keyterms> |synthesize
I also wrote a screenshot extension for Chrome that lets you save a page when you find it interesting. The site is definitely not "done" but it's usable if you want to try it. Some info in help and in commands is inaccurate/broken, so it is what it is for now.
It does the !google <search term> and !ddg <search term> thing to find pages to save to the index. There are a bunch of other commands I added, and there's an ability for others to write commands and submit them to a Github repo: https://github.com/kordless/mitta-community
!xkcd was fun to write. It shows comics. The rest of the commands can be viewed from !help or just !<tab>
I've been working on pivoting the site to do prompt management for GPT-3 developers and have been kicking around Open Sourcing the other version for use as a personal search engine for bookmarked pages.
- Beltalowda – no results (for reference: it's a term to refer to "people from the [asteroid] belt" used in the The Expanse books and TV series).
- The Expanse – bunch of results, but none are what I'm looking for (the TV series or books). It looks like it may drop the "the" in there?
- Star Trek – a bunch of results, but ordered very curiously; the first is the Wikipedia page for "Star Trek Star Fleet Technical Manual", and lots of pages like "Weapons in Star Trek" and such.
- NGC 3623 – lists "Messier object" and "Messier 65", in that order, which is somewhat wrong as NGC 3623 refers to Messier 65 specifically.
- NGC3623 (same as previous, but without a space) – no results.
- vim map key – pretty useless results, most of which have no bearing on Vim at all, much less mapping keys in Vim.
- python print list – the same; The Go type parameters proposal is the first result; automake the second, etc.
Conclusion: "this product is experimental and incomplete" is an understatement.
Give
Up
This requires less paying attention to negative emotion and more “water off my back”.
Tweak it. tweak it some more.
Focus on the goal, notably one tiny sub-goal at a time.
Good luck, entrepreneurial spirit is a tough beast to attain.
Whatever you stake on,
NEVER
EVER
GIVE
UP
Rather than maintaining a whole separate index for myself, I'd love to self-host an instance of this, only indexing sites that aren't in the main index, and then falling back to the main index / merging it with my index to answer queries. I wonder how easy that would be with the current architecture.
Does Sourcegraph index Sourceht projects? It is a proper code search engine and very good.
In other news we now index 87k Rust packages on crates.io
https://sourcegraph.com/search?q=context:global+repo:%5Ecrat...
Wikipedia: List of Search engines
Drew's blog: We can do better than DuckDuckGo (perhaps the impetus for this project)
Wikipedia: List of free and open source projects
Wikipedia: Internet Privacy
Given it’s not even ready for release either.
The best you could have hoped for is a GitHub link, but GitHub isn’t being crawled right now.
So I’m not sure what you’re getting at. Your expectations for an alpha level software that wasn’t supposed to be announced is far too high.
Are there already plans to expand it as a service? E.g. subreddits could maintain their preferred lists of domains.
Also, selfish plug I think it would be cool if you added Hackernoon to that list.
How many requests per minute / hour are acceptable?
SearchHut: The first result is django which is not the most popular web server.
Google: Shows an answer box with the market share of various web servers.
The part that Google seem to have unfortunately skimmed over is that the answers need to be relevant, exact & correct.
But I'm glad people are trying to build alternatives. I'd love a search engine that ignores sites with antipatterns like required registration for any kind of usage, and this is the first step.
[citation needed]
The quality of the results right now are not very high, and in theory I don't understand why one would believe a search engine with a hand picked set of domains would be expected to outcompete a search engine that can crawl the entire web and determines reputation by itself. This also ignores the fact that a lot of domains have a mix of high quality content and low quality content, for example twitter or medium. If you are going to rely on domain-level reputation then your search engine is going to be way behind the search engines that can judge content more specifically, which is all of the other search engines.
If you were to tell me curated domains is just a bootstrapping method and as the search engine evolves it will change, fine, but right now the search engine is so simplistic that the theory of how it might be good is really the only point. And if that underlying theory is dubious, and the infrastructure is simplistic and obviously won't scale, then I don't know what is interesting or novel about this right now. Doesn't seem worthy of reaching the top of HN.
Then why do Google and DuckDuckGo return 90% garbage for most queries?
"All of the other search engines" have completely failed to keep pages from the results that are not only low-quality, but outright spam.
Instead, the top result was W3Schools. Then came the Python docs, then 5 pages somewhere between blogspam and poor-quality tutorials. Then a ReadTheDocs page dating to 2015. And that was it. No more official Python resources, no StackOverflow. In the middle of the results some worthless "Google Q&A" dropdowns that lead to more garbage quality content.
So for this query, using my definition of "garbage", the "garbage percentage" is somewhere between 80% and 90+%, depending on how many Q&A dropdowns you waste your time opening.
So at least some official Python docs are indexed.
You could instead search for "python string" to find more information about python strings.
Even then the very first result for Python str is actually relevant for me (Python documentation about built in types.
It does have it:
$ python3 -c 'print(type(""))'
<class 'str'>If you can give me a list of 10 normal-ish queries where 9 out of the first 10 results on Google or DDG are "garbage", then I'll concede your point.
I think you are creating an impossible standard for search engines, then using it to deem the current ones as failures. While at the same time ignoring that this new search engine is, as present, unusable with no realistic argument for why it might eventually be better.
Seems like your expectations are misplaced. Being at top of HN is not an indicator of quality, just interest.
Because SEO manipulation is a well developed field, ensuring that the search engines trying to determine reputation automatically will (and does) end up with bad results.
This makes me think of a possible approach. Curate a giant set of domains that almost exclusively host high quality content. Crawl said domains. Use all of the crawled data as a training set to create a model with which to ascertain the quality of random Web pages from other domains. Then spider everything and run it against the model.
If you don’t want something disclosed , don’t disclose it.
Only way for three people to keep a secret is if two of them are dead.
A thing is in the world. Let it be in the world. Harness the collective power and focus it into a force multiplier.
Or don’t.
Develop a thick skin or don’t read the comments lol!
He chose to disclose it to a few people. Word spreads. That’s what happens.
Execute NDAs and have a security program if you don’t want stuff getting out.