There was a time when search engines were a thing, and it seems they still are
boston.conman.org
boston.conman.org
By that metric there are no good search engines at the moment and the older the pages the worse this effect gets. It's really nice to see Google do lots of 'moonshots' and interesting tech demos but I'd be far happier if they fixed search and kept their focus on that.
If a page doesn't show up in either Google or Bing for sensible queries then that page effectively ceases to exist. The perverse incentive that these companies have to avoid you going to a page with relevant results as long as you spend more time on pages with their advertising on it ensures that more and more content will end up missing in action.
I'd be happy to pay for a search engine that:
- actually really works
- also allows you to search past page 10
- has a working API with reasonable limits
Funny, Google has some perverse incentives here: it might be nice to have a good "history search" built into a browser, but as a search engine provider they won't build it into Chrome.
Now that we live in the future I guess you need never "clear your cache" except maybe for privacy reasons. You could keep full page text for just about any site you visit (so long as the authors don't consider their site a "web app" I guess.)
I miss this feature all the time on other browsers.
Chrome://history
I get history results when typing into the address bar.
You can craft a google search with greater specificity but it's very difficult to obtain the sort of search behavior that google used to have. Now google treats your search terms as sort of a grab bag, it mutates them into a cloud of synonyms and related words, then it picks a subset of the grab bag that it decides is valuable and gives you results that are tuned by about a zillion arcane heuristics. This works great for giving you "magically" accurate answers for the most common search queries. It works terribly for giving you highly specific answers to highly specific queries. The way google used to work was by providing results that matched all of your search terms, and being smart enough to include different variations of each word but not vaguely related words. That sometimes made it hard to find the right thing if you didn't get the right words but now we're in a state where you can't find the right thing even if you do have all the right words.
Duckduckgo is a metasearch engine, technically, but mostly it delegates to Bing.
As far as I can tell, there are only two and a half real search engines that still exist: Bing, Google, and Wolfram Alpha. (I count Alpha as a half because it's not really what most people are looking for.) I'm curious if anyone else knows of other real search engines still in existence.
Yahoo was running user studies where they would put Google results and Yahoo results side by side but switch the branding; while Yahoo results were ranked better than Google for most of the tested queries, results with Google branding ranked better than with Yahoo branding, regardless of whose results they were.
The plan was to just use Google, but the DOJ (or FTC?) put out guidance that that would be anti-competitive, so Bing was it. This might have worked out anyway, but the expected cost savings from outsourcing search didn't actually happen that I saw, but I left in late 2011, and stopped following closely after that. Web search was also linked with search ads, which Bing did poorly at too.
One tough thing is there isn't one search quality metric. It's important to have the search results page look good with its snippets, and another thing to have people actually look at the linked pages and compare the usefulness of the linked pages.
Common vs. uncommon searches are also important. It's not difficult to write a search engine that badly over-fits on the most common searches. However, for market share, it's important to do well enough on the common searches that users don't leave, and do well enough on tough long-tail searches that you pick up users that leave other search engines on tough queries. The idea is to be pretty good at the common searches, but the best at the kinds of searches that cause people to try other search engines. Naive frequency-weighted metrics will get this totally wrong.
It's also more important to get useful information in the first 2 or 3 links. If Google links to the second-best link at result #1 and puts the best link off the first results page, but Yahoo puts the best link down at #7 and second-best at #8, the user may lose interest before following a really good link.
I don't think Google took the union of front-page search results between two competitors and asked humans to hand-order the (up to 40) pages for how well they fit the query. But, that seems like a good way to test the actual usefulness of search results. You'd probably especially want to keep track of the percentage of the top 3 search results that were filled by top-5 (guessing at 5) useful links.
Anyway, inside Google it was well-known that Yahoo was the competitor to worry about in terms of search quality.
Unless something has changed it seems bing still gets at least 51% of traffic.
Or a series of unique queries could leak private information.
As long as DDG are doing it properly (and I believe they are), Bing would only learn that the contents of each individual query are associated together, they would learn nothing about which other queries were performed by the same user.
Homomorphic encryption might do the trick (?), but it's too slow at the moment.
Maybe I'm missing something obvious here, but how is that any different from Google or DuckDuckGo seeing the same spike?
I just think there is a point to be made here. Even generally it's often opaque what third parties have what data and I don't really think GDPR has fixed that. It's surprising for people the Bing might have the contents of their DDG search history, somewhere in the huge dataset of DDG searches that pass through.
Also they might not want to help improve Bing search but I'm guessing they do inadvertently?
I wonder if they use mostly Google for the backend.
- origination date of the website
- number of theme changes
- whether ads are present
- estimated data usage / load time
- general website size
- type of website - personal, company, etc.
- last update time
- presence of javascript
- browser compatibility (well, feature usage detection & browser compatibility inference from that).
Mostly this is just me missing the websites of the early 2000's, and trying to figure out a way to rediscover them.
And I'd probably want content on top of this. (Edit: e.g. search by topics)
Lastly, it'd be nice to restrict things to sub-genres, but I'm not sure. E.g. when I'm doing a search I'd love to reference things related to micro-controllers, and so maybe I'd put in Arduino to get into the realm. Sort of what like google does for you without telling you. (tailoring your searches by some magic context).
A man can dream...
Edit 2: Search engines these days seem to be answer engines, I want a research engine.
I think of it as a form (as opposed to content) first search engine, less for precise searching and more for general topic exploring.
Apparently somethingawful and ebaumsworld are still around in some form, and Slashdot of course. The thing is, these sites have largely been replaced by better versions of themselves. That’s resulted in a lot of centralization into a few sites like reddit, which is a combination aggregator and blogging platform for people who are too embarrassed to attach their real name to what they write, which is apparently a good share of the population. Then there are sites like YouTube, LiveLeak, Facebook, that just offer something that no one could or did in the 2000s. And with mobile and apps, there’s a level of engagement that doesn’t leave much room for a thousand little sites with quirky, regular, custom content.
Right now Google thinks it knows what you want and when you search for things, it returns the same few sites (mostly). You used to come across people's personal sites into which they poured their soul. And while those exist less frequently now, I bet they still exist.
They also -- despite being one of the first ever search engines -- didn't do their own search in 2008. They outsourced to Yahoo. Though there was an effort at the time to become a search engine again. I don't know if anything came of it.
Edit:
It's hilarious they labeled Lycos as...
> Lycos—is still around!
Because even at the time I worked their the number one response I got from people when I told them that was "They still exist?"
But the short summary: I loved it. It was a really fun company with a lot of great people. And we got to launch some really great products. Most of whom were let go during the great recession (myself included) but it was fun while it lasted.
On the flip side, every time we launched a product the news media treated it as a novelty instead of a serious thing. Which was insanely frustrating. Some of our tech was way ahead of its time.
Also, I can think of a couple specialized engines that do exist: Google Scholar and Shodan. There are probably more I'm unfamiliar with.
You can get most of the way to this with the verbatim option, but I think they make it difficult to make the default.
> By using the Verbatim tool Google will not make the following changes: • Personalizing your search using websites you have visited before; • Including synonyms of your search terms; • Automatic spelling corrections; • Searching for words with the same stem e.g. “Shopping” when searched for “shop”; • Finding results that match similar terms to those in your query.
- alternativeto.net and similar - Google Scholar / Semantic Scholar - Every sandboxed social network (Facebook, Twitter, Tumblr) - A variety of Instagram searching sites - Alternative App stores for Android
And there's room for plenty more sites like this.
Attempting to return a good response for anything in the search box is a bottomless problem of unclear utility. Being more focused makes the work easier and the value for the user clearer.
You'd think that a monopoly could just break instantly if all it required was typing in a different URL.
But modern search engines are reliant on machine learning on mind-bogglingly enormous troves of real human interaction data.
If you truly outsmart Google by inventing a better mousetrap, it's worth fuck-all. Your solution will probably require more usage data than you will ever be able to collect, because nobody will use your search engine while it still produces poor results.
Looks like german-english bilingual logistics professionals are looking for truck parts vendors, while teenage Americans are looking for hidden locations of laser focusing and enhancement devices within the video game Wolfenstein The New Order.
Do you envision a second route to that answer?
I think that's what you meant by partitioned results.
Google computes this internally but I've never seen them use it for anything other than having diversity in their top 10 results.
As I said, Google is pretty conservative about this and other entity-based features, so they definitely wouldn't do it got something like "lkw attachment". My question was whether Izik triggered this feature in such cases or not.
Izik would show you film-related website results for the film category.
Setup grants for a domestic internet fund to incubate tech, ban competitors and blamo!
Works miracles in China.
The same underpinning is in many Francophone nations.
It is true, however that many companies don't bother to localize.
It seems they are content with raiding the Canadian film board.
Their goal is ad clickthrough, not accurate search results.
Let's say that someone creates a better search engine, how would people actually use it? I'll tell you how most people who bothered to use it would integrate into their routines. Firstly they would still use google for everything day to day. They would still use google maps for directions. They would still use gmail for mail. They would still use google search as the semantic front-end for their browsing. Only after they performed a search using google that produced unsatisfactory results would they then pull out the better search engine and make use of it for that one isolated search. And that's the problem, because that scenario is very hard for the better search engine maker to monetize while google would continue reaping the major monitization haul for the vast majority of search uses for that user. And going from zero to completely integrating into a user's experience in the same way that google does now is not a realistic prospect for most startups.
As always I can shill for my utzoo early Usenet search that combines AltaVista desktop + a few hacks to make a specialized search
http://altavista.superglobalmegacorp.com/altavista
I know it's niche, but it's great for anything historical from 1981 to early 1991 on the internet
I'm working on a project now which has indexed billions of pages and answers queries similar to a web search engine like Google: https://www.AtSign.co/
The only difference is that it's a keyword + location based business contact information engine but operates on the same principles as a real web search engine client.
We're a small team and it would have been unthinkable even a years back to launch something of this scale effectively ... But here we are! Amazing space to be in right now
Also, ours is keyword based. We index the site similar to how Google does. So you can get very specific company matches and then export to CSV.
Regardless, I just looked at our results for tech support in CT and I agree we need to work harder on our results, but comparing to Google, they only had tech support jobs (not even business listings)... which makes sense in their product use case.
I can see what you're saying and Google is by far, the industry benchmark, but it's also difficult to compare results sometimes ... it's like Apples and Oranges.
But it can get really complicated, for example, sometimes there just aren't relevant documents in the index in the first place ... so, you can't really blame your ranking factors too much. The opposite can happen too, where a word occurs too frequently in which case you might resort to other kinds of ranking factors (most notably pagerank).
In our case, we're focused on broadening our state/country level coverage right now for keywords (more listings), then we're going to focus on making sure our location accuracy is a lot better (it needs work). Overtime, you should get the results you're expecting more often :)
Most of these used BOSS (aka Yahoo!'s old build your own search service API) which was served off Bing as its index, although Google has started paying more and more people to send their search traffic to Google.
Bing charges $7/thousand [1] for their "quality" searches and $3/thousand for their so-so searches (not as current, the index doesn't go as deep, this used to be what BOSS called until they turned it off in 2016).
That $7/thousand lets you give them up to 250 queries per second. For reference that is about 1-5M uniques per day. It looks like 21M searches a day but for English most of the searches come during the day from Europe and The US so you're really only going to do 10 - 15M searches per day at that rate. If you are clever you can cache results so for the same search you can just re-use the cache rather than paying for another result. This is nominally frowned upon but hard to defend against. If you manage to make a deal with a phone supplier to be the 'standard' search engine a lot of queries will just be 'facebook' or 'reddit' so you don't really need to actually query those. You will want to find some ad networks to provide you ads. Bing will do that too, but you will quickly figure out that if you could make money reselling Bing results with Bing ads, that they could do that too so you've find the margins pretty thin and negative at times. You'll have to pay for a machine that is taking those queries, calling out to what ever ad networks you want, and then filling out a results page (SERP) and sending it back to the consumer. If you are just fronting Bing or Yandex that is pretty straight forward to do with an nginx server on an AWS "large" instance.
If you negotiate well and market well you can be a dogpile or a startpage with some schtick that makes you different than just going to Google or Bing. The more privacy you afford the clients the more margin you give up (because you can't sell that information as well).
Bottom line is that its a hard way to make a living.
[1] https://azure.microsoft.com/en-us/pricing/details/cognitive-...
And then Google happened ...