The Age of PageRank Is Over
blog.kagi.com
blog.kagi.com
What really excites me is that is that I'm paying them. That sounds odd, but seriously. It's incredibly refreshing to know that the company providing my search results has an incentive to make things better for _me_ and not a legion of advertisers. With Google I can't help thinking about every keypress being logged to optimize sales pitches at me. I just don't feel that with Kagi, because I'm paying them.
Sure, they might be logging every keypress (I don't actually think they are, but you never can tell) but even if they were, I could be reasonably certain they were doing it to retain my subscription, which probably means making my search better, not selling me other stuff.
It's a priveleged position to be in, and the economic argument isn't watertight, but in the "search as a brain extension" space it still _feels_ premium, because it creates trust. And that frees up brain space for other things - like where the hell was that article I was looking for?
Cable television and netflix have made it quite clear that payment does not mean “no ads”, it just means “no ads, yet”.
It still might be a better alternative in the short term, but the moment growth starts to peak, companies follow where the incentives lead them.
I think what's more interesting/concerning/insidious is hidden ad content like product placement. It's getting to the point where personalized product placement could be embedded in shows dynamically.
Take Spotify as another example. I pay them, along with 188 million other people. Does that mean they won't turn around and ask artists, record companies and rights-owning conglomerates for money (or rebates) for putting their stuff in front of me? Of course not. Paying them means they have some interest in pleasing me, but it's far from the only way.
On the off chance that Spotify was being scrupulous about not taking payola in any form, it would be impossible for me to verify. Which is in itself a reason for them to cheat; they don't get much economic benefit from goodwill if no one can actually see them being honest.
This is not a reason to not use Kagi, it's just a reminder of what forces we're up against. Kagi will need an unusual amount of transparency in everything they do, in order to stand a chance in the long term.
And it IS a reason to not get warm fuzzy feelings merely from the fact that you're paying them.
2. Create a new plan at the price of the old plan. With ads.
I'm sorry, but pretending that they're not adding ads to their "regular" pricing plan is semantic at best. This isn't something they dreamed up overnight. They've been planning this for years and increasing their prices accordingly.
I am paying and have no way to disable that visual pollution nonsense.
Extremely exasperating when you are trying to choose something adequate for toddlers and they keep seeing flashy stuff on the screen and saying they want to watch that.
So far. I imagine they will be expanding it over time.
for now....
The agreement could be carefully written by a skilled lawyer to define the things Kagi cannot do, the proof customers must present in order to proceed with a valid lawsuit, and even the maximum damages that the customer may sue for.
In that case, if Kagi was found at some point to be using customer data for these purposes, it could be sued very easily and by many parties.
People are calling for regulation for data privacy. In the meantime, Kagi can create its own regulation it will hold itself to for the benefit of its users, can it not?
No VC startup would do that but if possible would allow trust.
Companies have an incentive to make money off of their customers' data, and technology just means that they can make money at every level (your ISP, your smart TV, the apps you use on it and the websites your browse).
We're so accustomed to entities gathering data about us that it's become part of the background.
Why is it the default that my home is on Google Maps/Google Earth? I should have to opt in, not opt out.
Just because there is "public" information about me as a registered voter or home owner doesn't mean that anyone should be able get a copy of this data, put it online and connect it to other data?
Complaining that cable TV isn’t ad free because you pay for it is like saying all content over the internet isn’t Ad free because you pay for your internet connection.
'It's better to be a millionaire hero than a billionaire asshole' (Ben Elton, Gridlock)
Of course at $10/month even a 1% share of the search engine market would make the company worth billions.
Afaik [1], there's about 5 employees and the revenue only covers server expenses while they're still trying to get more headcount. Not an expert on bootstrapping but I'm pretty sure you don't want to expand faster than your revenue does otherwise you stop being able to make all the decisions for the company.
I think another factor is the paradox that the user who has enough money to be willing to pay to not see another ad is exactly the most valuable user to advertisers— they've pre-qualified themselves on multiple axes.
So there is always going to be an enormous moral hazard here with the ad-driven-freemium model— always a temptation to run an occasional high-value ad to the paying users, or to disguise the ad as a "recommendation" or position it instead as a "sponsorship", or whatever else. Other than users reacting swiftly and decisively against this kind of thing, it's hard to know how else to keep companies away from it.
An angry bird informed me about this :)
Of course this depends on how the supply demand curve looks like for their specific business.
Need more money? Economic model not quite working? Companies face this all the time. Some solve it innovatively, survive, and show everyone else the way forward. Others fold or pivot.
But ad revenue short-circuits the hard process of economic innovation by tempting tech companies to take an easy way out, even when the problem is solvable without it.
(All of this wouldn't really be a problem if the ads business model weren't inherently adversarial against users.)
The reason they advertise is because they can. It doesn't matter how much money they're making now, they can always make more by advertising. Paying money just makes your attention more valuable. The only way to stop them is to make it impossible or illegal to advertise. Blocking ads everywhere and making no exceptions.
EDIT: Their website says that they don’t plan to take VC money, so I have no idea what a SAFE means in this case. The point of a SAFE is that the investor is guaranteed the best terms when a priced round of investment happens.
I agree that on the whole, people will continue to go for the free option, but I'm just glad that some options are appearing to pay, which rebalances the incentives a bit.
Now the issues is not just privacy but making sure it does NOT optimize search results based on prior searches or other facts it knows about me.
Now the big problem is that one would think delivering the exact search results that I am looking for is a good thing. No. It is super dangerous. It can causes me to live in a bubble, blinding me of possible opposing opinions. It is feeding my ignorance.
So there is a goal conflict here. I don't know how to solve it. I wish there would be something like Wikipedia for search engines but even the wiki models has its problems and biases.
This made internet completely unusable for discovery. A big shiny window into another tiny room. Absolute BS.
Recently I discussed it here on HN and was advised to like, dislike, subscribe and so on on youtube (things I rarely have done on my old account) to break out of this bubble. Well, it doesn’t work. It just suggests the content I immediately dislike or ban entire channels and it all settles to the same echo chamber as before.
On the other hand, a lot of the search results bias stems from the search query itself. "Does X cause Y" for example will show results that echo the search, and not show opposite results. Simply because you did not ask the opposite or a more open/unbiased query like "X Y".
On the point about payments, there are some potential approaches to handling this, but they tend to be a bit complex and technical. I"m not sur ethis is really a technical problem though (or at least one where a technical solution is what's needed).
A challenge of taking traditional payments is that the identity of the payer is inherently "linked" (to some extent) to the transaction. You could maybe use a privacy.com type intermediary or similar.
What you could do however, is buy a "kagi voucher" or similar, which would be an RSA-blind-signed token that you generated client-side in JS. You'd send it to the server during the payment transaction, and be given back a signed version of the token, which you unblind client-side, then give a backup copy to the user to save, and set it as a cookie to authenticate payment has been made.
That should, in theory, at least un-link the payment from the user. Ultimately though, a bad actor can still trivially correlate your searches (and payment, if you used same IP or browser fingerprint), and all your searches are linked by a common cookie.
So you decide to create some kind of token fountain, which you can auth to with a blind-signed token, and get 50 "search passes" from. Those could be unique per-search, and just be "bearer tokens" you can use. You'd have to trust that this token fountain doesn't record which search passes were issued to which "blind token".
Ultimately, it feels like this problem then becomes "do you trust your search engine?" - I'm not sure this is a problem that technology needs to solve, as ultimately your search provider is seeing your queries and IP, and most users won't realistically change this. A provider that isn't incentized to act against your interests (i.e. something like Kagi which isn't advertising to you or selling data) seems to address this for most user scenarios.
This way you get anonimity and custom results
I'm sold, literally.
[0]: https://teclis.com/
240 searches cost them over $10.00. They will need to find additional revenue or reduce costs to even reach a break even point.
They will have to raise prices or make you pay per search or sell your data or better yet all three.
If their paying user base grows by 10x, do the searches drop from 4¢ each to 2¢ each?
.0125 per search. The teams rate is twice that.
* one does not need $200-300B ARR to run a search engine; and merely making a ton of money does not mean that the way that one has been making that money will continue forever
* using Kagi, one realizes how much the web has been perverted, in my opinion, by the obsession with the ad model; there are many nice, novel things that Kagi could introduce which would reduce ad revenue for a company like Google, but would be useful and novel for Kagi without loss in revenue (on the contrary — could help); ie google is incentivized to keep people on the site through various irritating means; a search engine that isn’t distracting and gives me what I want — I would be more than glad to pay for (and already do pay Kagi $10/month)
Other thoughts:
* Supposing $10/month (honestly I would consider $20) with 100M paying users; that is $1B MRR which is not bad at all and is more than plenty for a meaningfully sized team with meaningful salaries.
* Just like how lowered energy prices during and after the Industrial Revolution made manufacturing at scale feasible, the incredible amount of high quality OSS and infrastructure these days is making it increasingly feasible to do things Google did 20 years ago — something unthinkable for even the best engineers of that era. Not to mention the relative ease of collecting capital with payment services (let alone VCs, etc.) today — even with a looming recession factored in.
US attention is the most valuable because it is the richest large nation and the largest rich nation.
European attention comes a distant second. Together they have 40% of the global GDP, but likely something like 60-70% of the consumption of non-essential goods in nominal terms. Accordingly, average advertiser cost per clock differs ~10x between the cheapest and most expensive groups of countries.
35 dollar is maybe ok… …. Maybe better than spending 3500$ on … I dont know … stuff is good also ..
They have a full business suite, App Store, hardware sales, fiber internet, wireless cell and internet, and a server rental division.
Yep, but the main problem here is the perverse pressure to grow indefinitely.
And that comes from it being acceptable for public companies to pay no dividends and instead rely on share price increases to deliver a return on investment.
Something that slows down this growth spiral would fix things partially. Require profitable companies to pay a % of the profit as dividends to their shareholders?
I'm a user now, and potentially a customer if I hit their free limit.
Is there some zero-knowledge protocol that could help here? In particular, to establish that a user is currently subscribed, without revealing which subscription they have?
----
I have an idea (more here[1]) that the next addition to our civilization should be a distributed system to evaluate, reward and compensate externalities. Open source software is a great example. It's free, it has massive benefits to users and for people reusing it for all sorts of projects, it's public. Yet we still can't find a great way to fund it. There are donations, and maybe that's all we need in some sufficiently enlightened society, but I think we should be addressing those more directly. And clearly there are technologies which generate friction and other social costs, that could be (differentially) quantified and prices. The same goes for pollution and a miriad of issues.
Laws can help here, but laws tend to lack softness and specificity in my opinion, and laws are strictly punitive. If you're generating a negative externality, and the legal system is well designed, maybe you'll have to compensante for it. But that relies on quick lawmaking and judges evaluating something they're not specialists at. And there's no dual for punishment (actually compensation and reeducation), there's no intrinsic reward system. There are prizes and grants, and I'm sure they generate immense value for what they are, but they're still very unsystematic and unreliable, I think.
I dream of a system where if you generate value for society, you will be recognized and make a comfortable living off it; if you need investment, you will also be supported. (I call this idea Elementalism; I'm not sure it's original)
I think we need to be a little more open to careful, thoughtful changes to our social-economic system that can improve things.
While Elementalism isn't a formal part of our system, each of us can do our part and give directly to what we understand needs the most. This is the core idea of Effective Altruism which I also extend to donating to Open Source (which I believe can be effective[2])
[1] https://news.ycombinator.com/item?id=29043752
[2] https://www.reddit.com/r/EffectiveAltruism/comments/v7ma0d/w...
Now, I will admit that for this particular query Kagi and Google results are pretty close. But my general experience is that when I search in Google I find that I have to look farther down the search results to look past the blogspam to find the authoritative reference.
Go search for something like "postgres cte" and you won't find anything useful until probably halfway down the page. And maybe not at all.
Now that you know about blogspam, we'll move on to the next topic: How to find blogspam. It's actually very easy to find blogspam. You are on the right page if you want to learn about finding blogspam..."
Instead of a dark forest (https://en.wikipedia.org/wiki/The_Dark_Forest) think of the outer internet as a fake forest. If you wander off the beaten path - the same dozen sites that everyone uses and complains about - you wander into an endless, trackless zone of fakes all of which are ultimately trying to sell you something.
https://www.postgresqltutorial.com/postgresql-tutorial/postg...
The first result was the official documentation.
What else are you expecting?
What I'm wondering is how much of your recognized fingerprint influences the results? What causes results to be different from user to user using the same search query?
[1] https://www.postgresql.org/docs/current/queries-with.html
Those training wheel search results are annoying but they're highly ranked probably because most people like and use them.
Neither of them offers bad results, but we are talking about google, it makes no sense to give example searches without saying what results you receive as Google is so heavily personalized.
Kagi does personalization, but it’s explicit. You yourself decide the region to search in (my default is international, though there are quick bangs to search in other regions), and you can up- or downrank domains, as well as block them completely.
https://kagi.com/search?q=postgres+cte&r=us&sh=5-n8GUySt5qmx...
For some reason, a lot of these search engines like to brag about the number of documents in their index, which never made sense to me. Maybe it was true in the past, but on the modern web, larger index !== better results. In fact, I'd argue the opposite since you're much more likely to serve SEO spam.
That was my motivation to start hacking around on CrowdView (https://crew-rho.vercel.app), a search engine specifically for forums and discussion content (e.g. forums, discords, twitter, reddit, etc). It has a curated index (today, curated by yours truly) to remove SEO junk and help you figure out "what does a real, genuine human think about this think?"
1. Enter the discussion on a human level
2. Hook onto existing context
3. Argue issues with existing context
4. Present your own software as a solution to these issues
It is like a date that only wants to offer MLM business opportunities, or a long-form joke that ends with a disappointing "better nate than lever" punchline.
You could equally say everyone's following Aristotle's Rhetoric.
I do think your idea is great and something with a lot of value for many use cases.
The only issue is that it won't surface certain things like restaurant contact info and Wikipedia that people often need to be a top result.
This is a good idea. Basically harkening back to the old internet directory days, combined with the powerful indexing tools we have now, and then use all the great language models and topic modeling tooling to make the querying great.
Said another way: moderate and protect the index, instead of trying to clean up the results of queries.
You'd need to monitor the whitelisted sites in various ways to make sure they don't bait and switch and turn to spam later, or get hacked and go to shit.
You might even be able to monetize on the publisher side by allowing sites to pay to be indexed faster/more regularly, which could potentially (but not totally) help with spam control too. Every publisher wants traffic.
You might be able to jump start the whitelist by getting folks using a plugin so you can know what domains they spend time on in the first place and index those when signal strength warrants it. Gotta be careful about that being gamed too though. Also I think this is actually the core principal of Brave's search engine come to think of it.
Who would be the curators? If you open up curation to volunteers, it could be gamed by bad actors. However if you have too small a team, the results will be limited and biased. For example, favoring the English-speaking or tech-sphere web while ignoring large sections of the web with which the curators are unfamiliar.
Perhaps machine learning could help with scale - start with a human curated dataset, then train a model on it. However that could end up getting gamed too.
The best way to compete with them IMO is to choose a slice of queries that they perform very poorly at (like non-SEO spam results) and build something 10x better. But that brings up the challenge of getting users to remember to come back to your search engine.
It's tricky for sure, with no clear answer. But for the first time in a long time, I sense that we're approaching a tipping point for Google's dominance over the web.
You might be interested in watching the series "Halt and Catch Fire".
They can conceivably use this information to do the same thing, curare the results and change their ordering
PageRank was never designed for adversarial scenarios.
It reminds me of KPIs like lines of code - it's only useful if it cannot be manipulated.
"When a measure becomes a target, it ceases to be a good measure."
That all seems pretty quaint these days, but even worms and viruses at the time were versions of "my goodness, who would ever write a script that emails itself to your address book, or copies itself across the network to all the other unsecured PCs with passwordless full hard drive access?"
The real consequence PageRank had was that it got web pages to stop linking to each other…. Google gaslighted people into removing the competition they could have had navigating from one page to another through links. (e.g. web directories)
I think the problem may more have been the lack of a sustainable revenue model. Getting volunteers to curate the directory is particularly destructive because the people who most want to volunteer either (1) want to promote something or (2) want to get paid so somebody else can promote something.
When there’s so much money at stake, people will go to great lengths to reverse-engineer the things that put them higher in the page, no matter how you try to hide the levers that make the rankings work.
Instead, you want to throw queries and user behavior into a blackbox algo and have it tell you a result then give it feedback on whether the result was good (did the user come back and ask the same question? did they click the top result and leave or did they have to come back and click through many more?). I think this is kind of how google works now, though results are frequently meh. Millions of backlinks will get you noticed but your ranking will just keep dropping if users don't appear to find your content useful (e.g. they hit the back button a lot and keep searching).
One form of personalization that they applied early on is move results you click on a lot up in your results you can’t trust the serps you see at all.
They also inject a lot of arbitrariness and randomness to make it impossible to make small changes to your site to incrementally improve it the way you would improve an advertising campaign. One reason the web seems so frozen is that if you have a successful site and make major changes to the layout, titles, link structure, etc. you will possibly trigger the ‘chaos monkey’ and wreck your ranking and it may never recover.
From a more macro perspective, I’ll believe Google is failing when a competitor starts eating their lunch. What I see right now are a bunch of would-be competitors who want to eat their lunch, including this company. The blog post is probably best understood as aspirational rather than descriptive.
As a user of search, Google results are frustrating at times, but is that because “pagerank is over?” Or because it’s an incredibly hard problem they’re working on? Google does not have to be objectively perfect to keep succeeding, they just need to be better than other search engines.
I whole heartedly recommend kagi. My favourite feature is the "blast from the past" that shows results that are not online anymore, but links to web.archive.org.
How does the workflow look like for sharing an album of pics with friends and family (who are on traditional social media platforms)?
Cloudflare access will handle auth and can use family members’ gmail or popular emails for authentication. It will also do https for you.
If you want the data encrypted from cloudflare, I’d probably install tailscale on family members’s machine and do ACL. Friends are gonna be challenging
Even in a hypothetical situation where ads would no longer be the driving force of algorithms, something else will. As soon as something gets large enough, it will be gamed. If not for commercial reasons, it might be cultural/political influence.
I will end by reminding ourselves that us techies should spend some more time with ordinary folks. I agree with everybody here that Google Search has been getting worse and worse for years now, especially for our niche searches.
It's a mistake though to think that this is widely experienced as such. My mum looks up an unknown ingredient in a recipe and within a second sees a picture of it. That means it works. All kinds of personal data might be shared in the process but since you can't really see that, it didn't happen. From her point of view, Google Search works extremely well and is close to magic.
The point being, Google doesn't give a shit that you don't get the best answer for your query on JavaScript closures. Nobody searches for that, and those that do, block ads.
Sure maybe it would be good if Google prioritised these results by default but that would just result in reddit becoming more manipulated since whatever the default search shows will be gamed.
One that annoyed me today: on mobile, reddit search results will all say like "3 days ago" or "6 days ago" and then the actual link takes you to a six year old post.
People don't want to search for a bread recipe, click to see it and be forced to scroll through 3 folds of ads and the history of bread in eastern europe, to get to the list of ingredients.
When didn't people want "direct quick answers"? Google got popular because that's what it used to give you - the answer you want used to always be the first or second...
While this sounds true, I think it actually needs motivation.
As far as I can tell, this type of analysis typically bases itself on only a few examples.
Actually Google is already doing this. Results are personalized and contextual. It can't really know what I feel like eating this evening because I don't, but it can guess.
I applaud the author for trying, but I don't see an alternative to PageRank being proposed. How exactly is Kagi proposing to rank results?
And no, I don't want an AI generated summary when I search for the best tutorial to do X. I want a list of tutorials. The question is how to rank that list, and I've yet to see anyone do a better job than Google.
Google is far more than PageRank. Its an AI ranking model that has the largest training dataset (queries and clicks).
Ads may suck, but simply charging for the same service isn't really innovative. Would I pay for a version of Google without ads? Probably not. But that's just me - I actually like to know who is advertising for particular searches. A company with an ad budget to rank at the top of ads is probably more trustworthy than an anonymous website.
Google already does this with YT; YouTube Premium[1] offers you YT experience without ads plus extra perks. And YouTube Premium has apparently more than 25 million subscribers in US alone[2] and 80 million subscribers globally[3]. My thinking is power users are ready to pay for ad-free experience and casual users probably not because they are not heavy users.
>But that's just me - I actually like to know who is advertising for particular searches. A company with an ad budget to rank at the top of ads is probably more trustworthy than an anonymous website.
A lot of spammers and fraudsters see advertising as their most effective gateway and tactic to scamming people so I wouldn't count on reliability and safety of all ads that you see.
[1] https://www.youtube.com/premium
[2] https://www.statista.com/statistics/1261865/youtube-premium-...
[3] https://blog.youtube/news-and-events/youtube-music-premium-8...
Google, less so - but I pay for Kagi as I like the results better, and prefer how they say they're tracking me.
YT Premium is a reaction to the infeasibility of subsidizing the extreme cost of hosting video streaming with ads alone. The scale of ads needed to cover hours of streaming video vs. ads to cover search results are magnitudes in difference, and people are more willing to pay to make those streaming ads disappear.
Even so, with a paid option, the vast majority of YT users still use Youtube w/o the Premium subscription.
That doesn't necessarily make it better for HN users. A more selective dataset could make it more useful to specific niches, like say, HN users?
It’s mostly to find an image, a video, a location on maps, etc.
I wonder if Google could turn things around by rethinking the search experience. They’re still the fastest first stop on your journey to find something on the web. But the behavior has changed dramatically that a change in their UX alone could really make a big impact. Maybe moreso than just an algorithm change.
I once redesigned the search experience of a large global ecommerce company. The UX changes alone grew their revenue quite a substantially and reduced customer complaints.
By the way, I've been a satisfied Kagi customer since the day the beta ended, and I have nothing bad to say about the service. Okay, it would be nice if they remembered I like to view temperature in Fahrenheit, but whatever.
but it takes a huge amount of training/test data to do it and takes a lot of computational work after you've got that data. What TREC revealed is that most of the things that would obviously improve search relevance don't, and it took 5 whole years of 20 teams working on it before a useful discovery was made.
There was an article about TikTok that revealed just how wrongheaded the viewpoint of the current web is. As much as Google fetishizes data, the data collected by sites like YouTube is useless because they offer you too many things to click on so you click around like a chicken with its head cut off. There are maybe 5 things that you like out of 20 that they show you, but which one you click on is random so the signal is mostly noise and worst of all they can't come to the conclusion that you didn't like any of the other 19 things they showed you.
TikTok shows you exactly one thing at a time so the opinion that they capture is meaningful.
It reminds me of one of the first "learning to rank" papers where I talked the management at the CU Library to let the Thorsten Joachims group run our search engine and we realized just how poor of a signal you get from search engine usage and how challenging it is to feed it back to improve your results.
Really marking up judgements for all the results that turn up and using "pooling" to add new results to the original set when your search engine is essential to make really better search. Otherwise you can be just another one of those guys who posts to Medium about the "semantic search" engine he built but can't tell you if the results are any better than any other search engine.
Plenty of excellent youtube channels that 10 years ago would've been websites.
I don't really have a problem with this, all things change.
Compare written instructions for e.g. fixing a car vs a video of a person showing you how to do it. Latter is much more helpful
This is true IFF this software is open source and running on my own server. Hard to see how this can happen.
Since Google turned to the dark side on August 9, 2006[1], ranking by links weight has been hopeless. But the alternatives are not much better.
I suspect that was what you were seeing, and if so, updating should have addressed it.
This reminds me of posts from 10 years ago where "the future of search of social" (this post says "user-centric" instead). Even Google feared that outcome (given the then-meteoric rise of Facebook; oh how the turns have tabled). Personally I find the musings in this post to border on nonsencial.
But here's the interesting part: certain ideas are sticky just people want them to be true. The core one here is that "search is broken" or even the lesser "Google is getting worse". It's one many on HN like to push. I don't think theres any quantitative empirical evidence of this. It's people wanting it to be true and you'll see it on every thread about DDG (as one notable example).
Silicon Valley is littered with the corpses of "Google killers" who seem to have all fallen into different versions of this trap.
Google's algorithm may be the best out there, but it still falls short of what is needed. Anecdotally, I often find it fails to conduct word-sense disambiguation, handle negation or answer edge case. My general impression is that Google's algorithm focuses on topical similarity (pages that talk about the query's topic) rather than answering the query.
This is exactly what Google is trying to do....unsuccessfully. I doubt you will have the scale to outperform Google if you try to replicate Google's AI model and replace ad based search with subscription based search. My point is don't do what Google is doing, think differently.
This is what Google is doing under the hood and they have the upper hand because they collect absolutely everything about the user. Yea there is a lot of noise but their advantage is enormous because theoretically the more they know about you the better search results you will get.
I respect the policy of letting users decide what data they want to hand in(volunteering their data) but I don't see how this can scale in the long term because users most probably don't even know what they know and they don't know what exactly they want.
My argument is that this is exactly the main reason for the mess on the web we are in now, and that the way forward is not doing more of the same thing that got us in this place, but doing the opposite, regardless of how hard and not scalable it is. Alas, we will at least try.
Better from whose point of view? ;)
> “Currently, the predominant business model for commercial search engines is advertising. The goals of the advertising business model do not always correspond to providing quality search to users. For example, in our prototype search engine one of the top results for cellular phone is “The Effect of Cellular Phone Use Upon Driver Attention”, a study which explains in great detail the distractions and risk associated with conversing on a cell phone while driving. This search result came up first because of its high importance as judged by the PageRank algorithm, an approximation of citation importance on the web [Page, 98].
A search for cellular phone now returns a page of links to plans, and one link to Wikipedia.
"And yes, the non-zero price point will mean you have to budget it with your other costs. But faster access to higher quality information will make you much more competitive globally, so you can decide if the investment will be worth it, like any other purchase you make. This will in turn incentivize these products to be even better, a positive feedback loop driven by entirely aligned incentives."
For businesses I think this effect is only driven by competition regardless of revenue model.
The more the economic incentives of a business are aligned with the customer's interest, the better. Of course, it's not perfect, and business still might have an incentive to show ads on top of charging you, but it's certainly much better than having hugely misaligned incentives as with Google and social media today.
> Why does everything in the future-web have a price tag attached?
Because it uses real resources like humans and hardware that cost real money in the real world. You don't ask why you need to pay money to buy an apple at the grocery store, do you?
Jerusalem is a city like any other. Everybody has food. You can buy bread there. And while the city has big problems, not many people are starving (not quite zero, but which city can say that). Economic theory dictates that bread would follow the law of supply and demand. Is that so?
Let's say the price of bread halves. Will tomorrow see more bread sales in Jerusalem? Here's the critical piece of evidence: NO, it won't. It won't even shift from one kind to a fancier kind of bread. Let's say the price doubles. Same thing. We're in a ZERO price elasticity regime. Sure there are prices where supply and demand would work, but we're far away from those prices. But it gets worse.
Let's look at supply. Bread is incredibly cheap to make. There is at least a 80% or so margin on bread (it's quite expensive, or profitable on the other side, if you look at it from a COGS perspective). Let's say the price of bread doubles (you could argue it has, many times even). Will more bread be produced in Jerusalem? NO. Again there are prices where we'd see supply tightly tied to prices, but we're very far away from those prices.
BOTH on the supply and demand side we see ZERO price elasticity. If you look up in your economics book you will not find this situation, and even in academia people aren't exactly comfortable discussing it. Price elasticity is ZERO for bread, both on the demand and the supply side, if you belong to the economists that at least split that up.
So what makes the difference? You can make a good living selling bread, but you can't attract customers using prices ... nor can you change your profits much by playing with prices (in fact the very act of trying out prices will lose you a lot of customers). So something else needs to make a difference. What determines which bread you buy? Google does (and Facebook, and ...)
And this applies not just to bread. This applies to the vast majority of foods. One of the big complaints in European countries these days is that entire economies are running headfirst into this problem. Even labour! People can't change jobs (supply of labor is regulated by employer demands, or outright regulated with laws) and employers can't change the pool of labor by much by raising wages. Unless you're exceptional or outside of the norm (yes, programmers are, so this is the wrong forum to make that particular argument), you cannot change your income much by providing more labor! Given "no skills" labor even the difference between working and not working ... is not very large and dropping.
This, obviously, enables FANG companies to extract the profits of normal businesses. They can choose who gets the profits of many products, and demand a growing share of the profits ...
https://en.wikipedia.org/wiki/Wikipedia:Funding_Wikipedia_th...
I don't believe direct sales fixes everything. There's no single party whose ultimate wishes are fully aligned with what's best for everyone. I do think, however, that direct sales and donations are preferable to ads in this regard.
There's reasons that the incentive to keep the users satisfied may not be followed, but it's mostly tangential, and relates more to other forces than it does where the money comes from.
Selling more product / continued subscriptions is the incentive.
There are so many ways to use Google to obtain the information you're looking for, and so many people do it literally on a minute-to-minute basis that it's flat absurd to call that ineffective.
It's trivially easy to learn something by typing a question into Google; maybe when writing an article try gut checking it against obvious observed reality first.
I simply cannot fathom a noble reason why the author would decide to publish this.
Edit: Oh wait I found it; the author is Vladimir Prelovac, CEO, Kagi Inc. What does Kagi Inc. do? I'll leave that as an exercise to the reader, but I bet you guess it in one try...
Point taken. But it's not as though this was concealed. The top line of the page is "Tales from Kagi." The piece is hosted on blog.kagi.com. The author's signature at the end of the piece reads "CEO, Kagi Inc." The final line of the post is "I hope you join us [i.e., Kagi] on this journey."
Maybe it doesn't bother you, just thought it was worth calling out explicitly.
It is marketing and trying to influence the discourse, but I don't see how you can call it SEO.
Also it's written by a CEO of a competing search company. It's literally the very SEO he's railing against...
Google is a horrible search engine, you routinely have to hack it to get searches not linking to the centralised web. The first 5 results are ads, followed by Youtube links or a Quora/Reddit on the topic.
If you used the web 20 years ago you know what a good search engine was like.
Especially if you're looking for Google users' private information, and you're a business partner of Google's.
> It costs us about $1 to process 80 searches.
This caught me by surprise. It would be good to know if this is pure hardware and bandwidth cost, or if they are factoring in all the costs associated with running a business. I wonder if costs for Google are also on the same level.
There team's plan seems to quite expensive considering we are so used to free (ad supported) search.
He then described how people will first do their usual _web searches_ via FB search bar. To my knowledge, that didn't happen. And he has since pivoted to "the future is private".
[0] https://techcrunch.com/2010/04/21/zuckerbergs-buildin-web-de...
PageRank could have been fine if the intrinsic motivations were to educate, enlighten, and help people.
Just because a search technology is wrapped up to look like a personalized AI doesn't mean it is immune to the same intrinsic motivations.
Who is to say that Google won't release a 'Personalized Search AI' that is free for the user and it continues to follow the same motivations that their PageRank solution does.
I think custom AI's tuned for each person will be the defining technical foundation for the next few decades, but I'm extremely concerned that the fundamental business drivers that compete for sales, view time, and divisiveness will not be resolved.
Nothing will stop an authoritarian government, an ad driven company, an entertainment company, or a political research company from giving out 'free' personal AI's.
If the history of the Internet has anything to teach us then it is this:
The free ones will always be more successful than the paid ones.
It's often pretty fustrating to try and find something you know you've seen before but don't remember quite the name of or exact keywords.
Ideal would probably be face/eye tracking to check attention/response but thats way too creepy to work.
It sounds all great and I agree that we should be prepared to pay for quality services instead of expecting everything to be free (so I will have a look at kagi). But this promise makes me skeptical: Twitter is apparently going to lose money on their "$8/m for half the ads" deal with US users because the ads make more money than that for them. Sure, paying for your services is going to mitigate the incentive problem, but it might not fully eliminate it. Of course, kagi has a reputation to lose with its users, so that is another line of defence, but I guess the best way to build such trust would be to be maximally transparent, even about the incentive structures.
More and more I find myself searching through moderated social media like Reddit or Facebook Groups/Marketplace for product recommendations or local services.
Gave Kagi a go, but it seemed to be even worse than Google for what I tried.
And moderated groups don't sound like a solution to the 'I have a concrete question that I want an answer for right now' problem either (unless you're thinking of massive live chats with many lurkers, which sounds more like the web of the past, tbh).
so many of the complaints about google search results seem to be along the lines of "i asked google what X i should purchase, and they served me an ad", whether that's a first-party ad on google or search results that are external ads. If you're only attempting to use a search engine to find facts, it's a much more satisfactory experience.
search results for opinions are and have always been trash - they don't serve you the best opinion, or the most trustworthy, or the most well-reasoned. they serve you the opinion most relevant to your search terms. and that's probably not what you want.
It took years for peer-reviewed papers to show any benefit from PageRank at all, part of it is that a real relevance function has to balance the keyword x document influence vs the document influence and you don’t come out ahead ranking an irrelevant document highly if it has a high page rank. (E.g. a popular document that is irrelevant is… irrelevant)
If you believe the original paper, PageRank is simulating the density of a random walk over web pages and Google has been able to sample that density directly w/ Google Analytics, Chrome browser telemetry and all their other web bugs.
That being said, I'm beginning to think that having a price tag as entry barrier is a good thing for a search engine.
One of the reasons for Google's downfall is that they had to cater to dumb and easily controllable masses.
Also, the aggressive SEO that made the search results so bad made economical sense as soon as those masses flocked at the platform.
Kagi may only stay this good if it does not become too popular. I hope the developers think about the fact that their best strategy might be focusing on their current sizeable niche market of nerds and intellectually curious people.
Two improvements that I'd like to see in search results are: (1) the ability to exclude particular domains from search results; and (2) the ability to over-weight certain referrers (i.e. a link from Encyclopedia Britannica is worth more than a link from Wikipedia, which is worth more than a link from MySpace.)
You are correct but in the scientific community there are no armies of spammers who persistently try to outperform other scientists by writing fake research papers and spamming citations. There are cases of fraud here and there but not at the scale of the Internet or to be exact the World Wide Web(WWW).
>the ability to over-weight certain referrers (i.e. a link from Encyclopedia Britannica is worth more than a link from Wikipedia, which is worth more than a link from MySpace)
Google probably already does this but exactly how nobody knows because of their blackbox algorithm policy. Speaking about this suggestion; it depends on what you search query is....you need to assign specific weight values to certain referrers in relation to search queries from X,Y,Z categories.
https://dkb.io/post/google-search-is-dying
Big discussion here and on Reddit when that article was posted. It doesn't set out to say Google search is worse/getting worse, but that people don't trust its results as much as human generated content (thus appending 'Reddit' to queries).
You can mitigate it with curation, but then you’re bound to niches and bias, which is true with automation as well. No victory here.
But in general, especially when looking for specific phrases or trying to find discussions related to a topic, it's working extremely well. Whenever I used it, I was happy with it, and the search result quality was better than in Google.
It's good to note though that iirc Kagi does actually use commercial search services by Google/Microsoft behind the scenes, in addition to their own custom components.
Wait, what? It boggles my mind that technologically inclined people continue to use this user-hostile, walled garden nonsense.
This is a non-sequitur. Publishers could monetize their traffic without PageRank, it is not a necessary condition for the publishers to realize they could monetize the traffic. This thesis statement is the crux of the rest of their argument, which makes the rest of this not make any sense.
E.g. instead of training it to generate sentences given a prompt, train it to generate URLs given a query (?)
IMHO, Yandex feels like the old Google. It returns lots of small sites for most searches and does way better on controversial topics.
unless advancement is understood as necessarily attendant upon linear time unspooling, it is no ways inevitable
Instead they are going the other way. Sad.
Given Some social media platforms are moving towards some embrace of AI feeds, could those platforms retain their 2nd search engines role such as Youtube, Instagram, etc?
My own biased use case is that I started following a whole bunch of designers on Instagram to get new design news.
> Yet, despite being acutely aware of the dangers of ad-supported search, selling ads was adopted as the primary business model of the new search venture just a few years later.
Google was successful as a superior search engine, pushing out AltaVista and similar, earlier alternative search engines and Yahoo!-style portals curated by humans. However, Google wasn't a successful business for quite some time. Eventually, they adopted (some people would say "stole" - there was a lawsuit) Yahoo!-owned Overture's ad model, which changed everything. Yahoo! owned Overture, and Overture had a critical patent. Yahoo! made one critical mistake: they settled the lawsuit for relatively little money. The rest is history.
Now many people complain about decaying search result quality levels. That just means there is space for a new search engine, how exciting! The good news is it has never been easier and cheaper to start a full-text index of the Web and associated search. For about 50k (a Xoogler's estimate, not mine) you should be able to get going. Sites like Gigablast show it can even be done as a one-man show, which I would not recommend (to many complexities in "small" bits even HTML to plan text conversion, load balancing, incremental inverted index updating etc. - all requiring nowadays some specialist expertise in a game where you can't afford to reinvent the wheel because you don't know the scientific literature/state of the art). The one thing that is hard to get is initial user traffic. But I think HNers will be happy to give each new engine a try!
In summary, I think there never was an "age of PageRank". But you may say Google Web search is past its prime. Perhaps Google could change that if they wanted - it may be that it isn't much of a priority at the moment, hard to say (they are (too?) big now).
Edit: Here, I've interviewed Shadi Saleh, the architect of Syria's search engine (if you think it's impossible to get up and running with a small team): https://irsg.bcs.org/informer/2019/07/syrias-first-web-searc...
https://kagi.com/faq#Where-are-your-results-coming-from
"Our searching includes anonymized requests to traditional search indexes like Google and Bing as well"
Lol okay
Check the history section: https://en.wikipedia.org/wiki/PageRank
If we just trace back bibliographic references to decide on the inventor of PageRank, it will probably be Archimedes.
So what destroyed the relative purity of search engines and the early web? Money. How do we solve it? Money apparently.
Yeah good luck with that.
What an optimistic perspective.
I’ll stick with DuckDuckGo, thank you.
I am genuinely curious how did Kagi arrive at this cost.
If I were them I would be focusing on: How do we get to 8k searches for a dollar?
Why would you not take the chance to completely reimage search into something way better than the absolute random crap we have today?
Whatever the next web discovery engine looks like, it'll look and feel entirely different.
It won't be the equivalent of someone saying, "Hey Facebook is bad, look at MY social network!" and it being the exact same layout and user experience as Facebook, with different margins, padding, fonts and colors.
Whatever comes next, it'll be like comparing Twitter to Facebook, or TikTok to Instagram.
As the kids used to say, this ain't it, chief.
How about personalized pagerank search that assigns any URL I've bookmarked a high value for E?
Is there a reason no one has done this yet?
Then you can check google or a mainstream engine when needed.
Besides, Kagi Safari iOS extension is full open source: https://github.com/kagisearch/Kagi-Search-for-Safari-iOS
EDIT: I'm dumb. Please do not reply or I will Sylvia Plath myself.
You can’t be serious with a claim like this
Time to join a web ring.
I think the biggest issue is live news but I mean then just pay like 12 people to watch news/sports networks and you won't be that far behind. (Or just have an section of the index for live information and let users to go nfl/espn/oan/cnn).
Otherwise I might have given it a shot.