Designing a website without 404s
pillser.com
pillser.com
The "implications for SEO" will almost certainly not be "positive" for quite a few reasons. As a general rule in my experience and learning, Google's black box algorithm doesn't like anything like this, and I expect you will be penalized for it.
There are many good comments already and my suggestions would merely be repeating them, so, just adding my voice that this is likely to be a bad idea. Far, far, better to simply have a useful 404 page.
Edit just to add: if you do something like this, make sure you have a valid sitemap and are using canonical tags.
I don't understand why people don't just read Google's own published guidelines ( https://developers.google.com/search/docs/fundamentals/seo-s...) on how to properly SEO one's site.
It's not some dark art. Google themselves tell you exactly how they index and why. All you have to do is read the guidelines.
Then there are the vagueries of a search engine becoming ever more dependent on ai where not everything makes rational sense and sites can get de-indexed at the drop of a hat.
I did this during a migration of a site that used '/<pk-int>/' and was changed to '/<slug>/' with corresponding 301. The SEO not only didn't punish the migration, but it seemed to like the change (except the Bingbot, that after five years still request the old Urls).
The problem I see with the OP strategy is that the bot can hit a Url that still doesn't exist, get 301'ed with that technique, and the Url is reused afterwards for new content: following their example, somebody links to '/supplements/otherbrand-spore-probiotic', that gets a bot following that link 301'ed to '/supplements/spore-probiotic'. Later you actually add 'otherbrand-spore-probiotic', but that Url will never be visited by the bot.
[1] https://developers.google.com/search/docs/fundamentals/seo-s...
301s are what you are supposed to do with things like changed URLs, so I cannot see it being a problem in this case. A 301 is nor duplicate content - it is one of the things Google likes you do do to avoid duplicate content.
It would probably be good to add a threshold to the similarity to prevent urls redirecting to something very different.
Semantically, using a 303 redirect might be the most appropriate signal for what they are doing.
They could redirect to a specific result page if it's a clear, unambiguous match and they could redirect to a search results page if there are other possible matches.
This reminds me of something similar in the mobile world: Apple's App Review guidelines[1]. They're surprisingly clear and to the point. All you have to do is read and follow them. Yet, when I worked for an App developer, the product managers would act as though getting through App Store review was some kind of dark wizardry that nobody understood. Them: "Hey, we need to add Feature X [which is clearly against Apple's guidelines]." Me: "We're going to have trouble getting through app review if we add that." Them: "Nobody knows how app review works. It's a black box. There's no rhyme or reason. It's a maze. Just add the feature and we'll roll the dice as usual!" Me: "OK. I've added the feature and submitted it." App gets rejected. Them: "Shocked Pikachu!"
I'm with your colleague on this, at least in a historical context.
I was also well aware of and accepting of the risk in my case. I wanted to learn more about certain iOS features, and I was able to do so scratching an itch I felt like scratching. I would have still gained from the process even if it had been rejected. However, as it got approved, contrary to the guidelines, I also got a good laugh at how inconsistent review was and a nice cash bonus on top!
You win some you lose some. Such is life.
Going through any forum thread of "reasons you've been rejected" and trying to find the reasons here isn't obvious.
I couldn't find anything in there pertaining to why we were rejected last month (mentioned Android in our patch notes as it's a third party device we integrate with).
A 404 doesn't cause nearly as much dissonance, because the website is at least telling me why what I'm seeing is not what I expected.
Not only did it improve our page rank once we fixed it, it also reduced the amount of bot traffic we were fielding.
That said, typically most people are going to find a site via search or clicking a link, which is why I think basically no one bothers doing this.
1) Users may (inevitably will) share links to the URLs on social media / wherever, and Googlebot et al will find those links and spider them.
2) The URLs will be sent to Google for any Chrome user with "Make searches & browsing better" setting enabled.
https://support.google.com/chrome/answer/13844634?#make_sear...
So no, please don’t ever do this.
But what you can do is provide a useful 404 page. Say “there’s no page at this URL, but maybe you meant _____?” Although there’s a strong custom of 404 pages being useless static fluff, you are actually allowed to make useful 404 pages. Just leave it as a 404, not a 3xx.
(Also, care about your URLs and make sure that any URL that ever worked continues to work, if the content or an analogue still exists at all. Distressingly few people even attempt this seriously, when making major changes to a site.)
I might be wrong but for me it’s because IRL this isn’t an issue. Users shouldn’t be finding/using random URLs to navigate the site. Where is this broken URL traffic coming from anyway? Are you trying to solve for people that randomly edit the URL and expect it work, most people don’t care about those users getting 404 because they should expect 404. They’re not real users they’re just playing around.
However, If you purposely changed the URL format after a lot of people have the old format bookmarked or indexed on the web, then do a 301 redirect to the new URL.
I’m not sure of the SEO implications of the described solution, however it seems like only risk and no upside.
301s also serve to inform search engines that your URL should be different and they update their index accordingly. My transition went very smoothly but it took prep work on my part to understand that 404s were inevitable and that I needed a plan to migrate that traffic.
Whether it is positive or negative, I do appreciate it as it helps me to learn and improve the product. I really didn't expect this to get any attention, let alone dozens of comments!
To clarify: This was originally designed to help me auto migrate URL schema. I am learning as I develop this website, and SEO has been one of those vague topics where there are few hard rules. I wanted to leave space for experimentation. As I rolled it out, I became intrigued with how it functions and wanted to share my experiment with you to get feedback.
Based on the feedback, I plan to change the logic such that:
- I will track which URLs are associated with which products - If user hits 404, I will check if there was previously a product associated with that URL and redirect accordingly - If it is a new 404, I will display a 404 page which lists products with similar names
I appreciate everyone hopping in to share their perspective!
Best compromise is that. It is still a 404 but with a "Did you mean" element to it, which is still useful to end users
Maybe because it is not a good idea. Masking errors is harmful. The information "what you're looking for is not there" is very important, because it lets users identify that something is wrong. Smart redirection can be outright dangerous: what if I am buying a medication, and the smart website silently replaces the correct drug with something similar but wrong? Not to mention the pollution of the indexes of search engines with all the permutations of the same thing they may discover. Lastly, the U in URL stands for unique; the web is designed around unique locators and this rule shouldn't be broken without a very good reason.
Show the user (and the crawlers) a 404, and suggest your corrected URL in the content of.the 404, and let the user know that it's a guess, so they make an informed choice about the situation.
Answers that appear to be right can be worse than no answer at all.
- Keep the 404 page
- Use autocorrection only for minor typos
It doesn't in a literal sense and not even in a practical sense.
The web in no way is designed around unique locators (not even unique identifiers). And for the search engines you point to there is metadata like `link rel=canonical` to help them around the obvious reality that one webpage may have many different locators (and you will find that used on most major websites)
- https://pillser.com/supplements/trump-won-the-election
- https://pillser.com/supplements/hitler-did-nothing-wrong
- https://pillser.com/supplements/9-11-was-an-inside-job
etc etc etc
Someone could seriously mess up your site by simply publishing their own page with many invalid links to your site, basically a dictionary attack, and if Google was to crawl those links, they'll cache all the redirects, and you'll have a hard time rectifying that if you were wanting to then publish pages on those URLs.
Also to reiterate other suggestions - your idea is not great for many reasons already stated (even with 302s). As suggested, just simply have a 404 page with a "Did you mean [x]?" instead. Use your same logic to present [x] in that example, rather than redirect to it.
As described elsewhere, here is how the new logic will work:
- I will track which URLs are associated with which products
- If user hits 404, I will check if there was previously a product associated with that URL and redirect accordingly
- If it is a new 404, I will display a 404 page which lists products with similar names
That said iirc it came up because they didn't have a proper 301/302 configured
I don't see the problem that this is a solution for but I can see a couple of problems that this solution causes.
https://pillser.com/supplements/merotonin
I am surprised that more websites do not implement this!
Vendors change product names, hyperlinks break! Fix bugs or change behavior, hyperlinks break! Do nothing, believe it or not, hyperlinks break!
I do think that it's ok for several URLs to point to the same content. In his example all three are fine. I also tried the product code (6066) without any of the text and it worked fine as well.
I've also noticed Lego's site does a version of this as well. https://www.lego.com/en-us/product/10334 will take you to the product page for the Retro Radio. However, I think Lego's site is just keying in on the product id as "retro-radio" doesn't work, but "ret-rad-10334" does.
But there are limits.
I put in the URL "https://pillser.com/supplements/go-fuck-yourself" to see what would happen. Now, I chose an offensive phrase to increase my chances of not coming close to any real product. I believe that URL should 404, but it took me to the page for the supplement "On Your Game" instead. If I had tried a real name and got taken to something with only the barest resemblance to the name I tried, I wouldn't be thinking "This must be the closest match". I'd think the site did something messed up or I typed something wrong or something malicious had happened.
> a product is renamed, or the logic used to generate the URL changes.
In that case you should store both urls and have a redirect_to_id or something similar to give search engines and users a proper 301. I don't see a use case for this fuzzy matching which will just make things not very explicit and unpredictable.
It is good because it will help most people and work for them.
It's bad because sometimes it will make people think they've found what they were looking for when what they had doesn't exist at all - but it gave them something that sounds similar.
I would at least have a "redirected from" banner at the top of the page when it triggers.
That would be true if all visitors would understand 404 pages exists, and they expect one for this website. I bet most people wouldn't know what to answer to the question "What you should see if there is no page for this supplement?"
It is true that there coul dbe some sort of message making the redirect explicit.
For everything else you can just have a nice 404 with suggestions of links that probably are a match.
The web by definition is a lazily-materialized query response graph.
Still, it's a good idea.
You can further this idea (especially when the slug returns nothing) by having this page also list "Best Bets" or what people most often come to your site for (regardless of any search query, perhaps, with their referrer, or on this day of the week etc)
And additionally, put the slug (bar the dashes) into a search box so it might be ammended (but tell them that you didn't find anything and they need to try something else).
Not sure how to best do that in postgres though, closest I can find is reserved connections per user. Idk maybe there's an extension or it's easy to do it in the webserver
https://www.postgresql.org/docs/current/runtime-config-conne...
Slightly OT but came into my mind when thinking about designing website related stuff: https://www.w3.org/Provider/Style/URI
Here is a partially fixed issue https://hatonthecat.github.io/Hurl/404.html
this is not going to end well...
"Hey that page doesn't exist, but there are some similar pages..."
https://pillser.com/engineering/2023-06-10-website-without-4...
You show me the links you think I want, on the 404 page.
clickable:
https://pillser.com/supplements/vitamin-1973-omg-i-can-type-...
Popular libraries like https://github.com/norman/friendly_id implement it like that too.