Somehow in Google search one of the unguessable pages is indexed. We have used Claude and Gemini to assist with some design aspects.
I'm thinking some aggressive data ingestion/indexing is happening by all the bots in the quest for frontier models.
Somehow in Google search one of the unguessable pages is indexed. We have used Claude and Gemini to assist with some design aspects.
I'm thinking some aggressive data ingestion/indexing is happening by all the bots in the quest for frontier models.
Especially if you have autocomplete-while-searching type of features on.
They control the entire browser surface, technically they can know everything, even TLS and E2E encrypted data, that they silently phone home…
If you think this is silly, consider that Microsoft Recall had been observing everything on people’s entire SCREENS and phoning home much of it. That is how a guy was caught recently: https://x.com/t3chfalcon/status/2074134314145489195
And it is actually much worse than even that:
https://community.qbix.com/t/increasing-state-of-surveillanc...
For some reason people are downvoting you, but yea, one day we'll likely see a lawsuit where they do exactly that.
Especially in managed environments where it’s someone’s job to keep systems autonomously updated…
And maybe have access to EVERY site actually, with “forgot password” type stuff in addition to providing oauth tokens…
Boy do I have news for you.
https://arstechnica.com/information-technology/2023/05/micro...
They call the signal „popularity“ and it is a successor of the Google Toolbar signal.
https://www.justice.gov/opa/pr/department-justice-wins-signi...
Why are you using weird quotes?
:root:lang(en) { quotes: '“' '”' '‘' '’' }
:root:lang(fr) { quotes: '«' '»' '«' '»' }
:root:lang(fr-CA) { quotes: '«' '»' '‹' '›' }
q:before { content: open-quote; }
q:after { content: close-quote; }
There's currently 26 quotation mark varieties by default (in Firefox):
https://searchfox.org/firefox-main/source/intl/locale/cldr-q...
which are sourced from the Unicode Common Locale Data Repository:
https://github.com/unicode-org/cldr/tree/main/common/mainEdit: also private browsing isn’t exactly private when you’re logged in to the browser.
[1] https://developers.cloudflare.com/cache/advanced-configurati...
If you don't use wildcard certs all of your subdomains can be scraped from the certificate transparency logs. Additionally, any domain+cert using HSTS with preload enabled end up in a big list at Google to speed up the initial connection from browser to site.
Creating Sitemaps, sharing it somewere public, putting the url in some 3th party service, server logs, some indirect path in javascript.
But if you never mention that url, it will not be found if not leaked by your server.
That sounds like a claim that security through obscurity is infallible, which is dubious. Don't get me wrong, it can be a reasonable part of defense-in-depth strategy, but like, brute force attacks are kinda a well known thing, especially if your URLs aren't truly random...
Google misusing chrome browser history as a hitlist for indexing sounds wild to me, so I tried to see if there's another way.
It also felt unlikely because there's multiple subdomains of mine that aren't indexed, and wildcards+no preload are the only precautions I've made myself for my private sites.
This might also be an EU vs rest of World thing, or my stuff isn't interesting enough to index(in retrospect the most likely reason I suppose)
But I think the other explanations take care of pages: cloudflare hints, chrome reporting addresses visited, etc.
HSTS preload is not for speed. It's to protect against SSL stripping on first connection. Modern browsers already try port 443 first or in parallel with 80.
Finding domains is easy, everybody uses CTL to find them.
Do you use a CMS or other tools that auto generate sitemap.xml? Perhaps you unknowingly told Google about those sub-pages.
Might have been an evil chrome extension, but ever since Google went IOK2BE ("It's OK to be Evil"), maybe it's just Chrome itself.
Also that browser setting to check urls are safe sends them out “sometimes“.
It’s a different story if it’s a subdomain though, OP wasn’t clear.
I have visited that page from a signed-in Chrome profile.
The information you mentioned is relevant. Unfortunately, it could be either Google/Chrome, or the LLM service you're using for development is misusing your data.