A search engine for your personal network of high-quality websites
grep.help
grep.help
I like the idea, but those are some pretty major red flags.
Additionally, the promise of "No AI-Content" appears to be snake oil unless the creators have discovered a magic formula for reliably distinguishing AI output from human output. With up to 7 degrees of separation, your search network is pretty much guaranteed to include AI-generated content, which will be indexed like any other content.
> Additionally, the promise of "No AI-Content" appears to be snake oil unless the creators have discovered a magic formula for reliably distinguishing AI output from human output.
From my testing, at least within your network, I did not notice any AI-generated content. Even if you find any low-quality sites, you can simply block the previous node that refers to the site (yet to be shipped). Grep will remove all the sites that the previous node was linking to, including any AI-generated ones.
Fair enough, but a privacy policy isn't a "nice to have", it's a legal requirement (in many jurisdictions) before you can start collecting email addresses and other PII. It's a very bad idea to push such things to later.
> From my testing, at least within your network, I did not notice any AI-generated content.
That doesn't mean anything. Humans cannot reliably recognize AI-generated content. Why do you think all kinds of institutions are so afraid of it? If it were obvious whether something is made by an AI, there would be no problem.
You can fairly effectively apply network topology heuristics to the link graph to separate the wheat from the chaff. Content farms tend to have different linking patterns from organic content. (Overall, few humans tend to link to them)
But, as other commenters have (imo correctly) pointed out: absolute head-scratcher of a name - especially for a _search _ engine. So much so that it makes me question their judgement more broadly. And, concealed pricing info is a flat-out non-starter for me.
Here is the reason why we liked the name. Grep is a great CLI tool that always gives the right answers if your query is correct.
I wanted Grep.help to do the same for your personal network.
> And, concealed pricing info is a flat-out non-starter for me.
I will mention this on the front page of grep.
Grep the open source project has spent years building up it's name and it's good will by being - as you point out - a great tool.
You are now attempting to claim some of that good will just by copying the name when you have nothing to do with grep.
If you did that to a commercial company they'd sue you for trademark infringement.
(Hopefully this is a serious enough point to not count as a "tangential annoyance")
Your serious point is critical and it must be considered.
The name "grep" brings up the imagery of an efficient search tool for most Linux users. That's based on years of work from the UNIX tool's contributors. A new web service shouldn't just swindle that; it feels wrong in many ways.
As others have already mentioned, I strongly suggest to change the name. Otherwise, it undermines your own message and importantly, does a disservice to the original grep. A lose-lose situation overall.
There is a problem here and that is how to stop "users" from spamming the project with marketing garbage. If the "solution" is to collect more personal data from users, e.g., e-mail addresses, mobile phone numbers, multi-factor authentication, etc., as a prerequisite to even submitting a search query, then that is not a solution. It is an additional problem.
Assuming the project is collecting personal data from users, e.g., peoples' search queries, where is the agreement that governs what the project is allowed to do with that data. If there is no agreement, then presumably the project can do whatever it wants.
We can guess what happens next. Search engine as userdata-gathering exercise and commercial entity. Been there done that.
A future is coming where everything that isn't behind a login wall will be widely abused. ToS mean nothing in the era of AI. Quality searches will stop being free and cost actual money because the alternative is to drown the users in ads and, yes, monetize their data
If you can't learn to establish trust with non-AI entities then..
I honestly think this may be what the future of search looks like. The only reason Google doesn't force you to sign in is they already know who you are.
Operating a truly public and anonymous search engine is a serious uphill battle with the background noise of bot traffic. Looking at what Cloudflare turns down, I'm currently on the receiving end of about 1.5 million fake queries per day from a botnet[1]. It's more traffic than you typically get from lingering on the HN front page for a day or so, and these aren't just favicon.ico requests, but actual search queries.
Websites submitted to HN regularly keel over from this amount of traffic. In any other business this would be a low grade DDoS attack. In search we call it Tuesday. Now I run my own hardware and the worst that happens is my site becomes slow or drops out. If I was cloud backed, I'd be destitute from paying the bills.
The usage limit of only 20 queries seems pretty low. That's barely enough for me to try this out and see what it can do.
It also seems that following some websites is difficult. For example, I wanted to follow Nomad List and Weather Spark, but neither came up in the list of options. I had to use a query to search for those sites before I could add them to my network.
>It also seems that following some websites is difficult. For example, I wanted to follow Nomad List and Weather Spark...
There is a feature on our roadmap where you can visit grep.help/nomadlist.com to follow the site, but it has not been shipped yet
It also seems like this is magic link authentication. It makes me bounce back and forth between my email client and the search engine I'm trying to use. Just let me set a password and my browser will auto-fill everything for secure one-click login.
Update, the quota of "10 free searches per month." is not enough for me to try out.
Still, interesting idea.
I'm so sick of services that hide their pricing or free trial restrictions until after you have signed up. These days, they are an immediate "NO" from me. If the team is that customer hostile from the start, how are they going to act later when you've become invested in their service?
Startups: Just be upfront about what you are offering and let me decide if it's worth it.
/rant
I say this to highlight that I really wasn’t expecting the pricing model and search method that’s in use here when I heard the site was “grep”
Here is the reason why we liked the name. Grep is a great CLI tool that always gives the right answers if your query is correct.
I wanted Grep.help to do the same for your personal network.
SoundCloud.help, say, a platform that leverages SC’s reputation in the DJ mixes community to advertise a platform for algorithmic audio compilations would be received rather poorly. If grep.help was some newbie-friendly url shortener for grep documentation or examples, I think we’d be more enthusiastic.
Currently it only work for a single user in self hosted manner.
In the future it may allow you to integrate the index of your trusted peer (similar to following another user on twitter).
The project is on https://github.com/beenotung/personal-search-engine
The service would take 2 regexes: 1 for the url and 1 for the content. It would then return the first 10,000 results of the pages that match both.
The latency would likely be in hours, and I don’t think advertising would work for monetization, so it would be a pay for use service, but for deep searches, this might be valuable.
A lot of work goes into making a search engine understand how a term relates to a website, and to rank the results accordingly.
You might get lucky and search for something that's in the goldie locks zone between underspecified and overspecified, and actually get good results, but in most cases you'd either get nothing or just noise.
Unix grep works well because you're typically grepping in hundreds or thousands of documents, not hundreds of millions or billions.
a) I want to import my bookmarks
b) I can't enter any of the domains I care about as I mostly have obscure tastes
due to the above two issues I was unable to actually see how the product felt.
Social booking site, rose to prominence a little over 10 years ago after del.icio.us was bought by Yahoo and started to suck. Wound up buying del.icio.us. The creator, Maciej Ceglowski, is a good writer with a blog at https://idlewords.com/
Very interested
It would have been much much smarter to choose a term that wasn't already broadly used. Even more mystifying is that someone would do this IF THEY KNEW ANYTHING ABOUT SEARCH ENGINES. Pick a different name, don't make your new offering compete with thousands of other search results explaining how GNU grep works.
An even bigger sin is the fact that grep is already used for searching and if even someone is looking for your site after forgetting the website, they are gonna have a fun time finding your offering vs scads of linux tutorials. This is BASIC marketing level SEO knowledge, you don't even need a search architect to tell you this info.
Good luck in any case. Do yourself a favor and pick another name if you want any traction.
I agree. The irony here is simply profound.
"Please don't complain about tangential annoyances—e.g. article or website formats, name collisions, or back-button breakage. They're too common to be interesting."
"Please don't pick the most provocative thing in an article or post to complain about in the thread. Find something interesting to respond to instead."
Anyway, I’m stopping there. I just thought this perspective worth mentioning.
First it was the term "crypto", then it's "LoRa", and now we have a "grep" that isn't.
This is quickly becoming a pet peeve of mine, as it's really polluting search results if you want to find the original things.
Unlike other disputes that people claim are ancient… the Apple dispute is still being fought, from Newton to the Beatles to genetically pure Honeycrisp.
Naming things is still the Lord’s most difficult commandment (probably a reference to RMS, I’m not sure anymore).
I'm not sure how to get that!
The first noun in Genesis is either "beginning" or "God", depending on whether you take בְּרֵאשִׁ֖ית as something like an adverb. (1:1)
The first noun mentioned as spoken is "light". (1:3)
The first noun mentioned as spoken in a conversation involving a human is "earth". (1:28)
The first noun mentioned as spoken by a human is either "time" or "bone", depending on whether you take הַפַּעַם as something like an adverb. (2:23)
(The above has been light-hearted gesturing to a story about naming things, and does not constitute actual theology.
Please consult your local serious person for further commentary; suggestion not valid in all nation-states)
;)
I hope this tool dies quick, and never even comes close to taking off. This is the only way the creator(s) of this tool learn in my honest, most humblest opinion.
Like, "crypto" is understandable, the technology is somewhat related and it was the non-technical people (who have very little exposure to the "cryptography" context) that shortened the long form into the current version, so not much that can be done by the tech side...
LoRa is a slightly worse but also understandable. The radio communication tech is Lo(ng) Ra(nge) (with just L and R capitalized). The ML stuff is Lo(w) R(ank) A(daptation) (with L, R, and A capitalized). Annoying, but probably still not intentional.
And then, there's "grep" here... Like, clearly the authors know about the older usage... and it's not like there's a case of some coincidental shortening...
Go is a very common word.
As a result people end up using rustlang and golang in searches, defeating the point of having a catchy name in the first place.