Startpage.com: Privacy-oriented search engine
startpage.com
startpage.com
Ref: https://www.reddit.com/r/privacy/comments/di5rn3/startpage_i...
- Startpage CEO Robert Beens discusses the investment from Privacy One / System1 [1]
- What is Startpage's relationship with Privacy One/System1 and what does this mean for my privacy protections? [2]
- What is the Startpage privacy-guarding data flow? [3]
Some further context [4].
[1] https://support.startpage.com/index.php?/Knowledgebase/Artic...
[2] https://support.startpage.com/index.php?/Knowledgebase/Artic...
[3] https://support.startpage.com/index.php?/Knowledgebase/Artic...
It's essentially a "proxy" search engine for many different ones. It has some really cool features aas well as a dark mode.
https://lifehacker.com/use-duckduckgo-lite-for-absurdly-fast...
Putting aside what happens if one allows Javascript and uses "modern" browsers, Startpage generally does not seem to require any more data from users than DDG. A small shell script can be used to search Startpage or DDG (or almost any other search engine) from the commandline without sending any unecessary data, like unecessary headers, cookies or hidden form variables. The best part is by not using the "modern" browser to send the search, one can easily automate editing the results page before viewing it in a browser, discarding all the cruft. I like to just return the URLs. (I notice that Startpage also (a) supports HTTP/1.1 pipelining, e.g., multiple page requests over a single TCP connection; Google does not and (b) allows bans to be overcome by solving an easily read captcha and this seems to prevent further bans; Google imposes automatic temporary bans that cannot be overcome by solving a captcha.)
The biggest problem I see with the major search engines and these minor search engines that repackage results from the major ones is that they are too often limiting the number of results returned. For example, Google limits to something like 200-300. In the early days of the web, search engines used to brag about how many pages were searched, and they proved their claim by how many results they returned. Today search engines want to localise and limit the results. Not to mention promoting their own websites. I also notice repeated searches where one is collecting the total results not simply the first page yield different results.
Not every query is a question and not every user is interested in an instantaneous "answer" or the most popular website. That type of quick searching certainly has its place but it is not "research" and will not lead users to learn much about what actually exists on the web, or how to think critically about the web's content. Some users may want to search for pages and then evaluate the pages themselves. Exploration and discovery. Those users are treated as "bots" in order to justify what can only be anti-competitive practices. The sad consequence of this "limiting" behaviour is to keep curious users from ever learning what actually exists on the web (versus what a "search engine" decides to promote, or demote).
I have been playing around with Common Crawl data and it seems woefully circumspect in its scope. A web index should be public information but these search engines sure as heck do not treat it as such.
But with or without a shell script, Startpage keeps your privacy. Our web app acts as a proxy between your endpoint and the rest of the web. We couldn't collect user data even if we wanted to. You deserve an explanation, and here's why:
Startpage is delighted that users are conscientious about our privacy practices, as they should be. Privacy-aware users are the kind of users we enjoy serving. Asking questions is always a good idea. As we’ve stated, System1 is interested in Startpage’s anonymous contextual advertising revenue, not in our data. Mainly because we don’t store any.
Even if they wanted to change our privacy policies, it wouldn’t be possible. Our co-owners and Surfboard Holding BV still have authority in our company. Our infrastructure is all in the European Union, where the strict GDPR legislation applies and the US Cloud law doesn’t.
Maintaining user privacy is our reason for being. We thank you for your curiosity and your vigilance. We hope you continue to ask questions and enjoy Startpage. We are gladly answering all your questions.
Even if they wanted to change our privacy policies, it wouldn’t be possible. Our co-owners and Surfboard Holding BV still have authority in our company. Our infrastructure is all in the European Union, where the strict GDPR legislation applies and the US Cloud law doesn’t.
Maintaining user privacy is our reason for being. We thank you for your curiosity and your vigilance. We hope you continue to ask questions and enjoy Startpage. We are gladly answering all your questions.
We support domain blocklist [3] natively and have !waves (similar to !bangs).
We're bootstrapped and not owned by an advertising company (startpage.com is owned by System1).
Just searched 4chan, no direct link to 4chan in first page, only "about" it.
Regarding results, I wrote an answer to that question here: https://news.ycombinator.com/item?id=25717814
That 4chan results page is embarrassing indeed. Currently many things in the pipeline, but will fix.
If you find any quirks, my email is in my profile.
- David
Paid plan has been on my mind for a while now.. and as you said, it's complicated. It's in the pipeline.
There's also things like https://coil.com/ who seem like they help support online content creators. I wonder if there's a way to treat search results like "content".
It is possible. I built the search engine [0] that was the first to integrate Coil as a monetization source. It is pretty small, but Coil payments do cover about 2% of the monthly cost to run the service.
Infinity Search also uses Coil. [1]
Here is an article with some thoughts around monetizing a privacy based search engine [2].
---------
[1] https://webmonetization.org/
[2] https://coil.com/p/runnaroo/Privacy-and-Search-Engine-Moneti...
It would be interesting to sell space against specific queries for a time duration vs per click or per impression.
This approach doesn’t lend itself as much to optimizing for every individual user action, but instead as the quality of the content and users as a whole.
I would focus on a few niches to start which have lots of ad spend that you could get a piece of. Maybe you could then pour those dollars directly into improving the organic results for those niches.
> Can't find what you're looking for.
js is not an option with many devices and useragents.
thank you for doing what you do.
In the pipeline!
[1] https://www.voxmarkets.co.uk/articles/kape-technologies-to-a...
In this case, your query is encrypted on the client side, passed through a proxy, decrypted at the engine, search is performed, and then results are encrypted, passed through the proxy, and the client side decrypts and displays the results.
USER Encrypted Search --- Proxy --- Search Engine Decrypts Search, Searches, Encrypts Search --- Proxy --- USER decrypts results and displays.
The search engine does not know your IP, and Private.SH does not know what you searched for.
but
"your query is encrypted on the client side"
and then
"the client side decrypts and displays the results"
So all this encryption/decryption code, where does it come from?
If the answer is Private.SH, then Private.SH can in fact know what the user searched for and the results they got by feeding the user code that sends that information (or even just the encryption keys) back to Private.SH
Also, I'm not clear on how the search engines are supposed to be able to decrypt something encrypted by the client. What actually happens there?
So you're using the search engine's public key to encrypt it, meaning the proxies can't decrypt it. But yes, you have to trust the client-side code, which is an insurmountable problem.
On the plus-side, the code is really short and easy to read. Perhaps a standalone app with reproducible builds could solve this, but that's much more of a pain than simply entering your query straight from the browser.
Edit: I was also going to mention that you can download the chrome/firefox extension by themselves, but the download link has an expired certificate which doesn't instill much confidence.
That depends on what you're trying to achieve, who you're willing to trust, and what you're willing to do.
If your goal is to do searches without having to trust client-side code from a search engine or Private.SH, then you could (assuming they have support for such a workflow) do your own encryption using a tool you do trust, such as gpg, then submit the encrypted query to Private.SH, which would hand it off to the search engine.
The search engine could then decrypt it, perform the query, and re-encrypt it to your public key (which would be contained in the encrypted query they got) and pass it back to Private.SH, which would then pass the encrypted query back to the user.
This way no code from Private.SH nor the search engine has to be trusted.
Of course, this does not help if Private.SH is secretly owned by, compromised by, or has a data-sharing agreement with some entity you don't want your data to be seen by (such as the search engine, hostile agency, data harvesting/reselling organization, etc).
This latter possibility is what I really don't see an easy way to mitigate.
For all we know any/all of these "privacy respecting" services might be owned by Google, Palantir, some other data harvesting corporation, government agency, intelligence service, etc.
If both are untrusted parties in cahoots with each other, then there's no getting around private.sh and gigablast sharing both the IP (courtesy of private.sh) and query (courtesy of gigablast) to each other.
For a time, I used https://brow.sh but its hosted html browser is not up anymore.
RMS has improved upon that if you are interested in privacy to that extent: https://lwn.net/Articles/262570/
You definitely need to whitelist cloudflare CDN and other popular CDNs like Amazon S3, things like jquery.com ,and maybe Google for the recaptcha (unless you whitelist google for individual websites)
YaCy: https://github.com/yacy/yacy_search_server (functional)
Seeks: https://github.com/beniz/seeks (defunct?)
---
There's also SearX, which isn't distributed but is a metasearch engine (pulls results from multiple search engines) that you can self-host [0] or use one of its many mirrors [1].
!sp
takes you to their home page.
!sp privacy
does that search on sp.
!sp duck duck go
does that search on sp.
EDIT: Ahem ...
!ddg
!ddg recursion