HNHacker News
TopNewBestAskShowJobs

anthonyhn

258 karma · joined July 27, 2022

Hello,

I run an search engine at https://ichi.do. Please feel free to test it out.

If you would like to get in touch, please feel free to contact me through github (https://github.com/anthmn) or send me an email (anthony.m.mancini@protonmail.com).

submissionscomments
anthonyhn··on Block AI bots, scrapers and crawlers with a single click
First off, I want to thank you and the other members of the CC Foundation, the CC data set is an incredible resource to everyone.

Much of the UA data, including CCBot, is from an upstream source[0]. I was torn on whether CCBot and other archival bots should be included in the configs, since these services are not AI bot scraping services. I've added an exclusion for CCBot[1] and the archival services from the recommended configs.

[0] https://darkvisitors.com/agents/ccbot

[1] https://github.com/anthmn/ai-bot-blocker/commit/ae0c2c40fd08...

anthonyhn··on Block AI bots, scrapers and crawlers with a single click
For those not using cloudflare but who have access to web server config files and want to block AI bots, I put together a set of prebuilt configs[0] (for Apache, Nginx, Lighttpd, and Caddy) that will block most AI bots from scraping contents. The configs are built on top of public data sources[1] with various adjustments.

[0] https://github.com/anthmn/ai-bot-blocker

[1] https://darkvisitors.com/

anthonyhn··on What We're Working on in Firefox
> search engine such an easy thing to produce?

Yes, I run my own independent search engine[0].

> successful

Now that's the challenging part, especially since Mozilla needs to fund browser development. The initial differentiation for a Mozilla search would have been difficult, but at their peak they had 30%+ market share, and if the default search on Firefox was Mozilla Search then they might have been able to make it all work financially. DuckDuckGo makes over $100 million in revenue per year[1], if a Mozilla Search made as much money annually, then even though it's below their current $500 million search contract with Google, with some fiscal responsibility they probably would have had enough to support a search engine and browser development concurrently.

[0] https://ichi.do/

[1] https://techreport.com/statistics/software-web/duckduckgo-st...

anthonyhn··on OS/2 Warp, PowerPC Edition
>The biggest issue is that an ancient version of Firefox is the only viable option for a web browser.

There was some work done on porting Palemoon, Otter Browser, and QtWebEngine to ArcaOS. A preview version of a QtWebEngine-based browser[0] was available for download and may work with newer websites (given QtWebEngine is based on Chromium), but I have not tested it myself.

[0] https://www.os2world.com/cms/index.php/past-news/80-news/sof...

anthonyhn··on Twenty years of blogging
> But I always find I have nothing to speak about.

Odds are you do have something interesting to speak about. Many people are experts in very niche things and don't even realize it. You may be very proficient with a niche piece of software that is not well documented, or may have created software to solve a very specific problem. Writing blog posts about your niche knowledge can be tremendously helpful; I can't tell you how many times a single blog post about an obscure problem has saved me hours (possibly even days or weeks) of research when I've encountered the same problem.

anthonyhn··on Ask HN: Go to Websites for Programmers / Hackers in 2023?
>Wish I had an invite.

They have an IRC chat room where you can request an invite[0]. They typically ask if you have a personal website, git repo, interesting projects, etc when inviting people.

[0] https://lobste.rs/chat

anthonyhn··on Ask HN: Why do apps use Lisp/Scheme as a scripting language instead of Python?
Lisp has very minimal syntax and so you can build a lisp interpreter fairly easily in many languages. Most languages don't have a readily available python interpreter that you can embed, and building one that is feature parity with CPython (the most popular python implementation) is not an easy task.
anthonyhn··on Ask HN: How do you document/index random thoughts, observations, & learnings?
Wikipedia has a page on single-user personal wikis that are great for organizing your notes[0].

If you're looking for something simple and unixy, Zim is a pretty good choice[1]. It's an offline GTK-based GUI application for creating personal wikis that saves all the wiki pages as Markdown files and can export your wiki as HTML using various templates. Zim has been available in many linux package repos for over a decade and is GPL-2.0 licensed.

[0] https://en.wikipedia.org/wiki/Personal_wiki?useskin=vector#S...

[1] https://zim-wiki.org/

anthonyhn··on Ask HN: Anyone else frustrated with “secure connection checks”?
Yes, it can be quite frustrating. On my search engine[0] I tag sites that use captcha services and browser checks so that the user has more transparency and can choose to avoid sites that use these services if they want.

[0] https://ichi.do

anthonyhn··on Ask HN: Are you worried your non-AI project might soon be obsolete?
I work on a search engine in my spare time as a side project. What I've learned from working on a search engine is that even though the recent GPT models are quite good for general purpose search, there are still ample opportunities in search, and generally not enough people looking at search to cover everything.
anthonyhn··on Ask HN: What happened to interest in FP recently?
If you don't mind me asking, what do you use Haskell for at work? Also was the company using Haskell before, or did you bring Haskell to the job? It's always fascinating hearing stories about people using less common languages like Common Lisp or Haskell on the job!
anthonyhn··on Ask HN: What's the best obscure but useful CLI tool?
GNU Parallel is pretty useful. You can use it to run command line programs in parallel. For example, if you wanted to gzip compress all individual files in a directory and limit the number of jobs equal to the number of CPU cores on your machine:

ls | parallel --dry-run --jobs `nproc` 'gzip {}'

(Note you have to remove the --dry-run switch to actually run this, --dry-run only prints out the commands parallel will run, not actually execute them)

anthonyhn··on Vim or Emacs for C++ Coding?
Just haven't tested it out yet, but am interested in testing out some different plugins and combinations in the future. Also neovim wasn't available in one of the platforms I use, so have been sticking with vim.
anthonyhn··on Vim or Emacs for C++ Coding?
I use vim for C++ coding, however it is a bit difficult to set up to make it productive. I use YouCompleteMe [0] for autocompletion, Vimspector [1] with the C++ plugin for debugging, ALE [2] for linting, along with a few other general plugins (such as NerdTREE for file view).

[0] https://github.com/ycm-core/YouCompleteMe

[1] https://github.com/puremourning/vimspector

[2] https://github.com/dense-analysis/ale

anthonyhn··on XUL Layout is gone
The Palemoon browser [0] also still uses XUL, and is in many ways a continuation of XUL browsers (was originally forked from FF 29, updated with various components from FF 50+, and with many other tweaks).

[0] https://palemoon.org

anthonyhn··on Ask HN: What's the best TTS engine you've heard?
For offline/local TTS, Coqui TTS [0] is quite good. It's essentially a continuation of Mozilla's TTS engine that Mozilla stopped working on ~2 years ago (and IIRC it's largely the same team that worked on Mozilla TTS).

[0] https://github.com/coqui-ai/TTS

anthonyhn··on Ask HN: RHEL vs. Ubuntu and Debian relevance in 2023?
Alpine Linux is also pretty relevant in the container space due to its small base image size. If you look at the downloads on Docker Hub [0], Alpine Linux, Ubuntu, Debian, and CentOS (although CentOS was deprecated by Red Hat) all have 1 billion+ downloads.

[0] https://hub.docker.com/search?q=

anthonyhn··on Ask HN: Is Firefox really privacy friendly?
You can disable some of firefox's background network requests by modifying the about:config key/value pairs of your firefox profile (for example, by using a user.js file). The Arkenfox user.js (https://github.com/arkenfox/user.js/) has some pretty good defaults that disable a lot of the background requests.

There are also firefox forks that disable background requests. Librewolf (https://librewolf.net/) is a popular fork that uses a combination of about:config tweaks, policies, and patches to disable all background requests (technically some requests still go through, but are replaced with a dummy url that you can block). Librewolf also downloads uBlock Origin by default, and is close to upstream. Overall it's a pretty good out-of-the-box solution.

anthonyhn··on Ask HN: Which Python Type Checker?
I typically use mypy. It's been around for over a decade, is available in many package repos, and has pretty good integration with vim lint plugins such as ALE (https://github.com/dense-analysis/ale).
anthonyhn··on Show HN: Ichido, search engine that tags sites using Google and Cloudflare
> I see you offer an opensearch.xml already - if you embed it as link node with the appropriate type it will be straightforward to add it to the browser as (default) search engine

Thanks for the heads up, I used to have a <link rel="search"> to the opensearch in a prior iteration of the site, must have removed it by mistake. Will add in the link in the next release.

anthonyhn··on Show HN: Ichido, search engine that tags sites using Google and Cloudflare
>Is it easy for you to rely on more search index providers, what are your options?

I have a few options:

* Switch index providers. For example, Mojeek has an index with 6bn pages and has a web search API; may be more sustainable to switch to their index in the long run.

* Build my own index. This is my preferred option and I've already started to work on this.

* Look for funding sources to offset the price hike.

anthonyhn··on Show HN: Ichido, search engine that tags sites using Google and Cloudflare
EDIT: I found some of the pages with links that include UTM tracking params. Let me know if you want me to send you the pages with those links, can send them through email (my email is on the contact page of the site).
anthonyhn··on Show HN: Ichido, search engine that tags sites using Google and Cloudflare
> what's wrong with webp?

Nothing wrong with the format in particular. However some may prefer formats such as PNG and JPEG since:

* A lot more software supports PNG and JPEG (backwards compatibility, better integration with one's existing system and tools).

* You can often get the same file size, visual quality, and performance with PNG and JPEG as you can with WEBP with optimization.

anthonyhn··on Show HN: Ichido, search engine that tags sites using Google and Cloudflare
Since spacehey includes user-submitted content, it's possible that:

* Someone uploaded a WEBP image to the site.

* Someone pasted a link with a utm_* param.

* The page was crawled when cloudflare was used.

Will look into it and see if I can find the pages that generated the tags. Search results are generally tagged by domain name (necessary since not all pages can be crawled, and even if the page the user connects to doesn't have, for example google trackers, a user would likely want to know if the site is using trackers elsewhere).

Also love the spacehey project, really captures the feel of Myspace!

anthonyhn··on Show HN: Ichido, search engine that tags sites using Google and Cloudflare
>Just added Ichido.

Thanks, much appreciated

anthonyhn··on Show HN: Ichido, search engine that tags sites using Google and Cloudflare
It's $4/1000 queries, but the rate is increasing in May to $18/1000 queries. The Bing API is available through Azure.
anthonyhn··on Show HN: Ichido, search engine that tags sites using Google and Cloudflare
>Edit: looks like the file type filter is dropped as well. Do add the arguments to the pagination links.

Thank you, great feedback! You're right, I forgot to include some of the params in the pagination, will have to include those in the next update.

anthonyhn··on Show HN: I made an early 2000s-inspired internet forum
>I don't remember which software it was but one of them would show, in the footer, a count of MySQL queries run.

It might have been Simple Machine Forum (SMF). The footer has the number of queries and number of seconds it took to generate the page (in my case, their community forum index page used 9 queries and took ~0.3 seconds to generate).

https://www.simplemachines.org/community/index.php#footer

anthonyhn··on The beauty of CGI and simple design
>The python folks thought so little of CGI that they ripped it out of the standard library.

IIRC, the issue was that they didn't have a maintainer for the cgi library anymore, and that is why they removed it.

anthonyhn··on Ask HN: Brave or Firefox and uBlock / Other Ext – Better?
>Seems like Firefox would be better but more work(?).

If you keep copy of your firefox profiles folder, its not too much work to maintain. Most of the work is an upfront cost of setting up the initial profile.

There is also Librewolf, a Firefox fork which includes uBlock by default and has other security hardened features.

>And now we have DuckDuckGo entering the scene and I'm wondering how they all stack up.

IIRC, DDG is using the system webview for their browser, so the rendering may differ depending on the platform.

Page 1 of 2Next →