I built a 500k-domain search engine for makers in a weekend for $10
alexmorleyfinch.github.io
alexmorleyfinch.github.io
1. read each site
2. rent a 4090 with https://vast.ai to run vllm
3. let llm model invent its own category and tag names freely
4. save 1KB of metadata each
a. a small local language model that reads each one and writes a name, two or three sentences, a category, and a handful of tags.
5. `code is going up as open source` soon (TM)Your impressions seem about right, but there are a few control steps it seems.
Since when has low effort become a selling point anyhow?
6. let llm write a blog post about this conversation
I have a 400 GB sqlite database with samples of rendered root document DOMs I use for ad detection in Marginalia Search I've been meaning to explore similar ideas using.
What a homebrewed solution lacks in coverage it excels in indexing and serving a small slice of the internet really really well.
Almost anyone can put together a basic search engine in a few thousand lines of code, it's just not very hard to make a program that will index a few million documents better than Confluence.
Then, between that first ansatz and a working scalable internet search engine, you have a pile of interesting problems touching every aspect of computer science and computer hardware and networking, enough so that hundreds of people will have gotten PhDs in narrow sub-problems of those problems you'll be facing.
It's great because you can just tackle the stuff you feel comfortable approaching and leave the rest for later.
There's about 200 million active domains currently. That's about 66% of all businesses worldwide of which there are around 300 million. Around 100 to 150 million have active webpages
Where do you get stats on this info? Not refuting your claims but it's interesting data. I would've thought a wider gap in all biz to active websites since some countries run nearly all their biz on whatsapp/telegram/wechat.
I did count the number of ICANN domains on their zone files on czds and there are a few services selling pre compiled domain lists that do crawling for marketers (it is forbidden by ICANN to use zone files for marketing) I don't remember how I got the active website count of 100 to 150 million, maybe I interpolated by checking some ICANN domains or there's a domain list service that does it. Domain list services usually say how many domains are active regardless of whether they have a website or not
Of course parked domains add up to the active website count even though they have no relevant content
Subject: I want all domains and subdomains https://groups.google.com/g/common-crawl/c/XC2QmOE-sdI?pli=1
or google for COMMON CRAWL
A few years ago, if you wanted translation, you'd use Google Translate. If you wanted to search the web, you'd use Google search.
But for a few gigabytes, you can now install nllb-200-distilled-600M, and get translations for almost any language locally. You can have your computer crawl the web, create abstracts and categorizations for websites, and build search exactly as you want it.
The main limiter now is hard drive space (and to an extent, local compute) -- but right now it feels like the 70s again where the terminal into a remote server turned into building applications locally.
You cannot just let a model run wild.
- why dont you tell me how you ll categorize this list with k-means?
And just like that, you recreated the internet of the 90s - early 00s. Brilliant!
Checks socials and trademark too
citation needed