Abusing AWS Lambda to make an Aussie search engine
boyter.org
boyter.org
A couple thoughts occurred to me as I read the post:
- Lambda functions deployed using Docker images can be up to 10GB.[1] Would that change your math here? I'm curious what the tradeoff would be vs parallelizing more function executions searching smaller datasets on both cost and performance.
- Great notes on the anti-competitive nature of the current market. If there was an open standard on crawling, maybe we'd see more innovation here.
- Cool use of a bloom filter!
1. https://aws.amazon.com/blogs/compute/working-with-lambda-lay...
> Incidentally searching around for prior art I found this blog post https://www.morling.dev/blog/how-i-built-a-serverless-search... about building something similar using lucene, but without storing the content and only on a single lambda.
See also:
Sqlite FTS (no server): https://news.ycombinator.com/item?id=27016630
EdgeSearch (Workers): https://github.com/wilsonzlin/edgesearch
EdgeSql (Workers): https://github.com/lspgn/edge-sql
InfiniCache (Lambda): https://news.ycombinator.com/item?id=25788893
quickwit (Lambda): https://news.ycombinator.com/item?id=27074481
> I had been working on a bloom filter based index based on the ideas of bitfunnel which was developed by Bob Goodwin, Michael Hopcroft, Dan Luu, Alex Clemmer, Mihaela Curmei, Sameh Elnikety and Yuxiong He and used in Microsoft Bing https://danluu.com/bitfunnel-sigir.pdf.
Reminds me of wavelet trees: https://alexbowe.com/succinct-debruijn-graphs/ https://alexbowe.com/wavelet-trees/
I tried 2 queries. Vegemite worked great:
https://bonzamate.com.au/?q=vegemite
The (unofficial) national anthem, less so.
https://bonzamate.com.au/?q=%22men+at+work%22+%22down+under%...
Have you seen the CommonCrawl dataset? It's a great source if you're looking for a full-web index to do some Big Data analysis on, and scale up your search engine to something closer to DuckDuckGo. The CommonCrawl data is 320 TB in size though, so I hope you've got plenty of free disk space!
You can also add "Show HN" to your titles in future, if it's a project you worked on. Then it'll show up in the "show" subset (linked from the orange top bar) and usually attract more encouragement because Hacker News readers usually really like original authors coming to hang out :)
I have seen common crawl. It's one of those things I should look into more when I get the time. Honestly thats the limiting factor for everything I do these days.
Supposedly stating you're a custom crawler [1] (like how Google has GoogleBot) and not a generic bot helps with this. I'm somehow very doubtful of this.
[1] JGC responded to a tweet a long time ago, can't find it now
But on a serious note - you can't exactly rely on Cloudflare to not bring down a business. For me, it is a very nice convenience to throw up in front of my WP sites for caching and to help limit the bad traffic I get, but that's about it
On the non-tech side:
>"In February 2021 the Australian Greens Party called for a publicly owned search engine to be created and be independent and accountable like the ABC."
Which is an interesting idea, given search has become something of a utility.
I think it should be picked up in your browser already as there is the open search definition in there. If you are using Chrome and visit chrome://settings/searchEngines you should see it already in other search engines, but the query URL needed is https://bonzamate.com.au/?q=%s
1. Put it in a layer instead. A bit more space and easier to maintain.
2. Put it in EFS and cache to /tmp. You get almost a half gig of temp storage which persists across executions.
I do know that the same index on a physical machine can process 200 million items in ~100ms just to give an idea. On lambda it was more like 100,000.
Keep in mind I stopped looking at that point. If someone is able to prove otherwise id be happy to switch over to this, but then again I don't have a real need currently.
When running locally on your physical machine, is your code only utilizing a single CPU core or does it make use of multiple cores? If it's the latter one, increasing the memory for the AWS Lambda function even further should improve performance as well.
Also if you hit a performance plateau with a memory size of less than 1769MB for your AWS Lambda functions, you could be bottlenecked by something else. I can imagine details like memory bandwidth playing a role there as well.
[1]: https://docs.aws.amazon.com/lambda/latest/dg/configuration-f...
My guess based on my tests is that memory access and cpu are the limiting factors in lambda for this but that because it encourages scale out less of an issue.
Pretty fast too, even viewing from the UK
I might break that rule for bom though. Although I wish they would just let myself add a cert for them. I'll even do it for free if they let me!