Uncurled – running and maintaining Open Source projects for three decades
un.curl.dev
un.curl.dev
I've been considering open sourcing my search engine. Search is like a fractal of interesting problems, and pretty much every aspect of the search engine has known areas of improvement, so I'm sure it would be a fun project to collaborate on.
I'm honestly a bit at a loss how to actually go about it, since it's not an application or a library where others are expected to run it, but a fairly bespoke piece of web service that requires specific hardware configurations to do anything useful (as well as extremely unwieldy datasets). The only rolemodel I can find is something like Wikipedia.
I'm curious if anyone knows good "role models"?
My email is on my profile if you'd like to chat directly. That offer extends to anybody else, too!
[0]: https://github.com/lunasec-io/lunasec/
[1]: https://www.lunasec.io/docs/blog/how-to-build-an-open-source...
Because if the answer is "knowledge sharing" then just document it and open it, that's it, whoever is interested will show up (and if nobody does you don't care, your goal is fulfilled). If instead your goal is to eventually build a community around it then you'll have to put more effort (for example talk in conferences).
You get my point, you will know which road to take as soon as you know where you want to go.
My 2c.
The problem is the logistics of an open search project that is a web service with serious hardware requirements. I think the minimum hardware requirements is about 14 Gb of RAM, that's without any real data loaded into the system. Testing is awkward and cumbersome, and the data logistics are a real headache even on the same network as the production instance.
To even run the search engine, you need a few hundred megabytes of language models, as well as a probably few gigabytes of website data to conduct meaningful testing. The production instance has a disk footprint of about half a terabyte.
My philosophy to open source is just do my personal projects in the open even if it’s not “accessible” and if someone comes along and likes it they can contribute or fork it and modify to their liking. Or they might just browse the code and pick up something in passing.
My ignorant suggestion towards sharing those gigabytes of data: have you considered bittorrents?