47 karma · joined October 24, 2025
Code: github.com/centamiv - Web: centamori.com - X: @centamiv (Italian content)
HNSW is just the indexing algorithm. It doesn't care where the vectors come from. You can generate them using Ollama (locally) HuggingFace, Gemini...
As long as you feed it an array of floats, it will index it. The dependency on OpenAI is purely in the example code, not in the engine logic.
Since it only stores the vectors, the actual size of the Markdown document is irrelevant; you just need to handle the embedding and chunking phases carefully (you can use a parser to extract code snippets).
RAM isn't an issue because I aim for random data access as much as possible. This avoids saturating PHP, since it wasn't exactly built for this kind of workload.
I'm glad you found the article and repo useful! If you use it and run into any problems, feel free to open an issue on GitHub.
I wrote a short post demonstrating how vector search works by building a D&D spell finder.
Instead of using a vector DB immediately, I implemented the cosine similarity math manually in PHP to demystify how embeddings actually work under the hood.
The code compares query vectors against a JSON dataset of spells to find matches by intent rather than keywords.