13 karma · joined March 22, 2016
I do like writing, and I can write. It just takes me a lot longer in English than in Italian, which is how I ended up handing the post over. But your point stands: if you already read that stuff all day at work, there's no reason to read more of it in your free time.
So I'll give it a go. The next write-up on Ferrox will be mine, rough English and all. Thanks for the comment.
What I actually want is MoE on machines that can't fit the model in VRAM, and specifically expert-level residency instead of layer offload: track which experts get hit during decode, keep those resident, evict the rest. Doing that well needs the router, the KV cache and the memory manager to be designed together, which is about the only good reason to write a runtime from scratch.
Yes, let’s see! You are welcome to contribute if you like!
The obvious question is “why, when llama.cpp already exists and is excellent.” The honest answer: I wanted to understand inference at a level deeper than “run the binary,” and I wanted a project where every performance claim had to be earned against a real, well-known baseline rather than asserted.
Key features:
- HNSW indexing with <1ms search times for 1M vectors
- ACID-compliant storage with Write-Ahead Log
- REST API + Python client
- Cross-platform releases (Linux, macOS, Windows)
- Perfect for RAG applications and local AI development
The motivation was simple: existing vector databases are either too complex for local development or too limited for production. VittoriaDB bridges that gap.