HNHacker News
TopNewBestAskShowJobs

antonellof

13 karma · joined March 22, 2016

submissionscomments
antonellof··on [dead]
Laya MLX typed decisions vs Dijkstra on seeded weighted mazes, side-by-side replay on Apple Silicon
antonellof··on Bonsai: A 27B reasoning model on a 16 GB M2 Mac, with Ferrox
Run a 27B reasoning model locally on a 16GB M2 Mac. Ferrox + ternary quantization delivers 5.9GB models without sacrificing much quality.
antonellof··on [dead]
Visual Git app inside your terminal: commit graph, diffs, stage by file or hunk, commit, branches, stash, push and pull, all with the mouse. Works in Ghostty, cmux, kitty and WezTerm, next to any CLI coding agent. One small Rust binary, no browser engine, works over SSH.
antonellof··on Building a Rust Inference Engine That Matches Llama.cpp
AI code quality is not bad, using vim and emacs do not elevate the code quality. do you prefer a typewriter or a computer? it's called future...
antonellof··on Building a Rust Inference Engine That Matches Llama.cpp
Contribute = Human ideas, testing, review etc. The monkey part of writing code by hand = obsolete. Do you still write code without and IDE? you remember all the programming language words? everything? you use stackoverflow? it's called evolution, btw: AI full disclosure This software is developed with strong assistance from Cursor, Grok 4.5, GPT 5.6, and Claude Fable 5, with humans leading the ideas, testing, and debugging. We say this openly because it shaped how the project was built. If you are not happy with AI-developed code, this software is not for you. The acknowledgement below is equally important: this would not exist without llama.cpp and GGML, largely written by hand.
antonellof··on Building a Rust Inference Engine That Matches Llama.cpp
That's a fair ask, and thanks for putting it that way.

I do like writing, and I can write. It just takes me a lot longer in English than in Italian, which is how I ended up handing the post over. But your point stands: if you already read that stuff all day at work, there's no reason to read more of it in your free time.

So I'll give it a go. The next write-up on Ferrox will be mine, rough English and all. Thanks for the comment.

antonellof··on Building a Rust Inference Engine That Matches Llama.cpp
Thanks for the comment. Parity with llama.cpp isn't my goal.

What I actually want is MoE on machines that can't fit the model in VRAM, and specifically expert-level residency instead of layer offload: track which experts get hit during decode, keep those resident, evict the rest. Doing that well needs the router, the KV cache and the memory manager to be designed together, which is about the only good reason to write a runtime from scratch.

Yes, let’s see! You are welcome to contribute if you like!

antonellof··on Building a Rust Inference Engine That Matches Llama.cpp
That’s the goal, an infinite loop of self improving inference engine that gets better and better by himself.
antonellof··on Building a Rust Inference Engine That Matches Llama.cpp
I have the skills, qualifications and experience to do this; ( see https://www.credly.com/users/antonello-fratepietro ) by using AI (I monitor what it does), I’m able to carry out a project like this! I’m a coder, not a blogger; it’s only natural that I use AI, just like everyone else, to write technical texts. I have over 20 years’ experience and I’m not ashamed to admit that I use Claude, Cursor and so on. The problem is not using it. It’s not the code or who writes it that matters, but the result: building my own inference engine. ( reply written with my brain ) ^_^
antonellof··on Building a Rust Inference Engine That Matches Llama.cpp
I’ve spent the last few days building Ferrox, a pure-Rust inference engine for running open LLMs locally — dense models and Mixture-of-Experts, on CPU, Apple Metal, or CUDA. No bindings to llama.cpp or ggml, no wrapping an existing runtime. Every kernel, every loader, every scheduling decision written from scratch.

The obvious question is “why, when llama.cpp already exists and is excellent.” The honest answer: I wanted to understand inference at a level deeper than “run the binary,” and I wanted a project where every performance claim had to be earned against a real, well-known baseline rather than asserted.

antonellof··on Show HN: Datapizza AI – Lightweight open source framework for GenAI apps
I tried it and created a full RAG application very easily. I followed the examples in the online documentation and it was really a piece of cake. My previous implementation of RAG (without Datapizza) used direct libraries such as OpenAI, Sentence Transformers, etc. Instead, I find it very useful to have a “framework” where you already have everything. In the past, I played around with Langchain, but I find it really immense, as complete and large as it is complex and distracting.
antonellof··on VittoriaDB: Zero-config local vector database in a single Go binary
Yes, thanks, good catch, i've just edit the gitignore, in case you are interested please open a PR or contribute, the project is very very fresh!
antonellof··on VittoriaDB: Zero-config local vector database in a single Go binary
I built VittoriaDB as a simple alternative to complex cloud vector databases. It's a single Go binary that works immediately after download - no Docker, no configuration, no cloud dependencies.

Key features:

- HNSW indexing with <1ms search times for 1M vectors

- ACID-compliant storage with Write-Ahead Log

- REST API + Python client

- Cross-platform releases (Linux, macOS, Windows)

- Perfect for RAG applications and local AI development

The motivation was simple: existing vector databases are either too complex for local development or too limited for production. VittoriaDB bridges that gap.