HNHacker News
TopNewBestAskShowJobs

vforno

416 karma · joined May 18, 2026

submissionscomments
vforno··on Brio mode in Colibri: Scoring a closed set instead of generating
Hi everyone, I’m Vincenzo, the founder of Colibrì. Today, we are adding a new mode to Colibrì called "Brio," accessible via both the TUI and the web interface. We were all impressed by Jev and its potential, so I decided to create a mode that allows you to use any model supported by Colibrì—in both chat and Brio modes. Brio returns a probability for each choice and the entropy of that distribution. No output tokens are generated, but inference work still takes place: the engine processes the context and scores the provided option tokens. You can use it via the dashboard, the terminal, or the POST /v1/brio endpoint.
vforno··on Show HN: OpenVurp – An open-source alternative to Grok Bot
Vurp means octopus in Taranto dialect
vforno··on Show HN: Openvurp – A wallet of AI agents that use tools and consult each other
Hi HN — I built openvurp because I was already using AI agents, but I couldn’t find a setup that made them fast and natural enough to use in my personal life.

For a while I had been thinking about a “wallet” of agents: one place where I could keep agents with specific jobs, choose a different engine for each one, and let them ask one another for help when a question falls outside their role.

In openvurp, each agent has a name, a job, and an engine. It runs on your computer, but the engine can be local or remote — Codex, Claude, Ollama, or an API. Agents can use real tools such as the shell, files, web search, and a browser, while their conversations and memory remain as files you can inspect and delete.

The project is still early. I’m sharing it because I want to understand whether this way of organizing agents could be useful outside my own setup too.

Why “Vurp”? In the dialect of Taranto it means octopus. I chose the name as a small tribute to Taranto, my hometown and its sea.

vforno··on Show HN: Lumabri – Run Moe Models on a P2P Swarm with Colibri
Thanks really thanks for support!
vforno··on Show HN: Lumabri – Run Moe Models on a P2P Swarm with Colibri
Because everything stays local and latency is tiny, even modest always-on devices can contribute. A few Raspberry Pi 5s, old mini-PCs, or stronger IoT-style boards can each hold and run a handful of experts. The protocol doesn’t care if the peer is a big GPU or a small ARM box, as long as it can load the expert weights and do the matmul. Pure busybox-class sensors are usually too limited in RAM and compute for current MoE experts, but the broader “every half-decent always-on box in the house joins the swarm” vision works well and keeps everything private inside your LAN.
vforno··on Show HN: Lumabri – Run Moe Models on a P2P Swarm with Colibri
That’s one of the strongest use-cases. On a local network (office, lab, home cluster) the RTT is a few milliseconds instead of 20-50 ms, so the expert-offloading becomes much more practical. You can spread the experts across several cheaper GPUs or even CPUs, keep only the dense parts + router on the machine you’re chatting from, and the whole thing stays private inside your LAN. No internet required, no cloud, just the machines you already have.
vforno··on Show HN: Lumabri – Run Moe Models on a P2P Swarm with Colibri
Ho thanks for the comment. Verification does not depend on temperature. Expert execution is deterministic (pure matmul). LUMABRI_VERIFY=N re-runs N% of the calls on a second replica and requires byte-identical output. Temperature (and sampling) happens only on the chatter, after the experts return their activations. So it can be any value (0, 0.7, 1.2…) without affecting the verification contract.
vforno··on Show HN: Lumabri – Run Moe Models on a P2P Swarm with Colibri
Thanks for the support!
vforno··on Show HN: Lumabri – What if LLMs worked like Napster?
It’s a wonderful world because everyone has their own thoughts on how to handle various situations. I respect your point of view. Thank you very much for your time.
vforno··on Show HN: Lumabri – What if LLMs worked like Napster?
Thanks for the comment. Extremely clear and detailed. To write this project, I researched a lot about methods and how peers should work compared to a server, as well as security concepts, which certainly have more to add. The goal for this type of project was to move from a single-machine Colibri to multiple-machine Lumabri on a LAN to a large number of machines working together in a Napster-style P2P Lumabri. For latency and other issues, we are studying every type of method that can improve it, and we are also writing and testing other things on our test server. Thank you very much, we will continue to improve.
vforno··on Show HN: Lumabri – What if LLMs worked like Napster?
Actually, no, in this case it would be possible that if other peers in the network give up processing or space you would have a speed that you wouldn't have as a single computer.
vforno··on Show HN: Lumabri – What if LLMs worked like Napster?
Adding information for this type of project might make the readme seem overloaded with information, but the goal was to show each test and explain it as best as possible. I'll work on it, thank you very much.
vforno··on Show HN: Lumabri – What if LLMs worked like Napster?
absolutely controls and other things will be part of everything for this type of project. Thanks for the comment and the goal is definitely to improve it more and more.
vforno··on Show HN: Getting GLM 5.2 running on my slow computer
Thanks for this benchmark have you used the last commit? We made some changes from mtp int 4 to much more! If you like, leave an issue so we can work on it! Thank you so much!
vforno··on Show HN: Getting GLM 5.2 running on my slow computer
Yes you can use ./coli chat
vforno··on Show HN: Getting GLM 5.2 running on my slow computer
Thanks for kind words!
vforno··on Show HN: Getting GLM 5.2 running on my slow computer
Maybe some from intel can read and we can try? :)
vforno··on Show HN: Getting GLM 5.2 running on my slow computer
Yes accurate!
vforno··on Show HN: Getting GLM 5.2 running on my slow computer
Really thanks!!
vforno··on Show HN: Getting GLM 5.2 running on my slow computer
Thanks We're working on it!
vforno··on Show HN: Getting GLM 5.2 running on my slow computer
I have a small laptop. If you have more disks available, you could really do some testing. When you have some benchmarks, submit a pull request or issue so we can maybe work on them. We are really happy for contribute!
vforno··on Show HN: Getting GLM 5.2 running on my slow computer
We're working on it right now with a pull request that will also arrive for opencode!
vforno··on Show HN: Getting GLM 5.2 running on my slow computer
That's possibly a good idea! We can work on it!
vforno··on Show HN: Getting GLM 5.2 running on my slow computer
Thanks really thanks!
vforno··on Show HN: Getting GLM 5.2 running on my slow computer
if you like, colibrì always needs to improve so if you have ideas or anything else you are welcome for pull request issues and also benchmarks!
vforno··on Show HN: Getting GLM 5.2 running on my slow computer
It’s a good question.

In theory MPI could distribute experts across nodes. In practice, for small clusters the added network latency usually hurts more than it helps.

Better suited for big clusters with fast interconnects. For now we're focusing on single-machine speed (caching, GPU hybrid, etc.).

vforno··on Show HN: Getting GLM 5.2 running on my slow computer
Really thanks!
vforno··on Show HN: Getting GLM 5.2 running on my slow computer
No because I have only 32gb of ram too low
vforno··on Show HN: Getting GLM 5.2 running on my slow computer
Maybe we can see some integration!
vforno··on Show HN: Getting GLM 5.2 running on my slow computer
Really thanks!!
Page 1 of 2Next →