HNHacker News
TopNewBestAskShowJobs

fabmilo

79 karma · joined September 20, 2014

I am interested in helping software development teams and individuals turn plain-language specifications into production-ready code using AI. My expertise bridges the gap between intent and implementation, accelerating software development.
submissionscomments
fabmilo··on Ask HN: Who wants to be hired? (June 2026)
Location: San Francisco, CA, USA Remote: Yes Willing to relocate: No

Technologies: Python, Go, TypeScript, PyTorch

Linkedin: http://www.linkedin.com/in/fabmilo

Hi, I’m Fabrizio Milo, a senior AI/ML engineer, large-scale systems architect, and former technical co-founder. I’ve spent my career building production AI, ML infrastructure, and high-scale backend systems across startups and growth-stage companies.

Most recently, I’ve been building AI platforms for LLM training and fine-tuning, RAG databases, local and remote LLM inference, semantic retrieval, and agentic orchestration for business intelligence and code generation. I’ve also contributed to open source ML projects including GPT-Neo, TensorFlow and published research on synthetic data from LLMs.

Previously, I was VP of Technology at ZELIG, where I led the virtual try-on AI research roadmap and managed a cross-functional team building ML/3D systems for fashion retail. Before that, I was Head of Machine Learning Engineering at Recurrency, where I hired and led a 6-person ML/platform team and shipped demand forecasting, dynamic pricing, and recommendation systems on AWS/Snowflake/SageMaker. I also co-founded Passio, where I built the technical foundation for an on-device Nutrition-AI SDK with real-time computer vision inference.

Earlier in my career, I built scalable systems at TheRealReal and Scopely, optimized CUDA kernels at NVIDIA, and worked on real-time market-data and high-performance systems. I’m strongest where AI research, production engineering, and startup execution meet: taking ambiguous technical/product goals and turning them into shipped systems, teams, and infrastructure.

I can architect and build anything you need given enough compute and time.

fabmilo··on Why I Write (1946)
Writing it thinking. We developed our brain together with our hands. It feels slow but is actually faster for the end goal.
fabmilo··on Scaling Karpathy's Autoresearch: What Happens When the Agent Gets a GPU Cluster
I am fascinated by this example of using AI to improve AI. I won a small prize using this technique on helion kernels at a pytorch hackathon in SF.

The next step are: - give the agent the whole deep learning literature research and do tree search over the various ideas that have been proposed in the past. - have some distributed notepad that any of these agents can read and improve upon.

fabmilo··on Show HN: Hacker News archive (47M+ items, 11.6GB) as Parquet, updated every 5m
Was thinking the same thing. probably once a day would be more than enough. if you really want a minute by minute probably a delta file from the previous day should be more than enough.
fabmilo··on Show HN: Better Hub – A better GitHub experience
indeed. make a loom showing us why is better.
fabmilo··on How I've run major projects (2025)
There is tons of good advice. This blog post can be easily turned into a skill for agents.
fabmilo··on Generative Modeling via Drifting
New generative modeling using a single inference step
fabmilo··on The Waymo World Model
Very impressive work from Waymo. The driving with a tornado in the horizon example kind of struck my imagination, many people actually panic in such scenarios. I wonder though the compute requirements to run these simulations and producing so many data points.
fabmilo··on Flux 2 Klein pure C inference
because of the principle: you only understand what you can create. You think you know something until you have to re-create it from scratch.
fabmilo··on Ask HN: What Are You Working On? (Nov 2025)
VAE for real time video generation, WAN 2.1 / Matrix Game 2.0
fabmilo··on Coral NPU: A full-stack platform for Edge AI
How much would cost to produce these ?
fabmilo··on Show HN: FleetCode – Open-source UI for running multiple coding agents
nice, didn't knew this tool either
fabmilo··on Claude Sonnet 4.5
Yeah I totally agree, we need time to completion of each step and the number of steps, sizes of prompts, number of tools, ... and better visualization of each run and break down based on the difficulty of the task
fabmilo··on Ask HN: What are you working on? (September 2025)
How does it work? is just a documentation specification like spec kit?
fabmilo··on Show HN: I built an MCP server using Cloudflare's code mode pattern
I was just reflecting on this blog post after reading it this morning. What do you think on code mode after implementing it? At this point would not be better to just have a sandboxed api environment with customizable api/tools endpoints? basically an RL environment :)
fabmilo··on The AI coding trap
one axis that is missing from the discussion is how fast they are improving. We need ~35 years to get a senior software engineer (from birth to education to experience). These things are not even 3.5 years old. I am very interested in this space, if you are too dm me on X:@fabmilo I am in SF.
fabmilo··on AI for Scientific Search
I like zotero, I started vibe coding some integration for my workflow, the project is a bit clunky to build and iterate the development specially with gemini & claude. But I think that is the direction to take instead of reinvent from scratch something
fabmilo··on Show HN: Defuddle, an HTML-to-Markdown alternative to Readability
reference to the library: https://trafilatura.readthedocs.io/en/latest/

for the curious: Trafilatura means "extrusion" in Italian.

| This method creates a porous surface that distinguishes pasta trafilata for its extraordinary way of holding the sauce. search maccheroni trafilati vs maccheroni lisci :)

(btw I think you meant trafilatura not trifatura)

fabmilo··on A Research Preview of Codex
more excited about the rust impl than the typescript one.
fabmilo··on Intellect-2 Release: The First 32B Model Trained Through Globally Distributed RL
The interesting delta here is that this proves that we can distribute the training and get a functioning model. The scaling factor is way bigger than datacenters
fabmilo··on Multi-Token Attention
I read the paper and the results don't really convince me that is the case. But the problem still remains of being able to use information from different part of the model without squishing it to a single value with the softmax.
fabmilo··on Multi-Token Attention
We have to move past tokenization for the next leap in capabilities. All this work done on tokens, specially in the RL optimization contest, is just local optimization alchemy.
fabmilo··on LIMO: Less Is More for Reasoning
I will believe reasoning architectures when the model knows how to store parametric information in an external memory out of the training loop.
fabmilo··on DeepSeek-R1
was genuinely excited when I read this but the github repo does not have any code.
fabmilo··on rStar-Math: Small LLMs Can Master Math Reasoning with Self-Evolved Deep Thinking
I was just about to submit this link and redirected me to this page. I am shocked that it received only four comments. If you are working in the LLMs/Agent space ( you are, right?) and you don't understand the significance of this paper, you are set for failure.
fabmilo··on Happy New Year 2025
Happy new year to everyone, hacker news is more than my home page. This community is awesome!
fabmilo··on Ask HN: Why isn't Alex Krizhevsky as famous as Ilya Sutskever or G.Hinton?
Doesn't make enough drama.
fabmilo··on Show HN: Open-sourcing my failed startup Buzee – A file search application
Thanks to make this open source! Doesn't seem to have any AI enabled search feature does it? Honestly I think you gave up too early, there are solid foundations in the app but definitely needs some more polishing for practical use. I managed to build it on my mac. npm install, npm run tauri build. What is your workflow to add features / debug?
fabmilo··on Byte Latent Transformer: Patches Scale Better Than Tokens
I am gonna read this paper and the other latent sentence later today. I always advocated for this kind of solutions together with latent sentence search should get to the next level of AI. Amazing work from Meta
fabmilo··on Training LLMs to Reason in a Continuous Latent Space
from my understanding that is what they do, see the paper: > We use a pre-trained GPT-2 (Radford et al., 2019) as the base model for all experiments. I agree the feedback is necessary, and the mechanism simple and cheap, but I don't think is optimal.
Page 1 of 3Next →