HNHacker News
TopNewBestAskShowJobs

carloslfu

120 karma · joined April 13, 2017

Working on Executable Rationality and efficient AI. https://carlosgalarza.com
submissionscomments
carloslfu··on 40 years of writing code. I didn't see the end coming
what does your company do?
carloslfu··on Show HN: Running 104GB Qwen3.8-Flash-Next on 48GB Mac with at ~12 tok/s
It wasn't either/or, the N-gram table is part of Qwen itself and stays on disk. I’ve now added its 1.5GB MTP draft head too, it gets 86% acceptance and about 1.24× faster decoding on my 48GB Mac.
carloslfu··on Show HN: Running 104GB Qwen3.8-Flash-Next on 48GB Mac with at ~12 tok/s
thanks! in part, I was wondering how you got the code into files. I guess you copy pasted it inside a file, am I right?
carloslfu··on Show HN: Running 104GB Qwen3.8-Flash-Next on 48GB Mac with at ~12 tok/s
This is next! in the works rn.
carloslfu··on Show HN: Running 104GB Qwen3.8-Flash-Next on 48GB Mac with at ~12 tok/s
thanks!

> ditched oLlama"

yeah! this is interesting.

> 8.1GB per slotserve process is a lot! Is that in your control?

yes, it is hard, but I agree the smaller the better. I'll work on that

> If it's local and open-weight, this could be marketed this way I think.

I like this!

> what's the actual, real use case for slotserve?

I'm working rn on an app on top of it that closes the loop and is a fully local AI app, an experiment. I'll publish it as soon as it is usable!

> built a small html hello world served via Python

What did you use as a harness here?

carloslfu··on Show HN: Running 104GB Qwen3.8-Flash-Next on 48GB Mac with at ~12 tok/s
great idea!! a native app would be awesome
carloslfu··on Show HN: Running 104GB Qwen3.8-Flash-Next on 48GB Mac with at ~12 tok/s
Both projects are different in scope. Think of slotstream as optimizing for memory and for this specific model for now, my intention is not to build an inference engine the same as oMLX
carloslfu··on Show HN: Running 104GB Qwen3.8-Flash-Next on 48GB Mac with at ~12 tok/s
Interesting! I'll check it out
carloslfu··on Show HN: Running 104GB Qwen3.8-Flash-Next on 48GB Mac with at ~12 tok/s
true
carloslfu··on Show HN: Running 104GB Qwen3.8-Flash-Next on 48GB Mac with at ~12 tok/s
for the record, I'm almost 35
carloslfu··on Show HN: Running 104GB Qwen3.8-Flash-Next on 48GB Mac with at ~12 tok/s
interesting! Yes, thermal is important. Pretty cool project man! Starred and checking it out!
carloslfu··on Show HN: Running 104GB Qwen3.8-Flash-Next on 48GB Mac with at ~12 tok/s
This is the best I could find: https://huggingface.co/Qwen/Qwen3.8-Flash-Next?utm_source=ch...

About the specifics, I have only anecdotal evidence, but I guess this info can be found somewhere

carloslfu··on Show HN: Running 104GB Qwen3.8-Flash-Next on 48GB Mac with at ~12 tok/s
I hope not! this is a new macbook lol!
carloslfu··on Show HN: Running 104GB Qwen3.8-Flash-Next on 48GB Mac with at ~12 tok/s
I don't know actually. I'll check haha. My best guess is it isn't.
carloslfu··on Show HN: Running 104GB Qwen3.8-Flash-Next on 48GB Mac with at ~12 tok/s
I see! yes, downloading the weights part is painful. I tried a couple fixes and it is as fast as it can get downloading from HuggingFace. I think the field is heading toward smaller, more capable models soon, so you won't have to wait that long!
carloslfu··on Show HN: Running 104GB Qwen3.8-Flash-Next on 48GB Mac with at ~12 tok/s
yes! I'm bullish on this. there is a lot of work to do. I've been experimenting with pruning, distillation, and retraining too. I'm sure your 32gb m6 will run a badass local model!
carloslfu··on Show HN: Running 104GB Qwen3.8-Flash-Next on 48GB Mac with at ~12 tok/s
I feel you! fix incomming
carloslfu··on Show HN: Running 104GB Qwen3.8-Flash-Next on 48GB Mac with at ~12 tok/s
thanks! I'll do!
carloslfu··on Show HN: Running 104GB Qwen3.8-Flash-Next on 48GB Mac with at ~12 tok/s
interesting!
carloslfu··on Show HN: Running 104GB Qwen3.8-Flash-Next on 48GB Mac with at ~12 tok/s
yes! I guess future hardware designs will have something like that!
carloslfu··on Show HN: Running 104GB Qwen3.8-Flash-Next on 48GB Mac with at ~12 tok/s
thanks!
carloslfu··on Show HN: Running 104GB Qwen3.8-Flash-Next on 48GB Mac with at ~12 tok/s
Ah! Yeah, I didn't invent anything (yet!). The goal is to see how far I can take it in terms of speed without consuming that much RAM.
carloslfu··on Show HN: Running 104GB Qwen3.8-Flash-Next on 48GB Mac with at ~12 tok/s
Good one! I haven't measured this. I'll include it!
carloslfu··on Show HN: Running 104GB Qwen3.8-Flash-Next on 48GB Mac with at ~12 tok/s
I agree with the sentiment, but have you seen those videos in which all men say other men are gay? This feels like the same, so much AI paranoia!

I genuinely want to contribute. And hey! I was doing oss this since 2014 so waay before AI was cool.

carloslfu··on Show HN: Running 104GB Qwen3.8-Flash-Next on 48GB Mac with at ~12 tok/s
I'm sorry this makes it seem like I didn't do my research. I did a TON. To fix it I'll add a benchmark/comparison table. Also, I wouldn't call it market research since this is not commercial AT ALL.
carloslfu··on Show HN: Running 104GB Qwen3.8-Flash-Next on 48GB Mac with at ~12 tok/s
Sorry, I don't get "NIH". what's that?
carloslfu··on Show HN: Running 104GB Qwen3.8-Flash-Next on 48GB Mac with at ~12 tok/s
Thanks for the feedback! I'll create a section with a benchmark and comparisons. This will hold the project accountable and speed things up imo
carloslfu··on Show HN: Running 104GB Qwen3.8-Flash-Next on 48GB Mac with at ~12 tok/s
I see your point. As an oss defender myself, I agree, however, the spirit of this is to see how fast I can make it. I'm sharing this with the community, which I think is aligned with the original oss spirit.

It's an experiment for myself but I am committing to maintain it. I've been an oss person for a loooong time, way before AI was a thing. Think about it as a new, from-scratch take at it, not as a re-reproduction.

carloslfu··on The Harness Doesn't Matter
isn't the system prompt and tools part of the harness?
carloslfu··on Show HN: We built open OpenRouter that turns usage into a better model
why text world models? Is this backed by any evidence or did you find it to be good in practice?
Page 1 of 2Next →