HNHacker News
TopNewBestAskShowJobs

sawfwair

7 karma · joined May 8, 2026

submissionscomments
sawfwair··on Show HN: Local text, image, video, music and 3D from one CLI, no Python
That's a great q and point. I built it to get at a broad approach to tying multimodal capability together in one place as I'm personally building products on top of it that stitch chat, image, video together and wanted to enable others to do the same. For just coding/chat/specialized models I'd say ollama or lm studio is still the better bet generally, but for some of the models I care about I'm trying to optimize pretty deeply and make fine-tuning really simple. Plus I'm often supporting and bouncing between Mac and Cuda boxes so wanted something where the same codebase ran on both.
sawfwair··on Show HN: Local text, image, video, music and 3D from one CLI, no Python
Yes, missed that completely... Great point!
sawfwair··on Show HN: Local text, image, video, music and 3D from one CLI, no Python
totally fair! i spent a while trying to get clever and compress what i wanted to say and finally just hit submit but prob lost too much - local inference runtime, one cli that runs image/video/music/speech/3d/etc on your own machine without package hell. would def edit if I still could after your feedback, but window is closed. Thank you!
sawfwair··on Show HN: Local text, image, video, music and 3D from one CLI, no Python
some real numbers on my m4 max -

image gen (zimage-nano), 1024x1024: 58s

image -> textured mesh (trellis.2): 2m 49s

sfx generate (5s clip): 3.6s

music generate (8s, ace-step): 15s

speech synth: 13s | transcribed back: 2.2s

video gen w/ audio(ltx unified-av) 4s,768x512: 2m 48s

text chat (laguna xs2.1): 102 tok/s - https://mlx.fast leaderboard