H3-metal – Native MiniMax-H3 inference for Apple Silicon
github.com
github.com
I had to modify the default ComfyUI workflows to use a GGUF quant (city96's ComfyUI-GGUF custom node, UnetLoaderGGUF in place of the stock loader) [0].
I use the model labeled Q5_K_M. There is Q8_0 available as well, which is 34GB and fits fine in 64GB unified memory if you keep resolution modest.
The main issue is speed, a ~9-second 480x864 clip at 20 steps takes me a bit over an hour. So this will be cool to try for the speed up alone.
There's a lot of great information and workflows available to follow on the r/StableDiffusion subreddit.
[0] https://huggingface.co/Abiray/MiniMax-H3-GGUF/tree/main/unet
that's rough. for comparison, i tried the exact same parameters on my 5090 RTX and it took 2 minutes to generate.
i believe diffusion models are primarily compute bound so the macs aren't really the ideal hardware for this kind of stuff
Seriously, very dumb model compared to what you can run locally, but holy moly is it FAST on one GPU, seriously impressive. Can't wait for those to be scaled up a bit to fit perfectly within 96GB VRAM, then they'll be competitive.
Put Codex to work on deploying it now, hoping the speed can improve quite a lot :-) Thanks anyway
That's crazy, a RTX Pro 6000 does that in in 2-3 minutes (give or take, depending on your exact settings). LLMs don't make the difference between standalone GPU vs unified memory + CPU so obvious as diffusion models seems to do.
Which codex?
> On the 128 GB M5 Max, clean end-to-end image+audio and embedded-video+audio renders completed in 74.58 and 76.99 seconds respectively, each with about a 40.1 GB peak physical footprint and zero swaps.
Looks like it uses 40GB? So your 96GB mac setup should work fine i guess (Model itself is 33B)
2) Since it's unified memory, you won't have 96GB available.
3) I offered a solution that is usually recommended to the "gpu poor", if he's concerned with how much memory he would need.
4) I stated, that people already pointed out how he should be fine and that "gpu poor" doesn't apply to him.
5) "gpu poor" depends on what model you are trying to use. If you want to run Kimi or GLM you are still "gpu poor" even if you have an RTX Pro 6000 with 96GB of VRAM.
Anyway, good input!
> This misconstruction is very common, included in print publications spanning several centuries. It might be considered an alternative spelling, albeit still a mistaken usage.
Thanks though, I never actually knew so was helpful :)
Your comment really sounds like "many other people would do A and B if they just had time and money to do so", but he's been doing so since time and money were major constraints.
What are some adult entertainment workflows in comfyui, I need best loras, best prompts to start with
and the communities, are they on telegram or something?
I’d personally steer clear of messaging platforms for this - who knows what one might stumble into there
Personally I have no interest, but sometime browse stuff out of curiosity. But this got more of my curiosity, what kind of "stuff" are you implying they might stumble upon on the open, public internet? Sure, some NSFW, horror and otherwise weird stuff is there, especially around AI generation, but hardly something that will leave you traumatized, unless I misunderstand what you're implying?
And no, including the word "girl" in H3 does not lead to CSAM in any way, shape or form, but it's a great example how FUD quickly spreads.
One thing that I had experience with, that led me to believe this might be true: it seems one of the earlier llama models was over-tuned to resist generating CSAM. Once I tried a somewhat sensitive prompt containing the word "girl" in it. Llama only ever generated refusals for this prompt, citing I was prompting for CSAM. GPTs and Claudes of that era had no issues with the same prompt.
But I haven't seen anything reliable about H3 being particularly special in ther regard.
even a Chinese model that thinks its Claude when asked would have inherited this association
And the r2va model can also use video input for motion reference.
For Minimax H3, more than most models, you should read (and, if you are using an LLM for prompt assistance, make it sure it has access to) the official prompt guidelines, as each of the main models (fl2va that handles text-to-video and first- and/or last-frame-to-video and r2va that handles more complex reference cases) has its own structured prompt format (with many common features).
-Optimal settings/configs examples for h3.c to help speed things up -Optimal recommended generation settings for each mac product, i am sure it is easy to do -Prompt generator assistant
Great work from the github author. The more i use it the more i realise i don't need a gui to generate video, just terminal.
My setup: Macbook #1 as a client Macbook #2 as server
-use macbook #1 terminal + ssh -run mactop in terminal tab to monitor Macbook #2's hardware during ai video generation -run h3.c on Macbook #2 via ssh session -transfer the output video file from macbook #2 to macbook #1 using terminal scp -view the video
I noticed on a bar TV the other day that some of the Chromecast screensaver landscape photo credits were to Peter Norvig. They were really lovely pictures.
Does this model work with ComfyUI easily? Can I just download it?