MiniMax Open-Sources H3, an Omni-Modal Video Model That Runs on a 3060
Original source
MiniMax H3 Day-0 Support in ComfyUI: Open Weights, Native Audio, and 2K Video
Hacker News →MiniMax has released H3, the third generation of its Hailuo video line and the first the company has shipped with open weights. ComfyUI added native support the same day. The model takes text, images, video, and audio as inputs and produces video up to 2K resolution and 15 seconds per clip, with stereo audio generated in the same pass rather than added afterward. Optimization work on the ComfyUI side means it can reportedly run locally on hardware as modest as an RTX 3060.
The pitch is consolidation: instead of separate pipelines for text-to-video, image-to-video, first-and-last-frame control, and reference-driven generation, H3 handles all of them through one model that reasons across modalities. Users describe how their inputs relate and what shot they want, and the model resolves the cross-modal relationships itself. For node-graph workflows, the notable features are motion transfer — pulling a camera move, performance, or editing rhythm from a reference clip while sourcing subject and style elsewhere — and in-place editing for iterating on a shot.
The significance is less any single capability than the packaging: open weights, native audio, and consumer-GPU viability arriving together on day zero in a widely used local tooling ecosystem. That lowers the barrier to running a competitive multimodal video generator entirely on local hardware, outside the hosted APIs that have dominated this space.
Read the full article
Continue reading at Hacker News →This is an AI-generated summary. Read the original for the full story.