fal, the AI infrastructure company, released H3 Max on August 27, a post-trained version of MiniMax’s open-weights H3 video model tuned specifically for generation speed. The product renders a five-second clip in under three seconds, and fal is selling access at half price through the first week to pull developers onto the new endpoint. That combination, speed plus a launch discount, makes this a distribution play as much as a model release.

fal Research built H3 Max by taking the open MiniMax H3 weights and putting the weights through extra post-training aimed at how closely output tracks a prompt and how good it looks, per the company’s own announcement. The inference team then rebuilt the serving stack around the resulting model instead of adding speed tricks after the fact, training and running the system on Nvidia’s GB200 NVL72 hardware.

fal reports that H3 Max ranked first in its own head-to-head preference testing against twelve video models, including MiniMax’s H3 endpoint, Google’s Gemini Omni Flash, Wan 3.0, Seedance 2.5, Kling 3, and Veo 3.1, scored across overall quality, prompt fidelity, and aesthetics. Those are fal’s evaluations of a product fal is discounting this week, and the announcement does not include an independent quality audit run by fal itself. It does point to outside benchmarking from Artificial Analysis and Design Arena, which fal says also placed H3 Max at the top of their video-model rankings, with Design Arena crediting the release as dramatically faster than the base H3 model at comparable quality.

The number that matters for builders is one fal is not marketing directly: cost per usable clip. A five-second generation that used to take a minute or more of wall-clock time on frontier video models now returns in a few seconds on fal’s endpoint. That changes how a video pipeline gets built. Teams that batched generations overnight and reviewed results the next morning can instead test a prompt, watch it fail, and rewrite it within the same minute, treating video the way image-generation tools have treated stills for the past two years.

That kind of latency also reshapes the economics of production, not just the workflow. When a single generation costs a few seconds of compute instead of a minute, a team can afford to throw away nine bad takes to keep the tenth, rather than accepting whatever the model returns on the first try. Iteration becomes cheap enough to substitute for prompt engineering, which is a different way of getting to a usable clip than optimizing the prompt up front.

fal is pricing H3 Max at 50 percent off its standard rate for the first week, reachable through the Playground, the fal Agent interface, or the API for both text-to-video and image-to-video generation. The discount is a customer-acquisition move timed to the launch, not a permanent price point, and fal has not said what the endpoint costs once the promotional window closes.

The comparison worth treating skeptically is the throughput claim. fal measures its speedup against MiniMax’s own H3 endpoint, a deployment fal does not operate, so the headline number compares fal’s optimized serving stack to a third party’s default setup rather than to the fastest alternative on the market. Teams evaluating H3 Max for production work should benchmark it against whatever video endpoint they run today, and lock in pricing before the launch discount expires and the standard rate applies.

fal published the H3 Max announcement on its own blog on August 27, 2026.