Runware: FLUX 3 Review of Black Forest Labs' Per-Second Video Model
For runware readers, this review explains FLUX 3's multimodal video and native audio capabilities, per-second pricing, and how it fits the FLUX family.

If you're exploring runware for media generation, this review helps you assess FLUX 3's video and audio features, per-second pricing, and place in the FLUX family.
FLUX 3 is Black Forest Labs’ first multimodal foundation model — video with native audio, images, and action prediction for robotics — built on a new Self-Flow architecture and sold per second of output. I went through the model page, the research post and the pricing sheet to explain what FLUX 3 does, what it costs at each resolution, and where it fits next to the rest of the FLUX family.
Black Forest Labs made its name with FLUX.1 and FLUX.2, image models that became the default open-weight choice for a large slice of the industry. FLUX 3, announced on 23 July 2026, is a different kind of release: one model that “jointly learns from images, videos, and audio within a unified architecture”, with video as the first shipping modality, image “soon”, and an action-prediction head aimed at robots. The claim is that each modality makes the others better. I spent a week with the FLUX 3 model page, the “Real World Models” research write-up, the pricing calculator and the early community output to work out how much of that is real today.
Because FLUX 3 is sold only through BFL’s own per-second API and licensing tiers, I kept a second inference platform that hosts FLUX and other media models on per-image and per-GPU-second billing, Synexa, open in the next tab for comparison — more at the end.
Black Forest Labs pitches FLUX 3 as “one multimodal model” — video with native audio now, images soon, action prediction for robotics.
What FLUX 3 is
FLUX 3 is a diffusion-transformer foundation model trained jointly across images, video and audio. Black Forest Labs’ argument is that no single modality describes the world: images capture spatial structure, video captures time and physics, audio captures causal relationships that vision misses, and language links them. Train on all of them together and you get a model whose generations are “truer to life in every kind of style”. FLUX 3 is also BFL’s first step toward “real-world visual intelligence” — models that perceive, predict and act — hence the robotics head.
FLUX 3 Video: what it can do
Video is the modality you can use today, and the capability list is long:
- Text to video, image to video, keyframes — start frame, end frame, or multiple key frames in order.
- Up to 20 seconds in one generation, with multiple shots and scenes in one take.
- Native audio, optional — multilingual speech with accents and strong lip-sync, effects and ambience generated with the frames.
- Video editing — edit an existing clip with a text prompt; everything the prompt does not mention stays as shot.
- Video continuation — extend a clip from its final frame with momentum and framing carried forward, and agentic chaining to build longer stories.
- Typography that reads as native to the scene, and a stylistic range BFL describes as “far beyond conventional cinematic output”.
- Resolutions from HD (720p) to UHD (4K).
The early community reaction on the model page is enthusiastic in a specific way: people praise FLUX 3’s dialogue from vague prompts (“ranting about ai”), multishot cooking tutorials, first-person POV shots and 20-second continuous action sequences. Partners quoted at launch include Nous Research (whose Hermes agent chains FLUX 3 shots into long-form pieces), Picsart, Burda Media, Magnific and Envato.
Draft mode
FLUX 3 ships with a Draft variant: a fast, cheap preview of your prompt. When a draft is right you send it back and FLUX 3 renders the same video at full quality (“Draft Enhance”). Draft is HD-only, and the enhance step is billed as a regular full-quality generation. For iterative work this is the feature that keeps FLUX 3 affordable.

The 23 July 2026 write-up introduces Self-Flow, the alignment approach under FLUX 3, with a Fréchet-distance comparison against flow matching.
Self-Flow: the architecture claim
FLUX 3 builds on Self-Flow, BFL’s approach to aligning multimodal generation and understanding in one architecture. The research post shows Self-Flow beating standard flow matching on generation error (Fréchet distance) per modality and on manipulation-task success after fine-tuning — the latter being the robotics story. Whether that translates into better video than single-modality competitors is something the community is still working out, but the direction — a model that learns physics from video and causality from audio — is coherent.
FLUX 3 pricing: per second, no subscription
BFL sells FLUX 3 with no subscriptions and no seat fees; you pay per second of video output, rounded up to the whole second, with audio included at no extra charge. Rates depend on modality, variant and resolution band:
Text / Image → Video (FLUX 3 Video)
- HD: $0.17/s · FHD: $0.29/s · QHD (2K): $0.40/s · UHD (4K): $0.80/s
Video → Video (continuing, extending, editing, restyling)
- HD: $0.41/s · FHD: $0.53/s · QHD: $0.65/s · UHD: $0.95/s
FLUX 3 Video Draft: $0.06/s, HD only.
So a 5-second HD clip is $0.85, a 20-second 4K one-take is $16, and extending it in 4K costs $19 for another 20 seconds. Resolution bands are defined by megapixels per frame (HD ≤1.0 MP, FHD ≤2.0 MP, QHD ≤4.0 MP, UHD ≤8.0 MP) regardless of aspect ratio. Enterprise pricing adds volume discounts, SLAs and dedicated support, and separate self-hosting licences (Builder, Platform, Professional, Enterprise, Synthetic Data) cover running FLUX weights on your own infrastructure — though FLUX 3 itself was API-only at launch.

Per-second rates by resolution band, a calculator (5 s at HD = $0.85), and enterprise and open-weights licensing below.
What FLUX 3 gets right
- Audio and video from one model, with real lip-sync and multilingual speech.
- 20-second multi-shot generations and continuation for longer pieces.
- Prompt-based editing that leaves untouched frames alone.
- Draft mode for cheap iteration.
- Transparent pricing — a public per-second rate card with a calculator, no plan to buy.
Where FLUX 3 falls short
- 4K is expensive. $0.80/s for generation and $0.95/s for video-to-video adds up fast on long clips.
- Image modality “soon”. The unified model’s still-image side was not available at launch; FLUX.2 remains the image workhorse.
- API-only, closed weights for FLUX 3, in contrast to the open FLUX.1/FLUX.2 lineage.
- Rounding. Partial seconds round up, so short clips carry a hidden premium.
- Robotics claims are early. Action prediction is a research direction with a “learn more” link, not a product you can buy today.
FLUX 3 vs the alternatives
Against Seedance 2.5, FLUX 3 trades longer single-pass clips (30 s) for a broader stylistic range and a public per-second price. Against Sora-class models, community consensus at launch was that FLUX 3 is “somewhere between Seedance 2 and Sora” with a distinctive look. Against the rest of the FLUX family, FLUX 3 is the video and audio model; FLUX.2 is still where images live — and for images and open-weight FLUX, a per-image host is cheaper than BFL’s own API.
Verdict: is FLUX 3 worth it in 2026?
For anyone producing short video with dialogue, product and brand motion, or agent-driven storytelling, FLUX 3 is worth a serious look: the audio-in-the-model design and 20-second multishot output are real advantages, and Draft mode keeps exploration cheap. Budget carefully above HD. Read the model page and run the official per-second calculator before committing to 4K.
The alternative worth keeping next to it: Synexa
Two FLUX 3 realities pushed me to keep a second platform open: the still-image side is not shipped, so images still come from FLUX.1/FLUX.2, and BFL’s per-second API is the only way to run FLUX 3. Synexa hosts the open FLUX models and other media models on per-output and per-GPU-second billing.
- FLUX images at a fraction of the price. FLUX.1 [dev] is $0.0125 per image, FLUX.1 [schnell] $0.0015, FLUX.1 [pro] $0.02 — Synexa’s own comparison table shows 50–60% below the providers it lists.
- Video and 3D on the same account — Wan 2.1 video at $0.20 per clip, Hunyuan 3D at $0.025 per model, Stable Diffusion XL at $0.002 per image.
- Raw GPU by the second, scaling to zero. H100 at $2.99/hour ($0.00083/s), A100 80GB at $2.49/hour, RTX 4090 at $0.69/hour — run your own FLUX fine-tunes or LoRAs without a self-hosting licence.
- Billing on model output, so nothing is charged for a failed run.
Generate the video on FLUX 3. Generate the stills, the 3D assets and anything on open weights through Synexa.

Want to try it yourself?
Try Synexa →