balanced · 4.1 s
The default in this site's runs. Roughly 2.6 seconds of overhead on top of a 10-second 480p Turbo job, and four of eight clips came back with at least one hard cut the prompt never asked for.
Preview endpoint · timed here on 3–4 September 2026
Two different products share the word Turbo. H3 Max Turbo is a preview endpoint fal published on 2 September 2026: hosted, billed per output second at half the H3 Max rate, and carrying no loras field in the schema read on 3 September. The local Turbo LoRA is a downloadable four-step adapter that runs inside your own ComfyUI graph on your own card. Neither is a version of the other, and this page is about the hosted one.
Independent guide; not affiliated with MiniMax or fal. Every latency figure below was timed by this site through fal's queue API on one account across two consecutive nights, so read them as one operator's experience rather than a service-level promise. Anything you generate on fal is governed by fal's own terms.
Pick Turbo at 480p for anything that has to keep up with a viewer: it finished a 10-second clip in 4.3 seconds at the median here and costs half of H3 Max. Pick H3 Max when you want 768p from the non-preview endpoint or when you need reference-to-video, which Turbo does not expose. Reach for Director only when a continuous session matters more than a per-clip bill. And if you are chasing latency rather than looks, turning prompt expansion off did more for wall time than the whole Turbo-versus-Max gap did.
This page is for builders choosing an endpoint for batch rendering or an interactive channel. It does not repeat what a post-trained hosted model is or how to write the first API call — those live on the H3 Max reference page, along with the JavaScript quickstart. What follows is only what this site could time, count or read off a card.
The page's main promise
Wall time here means the interval between the submit request leaving this machine and the result JSON coming back; the download is not counted. Inference is whatever fal put in its own timings.inference field. RTF divides wall time by the length of the clip, so anything under 1.0 rendered faster than it will play.
| Endpoint and tier | Wall p50 / p95 | Inference p50 | fal's card sample | $/second now → after 7 Sep | Per clip now → after | RTF p50 / p95 | n |
|---|---|---|---|---|---|---|---|
| Turbo 480p · 10 s | 4.3 s / 12.9 s | 1.10 s | 1.55 s | $0.00625 → $0.025 | $0.0625 → $0.25 | 0.43 / 1.29 | 24 |
| H3 Max 480p · 10 s | 6.0 s / 14.6 s | 1.75 s | 2.53 s | $0.0125 → $0.05 | $0.125 → $0.50 | 0.60 / 1.46 | 29 |
| Turbo 768p · 10 s | 8.9 s / 18.9 s | 4.66 s | not published | $0.01 → $0.04 | $0.10 → $0.40 | 0.89 / 1.89 | 30 |
| H3 Max reference-to-video 480p · 5 s | 6.9 s / not measured | 2.11 s | 6.09 s | $0.08 → $0.08 No discount was shown on this card. | $0.40 → $0.40 | 1.4 / not measured | 3 |
| Turbo 480p · 10 s · expansion disabled | 2.5 s / 4.2 s | 1.07 s | not published | $0.00625 → $0.025 The mode does not change the rate. | $0.0625 → $0.25 | 0.25 / 0.42 | 30 |
Three things fall out of that table. First, the gap between Turbo and Max is smaller than the price gap: 1.10 seconds of inference against 1.75, but 4.3 seconds of wall time against 6.0, because a fixed overhead of roughly 2.5 to 3 seconds sits on top of both and does not scale with clip length. Second, the 95th percentile on every balanced row is longer than the clip it produced, which is the number that decides a player's buffer. Third, fal's published card samples run slower than what arrived here — 1.55 against 1.10 on Turbo, 2.53 against 1.75 on Max — so the card is not an optimistic figure to be discounted; it is a different measurement of a different job.
Buffer rule this supports: under balanced expansion, start rendering two clips ahead rather than one, and consider making the second of those two a disabled job, whose 95th percentile of 4.2 seconds leaves real headroom inside a 10-second slot. Longer clips amortise the fixed overhead better: at 480p the same endpoint gave RTF 0.85 for a 5-second clip, 0.53 at 8 seconds and 0.43 at 10.
Choose by constraint, not by suffix
Four of these columns are fal endpoints billed per output second; the fifth is the local adapter that shares the Turbo name and runs on hardware you own. Director was read from its endpoint card and never called here, so its row is marked accordingly, and the local column carries its own clip length because a 3060 run and a hosted run are not the same job.
| Property | H3 Max Turbo | H3 Max | H3 Max reference-to-video | H3 Max Director | Local Turbo LoRA · RTX 3060 |
|---|---|---|---|---|---|
| Discounted rate | $0.00625 (480p) · $0.01 (768p), to 7 Sep | $0.0125 · $0.02, to 7 Sep | none shown | $0.02 per second, to 14 Sep | not applicable |
| Standard rate | $0.025 · $0.04 | $0.05 · $0.08 | $0.08 at either tier | $0.08 per second | Your own hardware and power |
| Resolution | 480p or 768p | 480p or 768p | 480p or 768p | Adjustable per session | 1344×768 in the measured run |
| Duration | 5–15 s per job | 5–15 s per job | 5–15 s per job | 60 s billed minimum; past two minutes needs fal's approval | 124 frames at 24 fps in the measured run |
| End frame | Yes, on image-to-video | Yes, on image-to-video | not measured | A continuous stream, not discrete clips | not measured |
| Reference input | absent from the schema | Through the separate reference-to-video endpoint | Up to 9 images, 3 videos, 3 audio clips | Not offered | A separate local reference-to-video workflow |
| LoRA input | no loras field | no loras field | Not offered | Not offered | This column is the LoRA |
| Continuous context | Not offered | Not offered | Not offered | Yes, up to two minutes, prompt changeable mid-session | Not offered |
| Downloadable weights | none found | none found | none found | none found | Yes, under territory terms |
| Measured here? | Yes, across three runs | Yes, across two runs | Yes, three jobs | spec read, not measured by this site | Yes, five formal runs |
| Best fit | Batch work and interactive slots at 480p, where the bill scales with every second you render | The 768p job that is not a preview endpoint, and the gateway to reference input | Holding one subject across shots when a text prompt cannot | A live session where continuity beats per-clip accounting | Runs you can inspect, repeat and keep, with no per-second meter |
Why the local column is not a speed comparison. The RTX 3060 run rendered 124 frames at 1344×768 with the four-step adapter in 526.7, 528.9 and 528.6 seconds, against 2,175.5 and 2,176.2 seconds with the adapter off at twenty steps. Those are wall times for a different canvas, a different frame count and a different machine than fal's backend, so this page states them beside the hosted figures and never divides one by the other. The full conditions sit on the Turbo LoRA page.
Four deadlines in three weeks
fal restated the discount on these endpoints more than once in the nine days this page covers, in both directions and with different end dates on different surfaces. Each row below is one reading with the date it was taken, because none of them cleanly supersedes the others.
| Date | What changed | Where it was read |
|---|---|---|
| 27 August 2026 | H3 Max launches with half off. The launch post called it the first week; the landing FAQ said fourteen days. | fal launch blog and landing FAQ, read 31 August |
| 31 August 2026 | The endpoint card showed $0.025 and $0.04 per second with the discount ending 1 September. | H3 Max text-to-video card, read 31 August |
| 3 September 2026 | The same card now showed 75% off ending 7 September: Max at $0.0125 and $0.02, Turbo at $0.00625 and $0.01. The deadline moved out, and the discount deepened. | Max and Turbo cards, read 3 September |
| 7 September 2026 | Standard rates resume: 480p Turbo returns to $0.025 and 768p Max to $0.08. Every batch costs four times what it did the previous day. | Deadline printed on both cards, read 3 September |
| 13 September 2026 | Vercel's AI Gateway half-price window on H3 and H3 Max closes. It ran from 30 August and never matched fal's own dates. | Vercel changelog, read 3 September |
| 14 September 2026 | Director's separate discount ends: $0.02 per second becomes $0.08, on a 60-second billing floor. | Director endpoint card, read 3–4 September |
Worth knowing before you multiply anything: the 165-clip batch behind this page — 1,520 output seconds across Turbo and Max — was estimated at $13.34 from the card rates and charged exactly $13.34 in the dashboard. Three consequences follow. Submissions the API rejected with a 403 were not charged. Queue time, the prompt rewriter and the download were not charged. And nothing but output seconds appeared on the bill, so a per-second estimate needs no hidden-fee cushion. The H3 Max page keeps the evergreen standard-rate table; this page keeps the moving parts.
The parameter that owns your latency
Before either endpoint renders anything, fal rewrites your prompt. A 111-character median input came back as a 2,153-character median storyboard with shot markers, timecodes, a soundscape paragraph and a music cue — a nineteen-fold expansion, in the same three-part format on all 165 records inspected. That rewrite is where most of a Turbo clip's wall time goes.
balanced · 4.1 sThe default in this site's runs. Roughly 2.6 seconds of overhead on top of a 10-second 480p Turbo job, and four of eight clips came back with at least one hard cut the prompt never asked for.
fast · 4.2 sNo measurable difference from balanced at this sample size: same overhead, same rewrite format, five of eight clips with a hard cut. If you were hoping this was the latency switch, it is not.
disabled · 2.4–2.5 sOverhead collapses to 0.9 seconds and expanded_prompt comes back null. Seventy clips, zero hard cuts. This is the latency switch.
quality · 17–41 sBetween 16.8 and 41.2 seconds of wall time for the same 10-second clip, almost all of it rewriting, with no more shots planned than balanced produced. Unusable for anything interactive.
fal's schema text names only balanced and quality. Submitting a probe value of none returned a 422 whose body enumerated the real set: disabled, fast, balanced, quality. The two undocumented values are the two that matter — disabled is the fastest and fast is indistinguishable from the default. Checked 4 September 2026; treat the enumeration as current until fal's schema catches up.
What turning it off costs you. The saving is not free, and the price is not money — fal bills output seconds, so a disabled job and a balanced job of the same length cost exactly the same. What changes is the picture. Without the rewriter the clips arrive as a single continuous take with the look of a CG toy render, where balanced clips read as cinematic live action, because the rewriter put "Cinematic" into all 165 expanded prompts inspected and "live-action" into 113 of them. The automatic soundscape and music paragraphs disappear too. And you cannot get the best of both by pre-expanding: feeding an already-expanded storyboard back in under balanced does not skip the rewrite, it rewrites the rewrite, growing a 1,940-character input to between 2,356 and 2,525. Only disabled actually skips it, which makes "write your own shot list, then send it with expansion off" the low-latency way to keep a rich prompt.
The rewriter is also your concurrency ceiling. An earlier run had concluded that one fal key gets about three useful lanes, because pushing concurrency to 10 sent queue time to a 72-second 95th percentile. Two identical batches twenty seconds apart settled where that serialisation lives: at concurrency 10 with balanced, overhead ran 8.4 seconds at the median and 15.5 at the 95th percentile; with disabled, the same ten parallel jobs held 0.9 seconds flat. Inference stayed at 1.07 to 1.09 seconds on both sides. It is the rewriter that serialises, not the video backend — with it off, thirty 10-second clips finished as a batch in 23.6 seconds with no queueing at all. Plan lanes accordingly: about three under balanced, at least ten under disabled, and nothing above ten has been tested here.
Seeds do not bring a clip back. Re-submitting byte-identical inputs at the same seed produced structural similarity of 0.49 to 0.56 and peak signal-to-noise around 15.5 to 16.7 dB against the originals — different footage, not a variation. With expansion disabled, where the rewriter's randomness is gone, two same-seed passes still landed at 0.48 to 0.79. So the seed is a draw, not an address. Cache the mp4 file itself; a prerender cache, a branch replay or an A/B comparison keyed on seed will not find its clip again.
Short version of a longer problem
If you are stringing clips into a channel or an episode, this decides your architecture more than latency does. The counts below come from the same batches timed above, classified by eye from three frames per clip.
Across 69 same-seed clips of one described character, 19 kept the shape the prompt asked for. Splitting by whether the prompt itself carried an identity cue explains all of it: clips whose text mentioned the single glowing eye held 14 of 15, clips without any such cue held 0 of 45. The drift pattern was the same on Turbo, on Max and at 768p, so it is not something Turbo's distillation introduced. Turning the rewriter off does not rescue it either — 30 disabled clips of the same story held 5, no better than balanced, and the cued scenes actually did worse. The model's prior for "a brass robot" is a two-eyed humanoid, and nothing in the prompt pipeline defends against that except the prompt.
Costs nothing and did the most work: 14 of 15 versus 0 of 45. Write the character out in full in every prompt of the sequence rather than referring back to it.
Feeding each clip's final frame in as the next clip's start image held the subject for four links on Turbo and two on Max. It breaks at a scene change or when a second character enters.
One reference image held all three scenes that text alone had lost, 3 of 3. It is H3 Max only, $0.08 per second with no discount, and a 5-second clip took 6.9 seconds.
The full ladder — including how to write shot lists that survive disabled mode, and how to size a buffer around each option — belongs to this site's interactive-channel tutorial, which is being written and is not published yet. This section deliberately stops at the choice between the three approaches above.
Same prompt, same seed, three tiers
Each strip below is one scene rendered on all three timed tiers, sampled at 1.5, 5 and 8.5 seconds into the clip. Read them as evidence about identity drift and about what each tier produced, not as a quality ranking: this page scores nothing and declares no winner. Note also that the same seed number does not mean the same picture across tiers — it is a fresh draw each time, so the three rows are three independent samples of one idea.
No video is embedded on this page and none is hosted here. The clip packs behind these strips are published as a GitHub release, and the strips themselves were cut from those files with ffmpeg at fixed timestamps — no colour correction, no crop, no re-ordering.
Reproduce it on your own key
Every figure above came out of one script that lives in this repository. It is a single file with no dependencies, it never sees a server, and your key stays in your shell.
A dry run prints the job list and the estimate without a key, so you can see the bill before it exists. The estimate carries the 3 September card rates and switches to standard rates after 7 September.
Put FAL_KEY in your shell environment only. The runner submits through the queue API, polls, and downloads the finished clips into a run directory.
Each job records its request id, seed, wall time, queue time, fal's reported inference and the estimated cost; a summary file holds the p50 and p95 that this page's table is built from.
Where it is: 03-dev/tools/fal-batch in this site's repository, with the prompt lists used for the runs above. The clips it produced are attached to the clips-2026-09 release. The estimate is arithmetic and the dashboard is the invoice; on the batch behind this page the two agreed exactly.
Where the sources disagree
Each row is a place where fal's own surfaces, or fal's surfaces and this site's measurements, say different things. They are printed as open discrepancies rather than folded into a single confident answer.
| Discrepancy | The readings | What this page does with it |
|---|---|---|
| Which values prompt expansion accepts | The schema description on the endpoint page names two modes. The API's 422 response, provoked on 4 September 2026 by submitting none, enumerated four. |
The four-value set is used throughout, on the ground that a rejection message from the running service outranks prose on a card. If fal narrows the set later, the disabled-mode figures on this page stop applying. |
| How long the discount lasts | Launch post: the first week. Landing FAQ: fourteen days. Endpoint card on 31 August: ends 1 September. Endpoint card on 3 September: 75% off, ends 7 September. Vercel's gateway: half price to 13 September. Director's card: to 14 September. | The endpoint card is treated as operative because it is what the billing system quotes, and the reading date is printed beside every price. Nothing here should be trusted after 7 September without re-reading the live card. |
| fal's sample timing against this site's | The Turbo card's sample response reported 1.55 seconds of inference; 24 timed jobs here reported 1.10. The Max card reported 2.53; 29 jobs here reported 1.75. The reference-to-video card reported 6.09 against 2.11 measured on a shorter clip. | Both are printed in the table without a ratio between them. The card sample is a single fal-side example of an unstated job; this site's column is one account's median on a stated job. They are not the same measurement and are not subtracted from one another. |
Conditions attached to every timing on this page. One fal account, two consecutive nights (3 September around 21:00 and 4 September around 01:00, US Pacific), one region, the queue API at queue.fal.run with a 1.5-second poll interval, concurrency 3 except where stated. That poll interval puts up to 1.5 seconds of quantisation error into the queue and overhead columns. Group sizes run from 3 to 30, no statistical significance is claimed, no picture-quality judgement was made, and downloads are excluded from wall time. During one stage the account balance ran out and fal returned intermittent 403s; those submissions were not billed and are excluded from every figure. Independent guide; not affiliated with MiniMax or fal.
FAQ
No. H3 Max Turbo is a hosted preview endpoint fal published on September 2, 2026 at half the H3 Max per-second rate, and its schema exposed no LoRA input when read on September 3. The Turbo LoRA is a downloadable four-step adapter you load into a ComfyUI graph on your own GPU. One is a billing relationship with fal, the other is a file on your disk.
Across 24 timed 10-second 480p jobs on the night of September 3, 2026, submit-to-result took 4.3 seconds at the median and 12.9 seconds at the 95th percentile, of which fal reported only 1.10 seconds as inference. The tail is longer than the clip, so a zero-wait player needs two clips of buffer, not one.
The endpoint card read on September 3, 2026 showed a 75% discount ending September 7. When it lapses, 480p Turbo goes from $0.00625 to $0.025 per output second and 768p from $0.01 to $0.04, so every batch quadruples in cost on the same day. Read the live card before committing a large run.
Yes. The parameter accepts disabled, fast, balanced and quality, although fal's schema text names only the last two; the API's own 422 error listed all four. Disabled dropped a 10-second 480p job to 2.5 seconds median and 4.2 at the 95th percentile, and it costs nothing extra, because fal bills output seconds only.
It did not here. Re-running identical inputs at the same seed produced structural similarity of 0.49 to 0.56 against the originals, and repeating with expansion disabled still gave 0.48 to 0.79. Treat the seed as a fresh draw and cache the finished mp4 file, because you cannot regenerate it later from its parameters.
Reference-to-video, which exists on H3 Max and not on Turbo. One reference image held the subject in all three scenes that text-only prompts had lost, at $0.08 per second with no discount and a 6.9-second wall time for a 5-second clip. Without an image, repeat the character description in every prompt: clips carrying that cue held 14 of 15, clips without it 0 of 45.
Where to go next
Turbo at 480p is the default for anything that renders per viewer or per second. Move to H3 Max when 768p or reference input decides the job, and move off fal entirely when you need weights you can keep.
Prices on this page were read on 3 September 2026 and one of them expires on the 7th. Open the endpoint card, confirm the number, then size your batch.
Open the H3 Max Turbo card on fal