Tap8 Benchmark · July 2026
Contents
Back to Research
Benchmark overview · 7 systems · 5 use cases · 25 intent-driven scenes · 2026

On information-dense video,
Tap8 leads the field.

A benchmark for video that must inform — instructional, editorial, explanatory — not entertain. On the information-dense scenes it targets, Tap8 takes the outright Factual Coverage lead, ties for first on TeachQuiz Information Delivery, and tops Overall Quality across seven systems. The independent, industry-standard automated metric, VBench, reproduces the human panel’s preference, confirming the eval’s rigor.

7 systems — code-render vs pixel-diffusion 5 use cases × 5 long videos each 6,760 human annotations >3.7M VBench frame evaluations fully-crossed human annotator panel + VBench automated metrics: VBench · TeachQuiz · T2VTextBench · T2VWorldBench
The verdict

First on the content that matters — balanced everywhere else

Information-dense scenes — professional, news, how-to, education — are the benchmark’s target domain, and they are where Tap8 wins. Across all 25 scenes, Tap8 is the clear 2nd of 7 on Informative Quality and the only code renderer strong on both informative axes.

Overall Quality — informative + aesthetic
75.2 T
#1 of 7 on information-dense work — the only system high on every axis at once.
Aesthetic Preference vs Fable 5, the informative leader
58.7%
Preferred over the field’s top informative system in side-by-side viewing — and significantly does-not-lose.
Factual Coverage — information-dense scenes
83.6%
The outright field lead. Significantly ahead of 5 of 6 competitors at getting required facts on screen.
TeachQuiz Information Delivery — information-dense scenes
84.5%
Tied for first. Viewers answer comprehension questions as well after Tap8 as after any system in the field.
Overall Quality — informative + aesthetic · information-dense scenes
(Content Fidelity + Visual Quality + Aesthetic Preference) / 3 · field-relative T-score (70 = field mean, 10 = 1 SD)

Both frames are reported everywhere. Headline claims use the information-dense frame (UC1–UC4) the benchmark targets — shown here, where Tap8 tops Overall Quality; on all 25 scenes Tap8 is 2nd of 7 on Informative Quality (T 78 vs Fable 5’s 85). Full scorecards with confidence intervals are in the appendix.

The benchmark

Built for videos that must inform, not entertain

Each of five use cases draws on real-world source material and is evaluated on five intent-driven scenes: the brief specifies the facts a viewer should walk away with, and human annotators measure whether they do. The target is high-information-density, clear-intent video — instructional, editorial, explanatory — not aesthetic or entertainment content.

The field — seven systems, two paradigms
SystemHarness / modelFamily
Tap8code-rendered videocode-render
Fable 5Claude Code · Fable 5code-render
Code-render HClaude Code · Sonnet 5code-render
Code-render RCodex · GPT-5.5code-render
Pixel-diffusion Lofficial studio · Pro mode, 1080ppixel-diffusion
SeedanceSeedance 2.0 · Jimeng · VIP mode, 720ppixel-diffusion
VeoVeo 3.1 · Google Flow · Quality mode, 720ppixel-diffusion

All systems accessed 2026 Jul 6–12: locally-run tools frozen at their latest release as of Jul 6, hosted systems at the latest flagship model on their official platform — every system at its maximum quality setting, 16:9. Exact version pins with sources are in appendix A.

Five use cases, real-world source material
Use caseSource material
Professional
Enterprise / self-learning
GDPval tasks — Professional & Technical Services, Government, Manufacturing, Finance
News
Editorial
2025–26 covers: Economist, FT, Bloomberg, Nature, Science
How-To
Lifestyle
Top #HowTo YouTube — craft/DIY, art, fitness, mindset, puzzles
Explainer
Education
Most-viewed “AP Exam explained” — history, psych, calculus, biology, geography
Marketing
Product / sales
NASDAQ most-valuable-company launch videos across sectors

Sampling rules fixed before generation. Each use case draws its five scenes by a predefined rule — cluster the source pool, randomly sample one per cluster (top industries, publications, content categories, AP subjects, NASDAQ sectors). Every brief is a fully-detailed video-generation prompt carrying the complete information payload, authored and frozen before any system generated anything; its required-fact checklist (168 facts across the benchmark) and comprehension questions (149) were authored at the same time and human-reviewed. A complete worked example is in appendix B.

How the field was run — one shot, same brief, top settings
  • Identical inputs: every system received the same brief verbatim — zero per-system prompt engineering.
  • Best-of-1: one generation per scene, never curated; hard failures retried until success (at most 3 retries observed).
  • Durations as generated: target lengths set by the brief; each system’s actual output used as-is (system averages 1.1–4.2 min).
  • The 173/175 accounting: Fable 5’s safety rules declined 2 of its 25 briefs. Rather than substitute its Opus fallback, we scored Fable 5 only on its own 23 generations — no competitor was quietly weakened, and missing cells are dropped from denominators, never imputed.
The annotator panel — independent, blind, fully crossed
  • 12 independent annotators, outsourced freelancers unaffiliated with any system’s team — each holding a master’s degree or above from a QS World Ranking top-50 university, with certified English competence.
  • Blind throughout: videos presented unlabeled in randomized order; annotators were never told which system produced a video.
  • Fully crossed: all 12 annotators rated all 7 systems, two raters per video — a strict or lenient rater moves every system equally and cannot bias a comparison.
  • Deliberate viewing, not click-work: ~20 minutes of annotation per video of under 5 minutes, paid at 4× local minimum wage. Annotated Jul 13–15, 2026.

Metrics grounded in published methods, measured by people

Every video is scored by a fully-crossed annotator panel — every annotator rated every system, so a strict or lenient rater moves all systems equally and cannot bias a comparison.
The panel produced 6,760 human annotations across five stations grounded in published methods — three of them peer-reviewed at ICML, WACV, and IEEE TPAMI / CVPR — and every video is also scored by the VBench automated model: >3.7 million frame-level model evaluations.

VBench is an independent, machine-based cross-check of our human annotations: its Frame Quality evaluation reproduces our human preference ranking, so the two instruments corroborate each other (Finding 05).

The rankings are robust to how you compute them: three independent estimators — raw rates, leave-one-out information gain, and mixed-effects logistic models — produce the identical ordering, and reliability-weighted and PCA variants leave it unchanged. Because Marketing behaves as a different task family from the four information-dense use cases, every result is reported in two frames: all 25 scenes, and the information-dense frame (Professional, News, How-To, Explainer). Full tables for both are in the appendix.

Finding 01 · The full-field frame

Informative video is a trade-off frontier — and Tap8 holds the balance point

WIN
The field is a trade-off frontier: on Informative Quality the order is Fable 5 > Tap8 > everyone else; on Aesthetic Preference the order inverts exactly.
Tap8 is the only code renderer high on both informative axes — Content Fidelity 77, Visual Quality 79, 2nd of 7 on each — the balance point of the frontier.

Aesthetic Preference is measured arena-style — two videos of the same brief, side by side — and captures something the informative rubrics do not: its correlation with every informative metric is negative-to-zero (−0.22 to −0.01). Code renderers buy informative quality; pixel-diffusion buys aesthetic appeal. The frontier below places all seven systems on that trade-off, in both frames.

The informative–aesthetic frontier
x: Informative Quality · y: Aesthetic Preference · field-relative T-scores (70 = field mean) · information-dense scenes

Tap8 is the only system in the green quadrant — highly informative and aesthetically preferred — in both frames. The frontier runs downhill: Informative Quality orders Fable 5 > Tap8 > the rest, and Aesthetic Preference inverts it end-for-end. On information-dense scenes Tap8 moves to the frontier’s elbow (Informative Quality 82, within 3.5 T of the leader) while the pixel-diffusion aesthetic edge shrinks to parity; on all 25 scenes Seedance and Pixel-diffusion L sit above the preference range shown. Aesthetic Preference is reported beside Informative Quality, never averaged into it.

Aesthetic win-rate of Tap8 — side-by-side viewing, all 25 scenes
share of side-by-side comparisons Tap8 wins (ties count half) · ×100

Above 50 = viewers prefer Tap8. Significance from the mixed-logit non-loss / decisive-win tests; full CIs and the non-loss frame are in the appendix.

Aesthetic win-rate of Tap8 — side-by-side viewing, information-dense scenes
share of side-by-side comparisons Tap8 wins (ties count half) · ×100

On the benchmark’s target domain the aesthetic deficits shrink to parity — only Seedance remains significantly preferred on decisive comparisons.

WIN
Shown side-by-side against Fable 5 — the field’s informative leader — viewers prefer Tap8.
Win-rate 58.7%, and Tap8 significantly does-not-lose (P = 0.72).
The system that beats Tap8 on quality composites loses to it in front of viewers.
WIN
On information-dense work, Tap8 concedes nothing on aesthetics while leading content.
Overall non-loss rises to 53.0% — parity — and no pixel-diffusion system holds a significant non-loss advantage there.
Against Google’s Veo: ahead on Informative Quality (78 vs 65), significantly ahead on coverage, and a preference toss-up.
LIMITATION
Across all 25 scenes, aesthetic preference is a real weakness. Tap8’s overall win-rate is 40.7% — significantly below the 50% no-preference line — and it is significantly less preferred than Pixel-diffusion L and Seedance, the bottom of the informative ranking. The inversion is the paradigm trade-off, concentrated in Marketing; the close-out plan addresses it directly.
Finding 02 · Factual Coverage

Getting facts on screen is the binding constraint — and Tap8 does it best

WIN
On information-dense scenes Tap8 puts 83.6% of required facts on screen — the outright field lead, significantly ahead of 5 of 6 competitors.
Across all 25 scenes it is statistically tied with the leader Fable 5 and significantly ahead of every pixel-diffusion system.

Coverage is the benchmark’s most discriminating metric: the field splits into two clean tiers, with the four code renderers (74–79% of facts on screen) sitting well clear of the three pixel-diffusion systems (54–61%).

Factual Coverage — share of required facts on screen
mean · 95% bootstrap CI over scenes · information-dense scenes · ×100

Info-dense frame: Tap8 83.6 [77.3–89.0] leads outright; the gap to every system except Fable 5 is statistically significant (mixed-logit 95% CI excludes 0). All-25 frame: Fable 5 79.0, Tap8 77.3 — a statistical tie — with Seedance, Veo and Pixel-diffusion L significantly behind in both frames.

Coverage drives Delivery — one point per video, 173 videos
x: Factual Coverage · y: TeachQuiz Information Delivery (correct-rate) · ×100
WIN
Once a fact is on screen, viewers acquire it ~95% of the time.
Coverage and delivery correlate at r = 0.70 across 173 videos, and the fitted line runs from a 36% no-fact floor to a 96% on-screen ceiling.
Comprehension is nearly a solved step — the game is coverage, and Tap8 plays it best.
Coverage by use case — model-predicted P(fact on screen)
mixed-logit predicted probability per use case · ×100
WIN
Tap8 leads or ties for the coverage lead in four of five use cases.
It leads News (81.9) and How-To (86.7) outright, and on Professional it is significantly ahead of four competitors — including the overall leader Fable 5.
LIMITATION
Marketing is the exception. Tap8’s coverage drops to 53.8% there — bottom of the code renderers, behind Fable 5, Code-render R and Code-render H. The same gap shapes every metric’s Marketing column; the close-out plan treats it as one problem.
Finding 03 · TeachQuiz Information Delivery

Viewers learn as much from Tap8 as from any system in the field

WIN
On information-dense scenes Tap8 ties for first on TeachQuiz Information Delivery — viewers answer 84.5% of comprehension questions correctly — and significantly out-delivers three systems (Code-render H, Seedance, Pixel-diffusion L).
Across all 25 scenes the top five are statistically inseparable; Tap8 is significantly ahead of Pixel-diffusion L and Seedance.

Delivery follows the TeachQuiz protocol: after watching, annotators answer multiple-choice questions about the facts the brief required, and each video is scored against a per-question field baseline. Three independent estimators — raw correct-rate, leave-one-out information gain, and a mixed-effects logit — agree on the ordering.

TeachQuiz Information Delivery — comprehension correct-rate
raw correct-rate · significance vs Tap8 from mixed-effects logit · information-dense scenes · ×100

Tap8 is top-tier in both frames. Full per-use-case win/tie/loss tables are in the appendix.

Delivery by use case — comprehension correct-rate
raw correct-rate per use case · bars with 95% bootstrap CI over the use case’s scenes · ×100

Tap8 is top-tier on News (88.3) and How-To (89.7) — per the mixed-logit inference it significantly out-delivers Code-render R, Pixel-diffusion L and Seedance on News and Veo on How-To, losing to none in either. Professional and Explainer are mid-field; Marketing is the gap the close-out plan targets. Each use case rests on 5 scenes, so the CIs are wide; significance calls come from the per-use-case mixed logit in the appendix.

T2VTextBench Text Accuracy — on-screen titles, numbers and labels
four-tier grade mapped to [0,1] · 95% bootstrap CI · information-dense scenes · ×100
WIN
Tap8 renders on-screen text best on information-dense scenes.
Across all 25 scenes the top of the field is statistically inseparable.
Finding 04 · Visual quality

Statistically tied with the leader on three of four visual dimensions

WIN
On the T2VWorldBench human evaluation, Tap8 is statistically tied with the leader Fable 5 on Realism, Relevance and Consistency — in both frames — and a clear 2nd of 7 on the four-dimension composite (86.4), significantly ahead of Seedance and Pixel-diffusion L.
Whatever Tap8 renders, it renders cleanly.

The four dimensions are rated 1–5 by the crossed annotator panel against per-dimension rubrics. The one significant visual gap to Fable 5 is Video Quality (86.0 vs 90.0); the other three dimensions are statistical ties.

T2VWorldBench — four dimensions, all 25 scenes
mean grade · 95% bootstrap CI over scenes · ×100 · * = significant vs Tap8

Tap8 is 2nd on Quality, Realism and Consistency, 3rd on Relevance (behind Fable 5 and Code-render R). Relevance — “does it say the right things” — behaves as a content dimension, and is analyzed with the content metrics above.

WIN
Where raters fully agree, Tap8 rates highest.
On the exact-agreement subset Tap8’s Video Quality rises to 0.97, overtaking Fable 5 — Tap8’s clear wins are unambiguous.
Finding 05 · Independent corroboration

Human annotations validated by VBench industry-standard automated metric

CHECK
Run blind on the same 25 scenes, VBench — the automated model behind the public HuggingFace leaderboard — scores every video independently of the human panel.
It orders the systems the same way our annotators’ side-by-side preference does.
Spearman ρ = +0.89 between VBench Frame Quality and human preference; +0.87 leaving any one system out.
The human preference axis is not an artifact of our raters — an independent, paper-grounded machine metric lands in the same order.

A widely-used automated model, computed with no knowledge of the human labels, agrees with the human panel on how the systems rank — evidence that the eval captures a real, reproducible signal rather than rater idiosyncrasy.

Automated and human instruments agree — every video shown
x: VBench Frame Quality per video (automated) · y: each system’s human win-rate · 141 system×scene videos, coloured by system · ×100

Each small dot is one video’s automated Frame Quality, coloured by system and placed at that system’s human win-rate; the large dots are the system means. The coloured clouds stack in win-rate order and their centres climb left-to-right — automated frame quality and human preference rank the systems together (Spearman ρ = +0.89). Systems are left unnamed: the point is the agreement of the two instruments, not any one system’s position.

SRC
VBench (VBench++, IEEE TPAMI 2025 · original VBench, CVPR 2024 Highlight) is the standard automated video-quality benchmark behind the public HuggingFace leaderboard, scoring six dimensions per clip. It was run with the official VBench-Long code path (commit e946a8d) on dedicated GPUs on Jul 13–14 — completed before any comparison against the human labels existed. The Frame Quality head (Aesthetic + Imaging Quality) is the aesthetic cross-check used here, computed quality-when-it-works so a failed generation does not distort a system’s score. Every composite was recomputed from the raw per-scene scores and matches the source report to the decimal; the full run protocol is in appendix A.
The next wins to close

Two targets: Marketing content, and aesthetic appeal without an informative trade

Both gaps are measured, localized, and separable from what already works. Neither calls for touching the coverage-and-delivery engine that wins the information-dense domain.

Tap8 by use case — the Marketing gap is content, not rendering
Content Fidelity vs Visual Quality composite · field-relative T-score · per use case

On Marketing, Tap8’s Visual Quality holds at 75 while Content Fidelity drops to 51 — the videos render cleanly but miss required facts. Everywhere else the two axes move together at 68–93.

The aesthetic program — move up the aesthetic axis, hold the informative lead
x: Informative Quality · y: Aesthetic Preference · T-scores · all 25 scenes

Because preference is orthogonal to every quality metric, the move is straight up — raising motion, polish and appeal shifts Tap8 toward the frontier’s empty top-right corner without spending any of its informative lead.

NEXT
Win to close #1 — Marketing content. The collapse is fact delivery in marketing scenes, not rendering — a targeted extension of the coverage engine that already leads four use cases, not a rebuild. Closing it converts Tap8’s information-dense coverage lead into a full-field one: excluding Marketing, Tap8 already tops five of six systems on coverage outright.
NEXT
Win to close #2 — aesthetic preference vs pixel-diffusion, without giving back the informative lead. Preference is empirically orthogonal to every quality metric, so closing it is its own program — motion, polish, appeal — not a trade forced against coverage or delivery. The information-dense frame shows the ceiling: where content is dense, Tap8 already reaches preference parity while holding the content lead.
Appendix

Frozen protocol, seeded statistics, every result in full

System versions are pinned to the day with sources; sampling rules and fact checklists were fixed before any video existed; the annotator panel was blind and fully crossed (A–B).
Three independent estimators agree on every ordering, all randomness is seeded, and an independent re-run reproduced every artifact (C–E).
Complete leaderboards with confidence intervals for every metric, in both frames (F–K).
Δ columns are mixed-logit log-odds contrasts vs Tap8 unless noted; WIN/LOSS = the 95% interval excludes zero, from Tap8’s perspective; TIE = it does not.

KEY
System naming. The body of this report refers to three systems by anonymized labels. Their real identities, stated here only:
Code-render H = Hyperframes · Code-render R = Remotion · Pixel-diffusion L = LTX.
The tables below use the real product names.
A · Experiment protocol — versions, timeline, annotation ledger

Timeline (2026). Briefs + fact checklists authored and frozen before Jul 6 → systems frozen Jul 6 → all generation Jul 6–12 → VBench-Long scoring Jul 13–14 → human annotation Jul 13–15 (first response Jul 13 09:41, last Jul 15 04:33) → statistics and human↔VBench alignment Jul 14–16. VBench scoring completed before any human-label comparison existed.

Version pins — locally-run tools at their latest release as of the Jul 6 freeze; hosted platforms at the latest flagship model during the window. All 16:9, maximum quality settings, storyboard/planning features as shipped by each platform.

SystemVersion · platform · settingsSource
Tap8internal build current at the access window
Fable 5Claude Code 2.1.202 · model claude-fable-5 (announced Jun 9, redeployed Jul 1) · defaultsannouncement · changelog
Hyperframesv0.7.37 · Claude Code · Sonnet 5 (released Jun 30) · defaultsrelease · Sonnet 5
Remotionv4.0.485 · Codex CLI 0.142.5 · GPT-5.5 (announced Apr 23; GPT-5.6 shipped Jul 9, after the freeze) · defaultsrelease · Codex · GPT-5.5
VeoVeo 3.1 (announced 2025-10-15) · Google Flow, storyboard planning by Flow · Quality mode, 720pannouncement · modes doc
LTXLTX-2.3 (released Mar 2026) · LTX Studio, storyboard planning by LTX Studio using GPT-Image-2 · Pro mode, 1080prelease · GPT-Image-2
SeedanceSeedance 2.0 (announced Feb 12) · Jimeng — ByteDance’s official consumer platform — Agent Mode, story-film skill · VIP mode (top quality), 720pannouncement · Jimeng

Prompt provenance. All 25 briefs and their fact checklists were authored with Claude Opus on Claude Code, human-reviewed, and frozen before generation began. If AI authorship biases the briefs toward any system’s idiom, the beneficiary would be Fable 5 — a competitor — not Tap8.

The 6,760-annotation ledger. One row per annotator × scene × system × item: fact-check 2,326 (168 facts) + TeachQuiz comprehension 2,062 (149 questions) + T2VWorldBench four dimensions 1,384 (346 video×rater units × 4) + text accuracy 692 (346 units × grade + junk flag) + side-by-side preference 296 = 6,760. Every count is recomputable from the raw export.

VBench run. Official VBench-Long code (vbench2_beta_long, commit e946a8d) on 5× NVIDIA L4, one shard per use case, ~6 h wall; environment pinned and reproduced from scratch on a fresh instance; per-cell result JSONs and run logs retained. Input normalization and the >3.7M evaluation arithmetic: appendix E.

B · Worked example — one brief through the whole pipeline

The brief (How-To scene 1): “The Paper That Pops — A No-Glue Origami Fidget Toy”, target 5:00, 16:9. Like every brief, it is a fully-specified production document, identical for all seven systems: art direction with an exact hex palette, lighting and camera language, a nine-scene shot list with full voice-over text and per-scene motion notes, and a single master prompt — e.g. “…the three color-coded papers (indigo 6×6 outer box, amber 5×5 inside button, violet 9×4 spring), the shared blintz fold that builds both boxes… the no-glue lock where flaps tuck under adjacent edges…” The full information payload lives in the brief, so coverage failures are attributable to the system, not the prompt.

Its required-fact checklist (authored with the brief, frozen before generation; station items rendered in English). A fact is credited only if it appears on screen verbatim and legible — annotators may replay frame-by-frame:

  • Outer box paper: 6×6 cm
  • Inner button paper: 5×5 cm
  • Spring paper: 9×4 cm
  • The no-glue promise: no glue, no tape, anywhere
  • The key fold, by name: blintz fold (corners to center)

A TeachQuiz comprehension item (answered after viewing; the “video did not provide this” escape is scored incorrect): “What is the key fold that both boxes share?” — valley fold / blintz fold (corners to center) / squash fold / the video did not provide this information.

How it scored. On this scene both of Tap8’s raters credited 5/5 facts and answered 12/12 comprehension questions correctly. The same items caught real misses elsewhere in the field on this scene — including the fold-name question — which is the discrimination the benchmark is built to measure.

C · Statistical methods — estimators, seeds, and independent re-verification

Estimators. Binary stations (coverage, delivery): Bayesian mixed logistic regression value ~ system + (1|scene) + (1|item), Tap8 as reference. Graded stations (text, dimensions): scene-cluster-robust OLS with annotator fixed effects. Delivery additionally via leave-one-out information gain. Three independent estimators — raw rates, information gain, mixed logit — produce the identical system ordering in both frames.

Uncertainty. 95% CIs from scene-level bootstrap (resample the 25 scenes, B = 3000, fixed seed); significance = the 95% interval excludes zero; results reported as tiers where intervals overlap, never false-precision ranks. VBench: Skillings–Mack omnibus per dimension (method differences significant at p = 8e-18 to 1e-9 over 23 complete scene blocks), post-hoc pairwise Wilcoxon signed-rank with Holm correction. Human↔VBench alignment: Spearman rank correlation, stress-tested by 2000× bootstrap, leave-one-system-out, and permutation tests (all seeded).

Reproducibility — independently re-verified. All analysis randomness is seeded, raw data (the full annotation export and per-scene VBench scores) ships with the analysis, and each metric has its own one-command reproduce script under pinned dependencies (Python 3.14.5 · numpy 2.5.1 · pandas 3.0.3 · scipy 1.18.0 · statsmodels 0.14.6 · matplotlib 3.11.0 · openpyxl 3.1.5). An independent re-run in a fresh environment regenerated all 55 committed result artifacts: every deterministic output byte-identical, all 15 figures pixel-identical, and the variational mixed-model columns matching to ≤6e-06 log-odds — with no significance verdict changing anywhere.

D · Reliability, method, and limitations

Rater robustness by design. The panel is fully crossed — every annotator rated every system — so rater strictness cannot bias a between-system comparison. Adding an annotator random effect moves contrasts by at most 0.16 (delivery) / 0.31 (coverage) log-odds with no significance flips.

Inter-rater agreement per metric. Delivery: 73.6% raw agreement, Gwet AC1 0.59. Coverage: 67.2%, AC1 0.42 — the annotator effect is the dominant nuisance source (sd 0.98), so the absolute coverage level is rater-sensitive even though the comparison is not. Text: within-one-grade 91%, weighted κ 0.27 — the noisiest rating. WorldBench dimensions: within-one 79–85%, weighted κ 0.24–0.35. Aesthetic Preference: weighted κ 0.07 — aggregate win-rates over 296 comparisons are meaningful; individual judgments are not, which also attenuates the orthogonality correlations toward zero.

Limitations. Delivery is relative to this seven-system field and not comparable across benchmark runs with a different field. Comprehension is administered open-book; the 0.86 median ceiling compresses system differences, and a closed-book run is the highest-leverage sharpening. 25 scenes (5 per use case) with 2 raters per video bound per-use-case resolution — the top tiers do not fully separate. Coverage credits a fact only when verbatim and legible. Mixed-logit intervals are variational (anti-conservative); the large significant contrasts are safe. The arena’s star topology supports no competitor-vs-competitor ranking and no absolute aesthetic score. T-scores are field-relative conveniences, not absolute grades. A facts-per-second bandwidth station is defined but was not computed in this round.

E · VBench ↔ human alignment — method

Instruments. VBench Frame Quality is the mean of the model’s Aesthetic Quality and Imaging Quality dimensions, computed quality-when-it-works (a failed generation is excluded, not scored 0). Human preference is the side-by-side win-rate over Tap8. The correlation is over the six competitor systems — Tap8 is the star-topology anchor of every comparison and so is not a plotted point.

Alignment (VBench Frame Quality ↔ human preference)Spearman ρ
Overall+0.89
Leave-one-system-out (worst case)+0.87
Professional+0.94
News+0.77
Marketing+0.70
How-To+0.48
Explainer−0.33

The overall sign survives 2000× bootstrap resampling and leave-one-system-out; n = 6 systems bounds formal significance. Explainer is the one segment where viewers reward motion over frame-polish, so the frame-quality metric tracks a different cue there. Every VBench composite was recomputed from the raw per-scene scores and matches the source report to the decimal.

Run provenance and the >3.7M count. VBench-Long (vbench2_beta_long at commit e946a8d; VBench++, IEEE TPAMI 2025) was run on NVIDIA L4 GPUs on Jul 13–14, 2026 — before any human-label comparison existed. Inputs were normalized per scene so no system gains an encoding artifact: 720p cap (never upscaled), one uniform 24 fps frame grid, near-lossless re-encode. Six dimensions per video; four of them score every frame (24 model forwards/s each), motion smoothness runs 11.5/s and dynamic degree 7.5/s — 116 model forwards per second of footage. Across the 32,141 seconds of scored video: 3.73 million frame-level model evaluations. A further consistency re-check on length-trimmed builds (not counted above) verified that video length does not confound the comparison: native−trim deltas all < 0.0025.

F · Composite scorecard — Informative Quality and the three axes

Field-relative T-scores: 70 = the mean of these seven systems, every 10 points = 1 SD. Informative Quality = equal-weight mean of the Content Fidelity and Visual Quality axes. Aesthetic Preference is arena-derived and reported beside the informative axes, never inside them — it runs inverse to them. Overall Quality, (Content + Visual + Aesthetic Preference)/3, is the named information-dense scenario headline.

All 25 scenes

SystemInformative Quality95% CIContentVisualAesthetic Preference
Fable 585.4[78.3, 92.1]82.688.252.3
Tap878.2[69.3, 86.6]77.179.261.7
Remotion72.3[61.5, 82.2]76.668.169.2
Hyperframes71.8[59.3, 83.0]73.370.268.1
Veo64.8[51.7, 77.0]66.063.674.6
Seedance61.3[46.8, 75.1]60.362.283.1
LTX56.3[41.8, 70.8]54.158.581.0

Information-dense scenes (Professional, News, How-To, Explainer)

SystemOverall Quality (C+V+P)/3ContentVisualAesthetic PreferenceInformative Quality
Tap875.283.680.361.782.0
Fable 575.082.888.154.285.5
Seedance73.571.769.779.170.7
LTX70.060.274.875.167.5
Remotion69.777.565.965.771.7
Hyperframes69.072.968.365.770.6
Veo68.471.061.772.466.3

Bootstrap CIs over the 25 scenes overlap widely: the tiers the data supports are Fable 5 · Tap8 · Remotion/Hyperframes · the pixel trio. The Overall Quality Tap8–Fable 5 gap (75.2 vs 75.0) is within noise. The ranking is robust to method (PCA vs axis-mean), reliability weighting, and where Relevance is grouped. Tap8’s Aesthetic Preference is 0.5-anchored by the arena’s star topology; its per-use-case variation lives in the competitors.

G · Factual Coverage — both frames

All 25 scenes — share of required facts on screen, verbatim and legible; 95% bootstrap CI; mixed-logit contrast.

SystemCoverage95% CIΔ vs Tap895% CIResult
Fable 50.790[0.717, 0.860]+0.094[−0.183, +0.370]TIE
Tap80.773[0.699, 0.843]
Remotion0.759[0.668, 0.845]−0.166[−0.417, +0.085]TIE
Hyperframes0.745[0.642, 0.842]−0.199[−0.448, +0.050]TIE
Seedance0.611[0.506, 0.710]−0.826[−1.052, −0.601]WIN
Veo0.596[0.494, 0.697]−0.867[−1.091, −0.642]WIN
LTX0.541[0.433, 0.658]−1.179[−1.400, −0.958]WIN

Information-dense scenes

SystemCoverage95% CIΔ vs Tap895% CIResult
Tap80.836[0.773, 0.890]
Fable 50.792[0.710, 0.863]−0.291[−0.601, +0.020]TIE
Remotion0.765[0.669, 0.857]−0.520[−0.801, −0.239]WIN
Hyperframes0.727[0.607, 0.835]−0.681[−0.952, −0.409]WIN
Seedance0.677[0.578, 0.770]−0.922[−1.182, −0.662]WIN
Veo0.601[0.497, 0.714]−1.213[−1.464, −0.963]WIN
LTX0.545[0.418, 0.671]−1.553[−1.798, −1.307]WIN

Coverage ↔ delivery: Pearson r = 0.70 across 173 videos; pooled OLS gives P(comprehended | on screen) = 0.96 and a 0.36 off-screen floor; a mixed logit agrees (0.945 / 0.30). Coverage’s absolute level is rater-sensitive (see D); the between-system comparison is not.

H · TeachQuiz Information Delivery — both frames

All 25 scenes — raw comprehension correct-rate, mixed-logit predicted P(correct) at a typical scene/question, and contrast.

SystemRaw ratePredicted PΔ vs Tap895% CIResult
Fable 50.8210.868+0.243[−0.091, +0.576]TIE
Remotion0.8120.857+0.147[−0.167, +0.461]TIE
Tap80.7950.838
Hyperframes0.7890.834−0.028[−0.330, +0.274]TIE
Veo0.7850.830−0.052[−0.353, +0.249]TIE
LTX0.7150.756−0.509[−0.786, −0.233]WIN
Seedance0.6910.730−0.647[−0.918, −0.377]WIN

Information-dense scenes

SystemRaw rateΔ vs Tap895% CIResult
Fable 50.846+0.056[−0.342, +0.453]TIE
Tap80.845
Veo0.840+0.005[−0.368, +0.378]TIE
Remotion0.832−0.067[−0.433, +0.299]TIE
Hyperframes0.790−0.393[−0.731, −0.054]WIN
Seedance0.786−0.423[−0.759, −0.087]WIN
LTX0.752−0.650[−0.971, −0.329]WIN

Delivery is relative to this seven-system field (leave-one-out gain / field-baselined logit); administration is open-book, which compresses differences toward the 0.86 median ceiling. Three estimators (raw, gain D, mixed logit) agree on the ordering in both frames.

I · T2VTextBench Text Accuracy and T2VWorldBench dimensions

T2VTextBench Text Accuracy — four-tier grade on [0,1]; cluster-robust contrast vs Tap8. * = 95% CI excludes zero.

SystemAll 2595% CIInfo-dense95% CIΔ all-25
Fable 50.783[0.714, 0.850]0.750[0.681, 0.825]+0.062
Tap80.740[0.660, 0.820]0.769[0.675, 0.856]
Hyperframes0.700[0.615, 0.785]0.706[0.613, 0.794]−0.009
Remotion0.695[0.610, 0.785]0.688[0.594, 0.781]−0.050
Seedance0.640[0.535, 0.740]0.700[0.594, 0.806]−0.087
Veo0.630[0.530, 0.730]0.650[0.531, 0.762]−0.090
LTX0.580[0.500, 0.660]0.619[0.537, 0.694]−0.145*

T2VWorldBench, all 25 scenes — mean grade per dimension and the paper composite. * = significant vs Tap8 (cluster-robust OLS).

SystemQualityRealismRelevanceConsistencyComposite
Fable 50.900*0.9090.8910.9090.902*
Tap80.8600.8800.8440.8720.864
Remotion0.8000.8600.8600.8240.836
Hyperframes0.8120.8520.8160.8440.831
Veo0.8000.8040.8080.8240.809
Seedance0.8080.8080.772*0.796*0.796*
LTX0.7920.784*0.688*0.792*0.764*

The composite is internally valid (Cronbach α = 0.91) but folds one content-facing axis: Relevance correlates with Delivery/Coverage (+0.77) more than with its sibling visual dimensions (0.61–0.66). Info-dense frame: Tap8 stays 2nd on every dimension; only the Quality and Composite gaps to Fable 5 remain significant.

J · Aesthetic Preference (arena-style) — both frames, both readings

Star topology: every comparison is Tap8 vs one competitor on the same brief (296 comparisons). Win-rate counts ties as half; non-loss counts ties for Tap8; both are reported. Verdicts from binomial mixed logits vs the 0.5 no-preference line.

All 25 scenes

vsWin-rate95% CINon-loss95% CIVerdict
Fable 50.587[0.482, 0.683]0.696[0.577, 0.808]Tap8 does-not-lose*
Hyperframes0.440[0.347, 0.531]0.560[0.458, 0.667]toss-up
Remotion0.430[0.338, 0.519]0.540[0.433, 0.643]toss-up
Veo0.380[0.297, 0.464]0.480[0.389, 0.571]toss-up
LTX0.320[0.217, 0.417]0.380[0.267, 0.500]Tap8 loses*
Seedance0.300[0.206, 0.393]0.400[0.286, 0.500]Tap8 loses*
Overall0.407[0.35, 0.47]0.507[0.44, 0.57]less preferred than the field*

Information-dense scenes — overall win-rate 0.432 [0.37, 0.50], non-loss 0.530: parity. The significant non-loss deficits vanish; on strict decisive comparisons only Seedance remains significantly preferred.

vsWin-rate95% CINon-loss95% CI
Fable 50.569[0.450, 0.692]0.667[0.542, 0.800]
Remotion0.463[0.354, 0.568]0.575[0.458, 0.692]
Hyperframes0.463[0.361, 0.563]0.600[0.500, 0.714]
Veo0.400[0.304, 0.500]0.500[0.385, 0.607]
LTX0.375[0.250, 0.500]0.425[0.292, 0.545]
Seedance0.338[0.225, 0.450]0.425[0.286, 0.567]

Preference is empirically orthogonal to the informative metrics: correlating Tap8’s per-video advantage on each metric with the arena outcome gives −0.22 to −0.01 (n = 148 cells) — all negative or near zero. Non-loss is the favourable reading (20% of comparisons are ties); both are shown for that reason.

K · Per-use-case standings — where Tap8 wins and loses

Significant wins/losses vs Tap8 per use case (95% interval excludes zero). Cells list competitors; “—” = none.

Use caseCoverage: Tap8 beatsloses toDelivery: Tap8 beatsloses to
ProfessionalFable 5, Hyperframes, Veo, LTXFable 5, LTXSeedance
NewsRemotion, Veo, LTX, SeedanceRemotion, LTX, Seedance
How-ToRemotion, Hyperframes, Veo, LTXVeo
ExplainerLTX, SeedanceFable 5SeedanceFable 5
MarketingSeedanceFable 5, Remotion, HyperframesSeedanceFable 5, Remotion, Hyperframes

No system wins everywhere: Seedance leads Professional delivery yet collapses on Marketing (P 0.30); Veo ties the News delivery lead yet is worst-in-field on How-To. Tap8’s arena win-rate by use case: Professional 0.450, News 0.433, Explainer 0.429, How-To 0.417, Marketing 0.308. Tap8’s composite Informative Quality by use case: News 87, How-To 89, Explainer 81, Professional 71, Marketing 63.

Tap8 Generation Benchmark · 7 systems · 5 use cases · 25 intent-driven scenes · 2026.
Metrics: TeachQuiz Information Delivery (Code2Video, ICML 2026 · arXiv:2510.01174) · Factual Coverage · T2VTextBench Text Accuracy (arXiv:2505.04946) · T2VWorldBench four dimensions (WACV 2026 · arXiv:2507.18107) · Aesthetic Preference (arena-style) · VBench-Long automated video-quality (VBench++, IEEE TPAMI 2025 · VBench, CVPR 2024 Highlight · HuggingFace leaderboard).
6,760 human annotations from a fully-crossed 12-annotator panel (two raters per video) · >3.7M VBench frame evaluations on the same scenes, scored before any human-label comparison.
Generation Jul 6–12 · annotation Jul 13–15 · 2026. All human and VBench analyses are deterministic and reproducible from the raw exports under pinned dependencies — independently re-verified byte-for-byte; synthesis via reproduce_overview.py.