A benchmark for video that must inform — instructional, editorial, explanatory — not entertain. On the information-dense scenes it targets, Tap8 takes the outright Factual Coverage lead, ties for first on TeachQuiz Information Delivery, and tops Overall Quality across seven systems. The independent, industry-standard automated metric, VBench, reproduces the human panel’s preference, confirming the eval’s rigor.
Information-dense scenes — professional, news, how-to, education — are the benchmark’s target domain, and they are where Tap8 wins. Across all 25 scenes, Tap8 is the clear 2nd of 7 on Informative Quality and the only code renderer strong on both informative axes.
Both frames are reported everywhere. Headline claims use the information-dense frame (UC1–UC4) the benchmark targets — shown here, where Tap8 tops Overall Quality; on all 25 scenes Tap8 is 2nd of 7 on Informative Quality (T 78 vs Fable 5’s 85). Full scorecards with confidence intervals are in the appendix.
Each of five use cases draws on real-world source material and is evaluated on five intent-driven scenes: the brief specifies the facts a viewer should walk away with, and human annotators measure whether they do. The target is high-information-density, clear-intent video — instructional, editorial, explanatory — not aesthetic or entertainment content.
| System | Harness / model | Family |
|---|---|---|
| Tap8 | code-rendered video | code-render |
| Fable 5 | Claude Code · Fable 5 | code-render |
| Code-render H | Claude Code · Sonnet 5 | code-render |
| Code-render R | Codex · GPT-5.5 | code-render |
| Pixel-diffusion L | official studio · Pro mode, 1080p | pixel-diffusion |
| Seedance | Seedance 2.0 · Jimeng · VIP mode, 720p | pixel-diffusion |
| Veo | Veo 3.1 · Google Flow · Quality mode, 720p | pixel-diffusion |
All systems accessed 2026 Jul 6–12: locally-run tools frozen at their latest release as of Jul 6, hosted systems at the latest flagship model on their official platform — every system at its maximum quality setting, 16:9. Exact version pins with sources are in appendix A.
| Use case | Source material |
|---|---|
| Professional Enterprise / self-learning | GDPval tasks — Professional & Technical Services, Government, Manufacturing, Finance |
| News Editorial | 2025–26 covers: Economist, FT, Bloomberg, Nature, Science |
| How-To Lifestyle | Top #HowTo YouTube — craft/DIY, art, fitness, mindset, puzzles |
| Explainer Education | Most-viewed “AP Exam explained” — history, psych, calculus, biology, geography |
| Marketing Product / sales | NASDAQ most-valuable-company launch videos across sectors |
Sampling rules fixed before generation. Each use case draws its five scenes by a predefined rule — cluster the source pool, randomly sample one per cluster (top industries, publications, content categories, AP subjects, NASDAQ sectors). Every brief is a fully-detailed video-generation prompt carrying the complete information payload, authored and frozen before any system generated anything; its required-fact checklist (168 facts across the benchmark) and comprehension questions (149) were authored at the same time and human-reviewed. A complete worked example is in appendix B.
Every video is scored by a fully-crossed annotator panel — every annotator rated every system, so a strict or lenient rater moves all systems equally and cannot bias a comparison.
The panel produced 6,760 human annotations across five stations grounded in published methods — three of them peer-reviewed at ICML, WACV, and IEEE TPAMI / CVPR — and every video is also scored by the VBench automated model: >3.7 million frame-level model evaluations.
VBench is an independent, machine-based cross-check of our human annotations: its Frame Quality evaluation reproduces our human preference ranking, so the two instruments corroborate each other (Finding 05).
The rankings are robust to how you compute them: three independent estimators — raw rates, leave-one-out information gain, and mixed-effects logistic models — produce the identical ordering, and reliability-weighted and PCA variants leave it unchanged. Because Marketing behaves as a different task family from the four information-dense use cases, every result is reported in two frames: all 25 scenes, and the information-dense frame (Professional, News, How-To, Explainer). Full tables for both are in the appendix.
Aesthetic Preference is measured arena-style — two videos of the same brief, side by side — and captures something the informative rubrics do not: its correlation with every informative metric is negative-to-zero (−0.22 to −0.01). Code renderers buy informative quality; pixel-diffusion buys aesthetic appeal. The frontier below places all seven systems on that trade-off, in both frames.
Tap8 is the only system in the green quadrant — highly informative and aesthetically preferred — in both frames. The frontier runs downhill: Informative Quality orders Fable 5 > Tap8 > the rest, and Aesthetic Preference inverts it end-for-end. On information-dense scenes Tap8 moves to the frontier’s elbow (Informative Quality 82, within 3.5 T of the leader) while the pixel-diffusion aesthetic edge shrinks to parity; on all 25 scenes Seedance and Pixel-diffusion L sit above the preference range shown. Aesthetic Preference is reported beside Informative Quality, never averaged into it.
Above 50 = viewers prefer Tap8. Significance from the mixed-logit non-loss / decisive-win tests; full CIs and the non-loss frame are in the appendix.
On the benchmark’s target domain the aesthetic deficits shrink to parity — only Seedance remains significantly preferred on decisive comparisons.
Coverage is the benchmark’s most discriminating metric: the field splits into two clean tiers, with the four code renderers (74–79% of facts on screen) sitting well clear of the three pixel-diffusion systems (54–61%).
Info-dense frame: Tap8 83.6 [77.3–89.0] leads outright; the gap to every system except Fable 5 is statistically significant (mixed-logit 95% CI excludes 0). All-25 frame: Fable 5 79.0, Tap8 77.3 — a statistical tie — with Seedance, Veo and Pixel-diffusion L significantly behind in both frames.
Delivery follows the TeachQuiz protocol: after watching, annotators answer multiple-choice questions about the facts the brief required, and each video is scored against a per-question field baseline. Three independent estimators — raw correct-rate, leave-one-out information gain, and a mixed-effects logit — agree on the ordering.
Tap8 is top-tier in both frames. Full per-use-case win/tie/loss tables are in the appendix.
Tap8 is top-tier on News (88.3) and How-To (89.7) — per the mixed-logit inference it significantly out-delivers Code-render R, Pixel-diffusion L and Seedance on News and Veo on How-To, losing to none in either. Professional and Explainer are mid-field; Marketing is the gap the close-out plan targets. Each use case rests on 5 scenes, so the CIs are wide; significance calls come from the per-use-case mixed logit in the appendix.
The four dimensions are rated 1–5 by the crossed annotator panel against per-dimension rubrics. The one significant visual gap to Fable 5 is Video Quality (86.0 vs 90.0); the other three dimensions are statistical ties.
Tap8 is 2nd on Quality, Realism and Consistency, 3rd on Relevance (behind Fable 5 and Code-render R). Relevance — “does it say the right things” — behaves as a content dimension, and is analyzed with the content metrics above.
A widely-used automated model, computed with no knowledge of the human labels, agrees with the human panel on how the systems rank — evidence that the eval captures a real, reproducible signal rather than rater idiosyncrasy.
Each small dot is one video’s automated Frame Quality, coloured by system and placed at that system’s human win-rate; the large dots are the system means. The coloured clouds stack in win-rate order and their centres climb left-to-right — automated frame quality and human preference rank the systems together (Spearman ρ = +0.89). Systems are left unnamed: the point is the agreement of the two instruments, not any one system’s position.
Both gaps are measured, localized, and separable from what already works. Neither calls for touching the coverage-and-delivery engine that wins the information-dense domain.
On Marketing, Tap8’s Visual Quality holds at 75 while Content Fidelity drops to 51 — the videos render cleanly but miss required facts. Everywhere else the two axes move together at 68–93.
Because preference is orthogonal to every quality metric, the move is straight up — raising motion, polish and appeal shifts Tap8 toward the frontier’s empty top-right corner without spending any of its informative lead.
System versions are pinned to the day with sources; sampling rules and fact checklists were fixed before any video existed; the annotator panel was blind and fully crossed (A–B).
Three independent estimators agree on every ordering, all randomness is seeded, and an independent re-run reproduced every artifact (C–E).
Complete leaderboards with confidence intervals for every metric, in both frames (F–K).
Δ columns are mixed-logit log-odds contrasts vs Tap8 unless noted; WIN/LOSS = the 95% interval excludes zero, from Tap8’s perspective; TIE = it does not.
Timeline (2026). Briefs + fact checklists authored and frozen before Jul 6 → systems frozen Jul 6 → all generation Jul 6–12 → VBench-Long scoring Jul 13–14 → human annotation Jul 13–15 (first response Jul 13 09:41, last Jul 15 04:33) → statistics and human↔VBench alignment Jul 14–16. VBench scoring completed before any human-label comparison existed.
Version pins — locally-run tools at their latest release as of the Jul 6 freeze; hosted platforms at the latest flagship model during the window. All 16:9, maximum quality settings, storyboard/planning features as shipped by each platform.
| System | Version · platform · settings | Source |
|---|---|---|
| Tap8 | internal build current at the access window | — |
| Fable 5 | Claude Code 2.1.202 · model claude-fable-5 (announced Jun 9, redeployed Jul 1) · defaults | announcement · changelog |
| Hyperframes | v0.7.37 · Claude Code · Sonnet 5 (released Jun 30) · defaults | release · Sonnet 5 |
| Remotion | v4.0.485 · Codex CLI 0.142.5 · GPT-5.5 (announced Apr 23; GPT-5.6 shipped Jul 9, after the freeze) · defaults | release · Codex · GPT-5.5 |
| Veo | Veo 3.1 (announced 2025-10-15) · Google Flow, storyboard planning by Flow · Quality mode, 720p | announcement · modes doc |
| LTX | LTX-2.3 (released Mar 2026) · LTX Studio, storyboard planning by LTX Studio using GPT-Image-2 · Pro mode, 1080p | release · GPT-Image-2 |
| Seedance | Seedance 2.0 (announced Feb 12) · Jimeng — ByteDance’s official consumer platform — Agent Mode, story-film skill · VIP mode (top quality), 720p | announcement · Jimeng |
Prompt provenance. All 25 briefs and their fact checklists were authored with Claude Opus on Claude Code, human-reviewed, and frozen before generation began. If AI authorship biases the briefs toward any system’s idiom, the beneficiary would be Fable 5 — a competitor — not Tap8.
The 6,760-annotation ledger. One row per annotator × scene × system × item: fact-check 2,326 (168 facts) + TeachQuiz comprehension 2,062 (149 questions) + T2VWorldBench four dimensions 1,384 (346 video×rater units × 4) + text accuracy 692 (346 units × grade + junk flag) + side-by-side preference 296 = 6,760. Every count is recomputable from the raw export.
VBench run. Official VBench-Long code (vbench2_beta_long, commit e946a8d) on 5× NVIDIA L4, one shard per use case, ~6 h wall; environment pinned and reproduced from scratch on a fresh instance; per-cell result JSONs and run logs retained. Input normalization and the >3.7M evaluation arithmetic: appendix E.
The brief (How-To scene 1): “The Paper That Pops — A No-Glue Origami Fidget Toy”, target 5:00, 16:9. Like every brief, it is a fully-specified production document, identical for all seven systems: art direction with an exact hex palette, lighting and camera language, a nine-scene shot list with full voice-over text and per-scene motion notes, and a single master prompt — e.g. “…the three color-coded papers (indigo 6×6 outer box, amber 5×5 inside button, violet 9×4 spring), the shared blintz fold that builds both boxes… the no-glue lock where flaps tuck under adjacent edges…” The full information payload lives in the brief, so coverage failures are attributable to the system, not the prompt.
Its required-fact checklist (authored with the brief, frozen before generation; station items rendered in English). A fact is credited only if it appears on screen verbatim and legible — annotators may replay frame-by-frame:
A TeachQuiz comprehension item (answered after viewing; the “video did not provide this” escape is scored incorrect): “What is the key fold that both boxes share?” — valley fold / blintz fold (corners to center) / squash fold / the video did not provide this information.
How it scored. On this scene both of Tap8’s raters credited 5/5 facts and answered 12/12 comprehension questions correctly. The same items caught real misses elsewhere in the field on this scene — including the fold-name question — which is the discrimination the benchmark is built to measure.
Estimators. Binary stations (coverage, delivery): Bayesian mixed logistic regression value ~ system + (1|scene) + (1|item), Tap8 as reference. Graded stations (text, dimensions): scene-cluster-robust OLS with annotator fixed effects. Delivery additionally via leave-one-out information gain. Three independent estimators — raw rates, information gain, mixed logit — produce the identical system ordering in both frames.
Uncertainty. 95% CIs from scene-level bootstrap (resample the 25 scenes, B = 3000, fixed seed); significance = the 95% interval excludes zero; results reported as tiers where intervals overlap, never false-precision ranks. VBench: Skillings–Mack omnibus per dimension (method differences significant at p = 8e-18 to 1e-9 over 23 complete scene blocks), post-hoc pairwise Wilcoxon signed-rank with Holm correction. Human↔VBench alignment: Spearman rank correlation, stress-tested by 2000× bootstrap, leave-one-system-out, and permutation tests (all seeded).
Reproducibility — independently re-verified. All analysis randomness is seeded, raw data (the full annotation export and per-scene VBench scores) ships with the analysis, and each metric has its own one-command reproduce script under pinned dependencies (Python 3.14.5 · numpy 2.5.1 · pandas 3.0.3 · scipy 1.18.0 · statsmodels 0.14.6 · matplotlib 3.11.0 · openpyxl 3.1.5). An independent re-run in a fresh environment regenerated all 55 committed result artifacts: every deterministic output byte-identical, all 15 figures pixel-identical, and the variational mixed-model columns matching to ≤6e-06 log-odds — with no significance verdict changing anywhere.
Rater robustness by design. The panel is fully crossed — every annotator rated every system — so rater strictness cannot bias a between-system comparison. Adding an annotator random effect moves contrasts by at most 0.16 (delivery) / 0.31 (coverage) log-odds with no significance flips.
Inter-rater agreement per metric. Delivery: 73.6% raw agreement, Gwet AC1 0.59. Coverage: 67.2%, AC1 0.42 — the annotator effect is the dominant nuisance source (sd 0.98), so the absolute coverage level is rater-sensitive even though the comparison is not. Text: within-one-grade 91%, weighted κ 0.27 — the noisiest rating. WorldBench dimensions: within-one 79–85%, weighted κ 0.24–0.35. Aesthetic Preference: weighted κ 0.07 — aggregate win-rates over 296 comparisons are meaningful; individual judgments are not, which also attenuates the orthogonality correlations toward zero.
Limitations. Delivery is relative to this seven-system field and not comparable across benchmark runs with a different field. Comprehension is administered open-book; the 0.86 median ceiling compresses system differences, and a closed-book run is the highest-leverage sharpening. 25 scenes (5 per use case) with 2 raters per video bound per-use-case resolution — the top tiers do not fully separate. Coverage credits a fact only when verbatim and legible. Mixed-logit intervals are variational (anti-conservative); the large significant contrasts are safe. The arena’s star topology supports no competitor-vs-competitor ranking and no absolute aesthetic score. T-scores are field-relative conveniences, not absolute grades. A facts-per-second bandwidth station is defined but was not computed in this round.
Instruments. VBench Frame Quality is the mean of the model’s Aesthetic Quality and Imaging Quality dimensions, computed quality-when-it-works (a failed generation is excluded, not scored 0). Human preference is the side-by-side win-rate over Tap8. The correlation is over the six competitor systems — Tap8 is the star-topology anchor of every comparison and so is not a plotted point.
| Alignment (VBench Frame Quality ↔ human preference) | Spearman ρ |
|---|---|
| Overall | +0.89 |
| Leave-one-system-out (worst case) | +0.87 |
| Professional | +0.94 |
| News | +0.77 |
| Marketing | +0.70 |
| How-To | +0.48 |
| Explainer | −0.33 |
The overall sign survives 2000× bootstrap resampling and leave-one-system-out; n = 6 systems bounds formal significance. Explainer is the one segment where viewers reward motion over frame-polish, so the frame-quality metric tracks a different cue there. Every VBench composite was recomputed from the raw per-scene scores and matches the source report to the decimal.
Run provenance and the >3.7M count. VBench-Long (vbench2_beta_long at commit e946a8d; VBench++, IEEE TPAMI 2025) was run on NVIDIA L4 GPUs on Jul 13–14, 2026 — before any human-label comparison existed. Inputs were normalized per scene so no system gains an encoding artifact: 720p cap (never upscaled), one uniform 24 fps frame grid, near-lossless re-encode. Six dimensions per video; four of them score every frame (24 model forwards/s each), motion smoothness runs 11.5/s and dynamic degree 7.5/s — 116 model forwards per second of footage. Across the 32,141 seconds of scored video: 3.73 million frame-level model evaluations. A further consistency re-check on length-trimmed builds (not counted above) verified that video length does not confound the comparison: native−trim deltas all < 0.0025.
Field-relative T-scores: 70 = the mean of these seven systems, every 10 points = 1 SD. Informative Quality = equal-weight mean of the Content Fidelity and Visual Quality axes. Aesthetic Preference is arena-derived and reported beside the informative axes, never inside them — it runs inverse to them. Overall Quality, (Content + Visual + Aesthetic Preference)/3, is the named information-dense scenario headline.
All 25 scenes
| System | Informative Quality | 95% CI | Content | Visual | Aesthetic Preference |
|---|---|---|---|---|---|
| Fable 5 | 85.4 | [78.3, 92.1] | 82.6 | 88.2 | 52.3 |
| Tap8 | 78.2 | [69.3, 86.6] | 77.1 | 79.2 | 61.7 |
| Remotion | 72.3 | [61.5, 82.2] | 76.6 | 68.1 | 69.2 |
| Hyperframes | 71.8 | [59.3, 83.0] | 73.3 | 70.2 | 68.1 |
| Veo | 64.8 | [51.7, 77.0] | 66.0 | 63.6 | 74.6 |
| Seedance | 61.3 | [46.8, 75.1] | 60.3 | 62.2 | 83.1 |
| LTX | 56.3 | [41.8, 70.8] | 54.1 | 58.5 | 81.0 |
Information-dense scenes (Professional, News, How-To, Explainer)
| System | Overall Quality (C+V+P)/3 | Content | Visual | Aesthetic Preference | Informative Quality |
|---|---|---|---|---|---|
| Tap8 | 75.2 | 83.6 | 80.3 | 61.7 | 82.0 |
| Fable 5 | 75.0 | 82.8 | 88.1 | 54.2 | 85.5 |
| Seedance | 73.5 | 71.7 | 69.7 | 79.1 | 70.7 |
| LTX | 70.0 | 60.2 | 74.8 | 75.1 | 67.5 |
| Remotion | 69.7 | 77.5 | 65.9 | 65.7 | 71.7 |
| Hyperframes | 69.0 | 72.9 | 68.3 | 65.7 | 70.6 |
| Veo | 68.4 | 71.0 | 61.7 | 72.4 | 66.3 |
Bootstrap CIs over the 25 scenes overlap widely: the tiers the data supports are Fable 5 · Tap8 · Remotion/Hyperframes · the pixel trio. The Overall Quality Tap8–Fable 5 gap (75.2 vs 75.0) is within noise. The ranking is robust to method (PCA vs axis-mean), reliability weighting, and where Relevance is grouped. Tap8’s Aesthetic Preference is 0.5-anchored by the arena’s star topology; its per-use-case variation lives in the competitors.
All 25 scenes — share of required facts on screen, verbatim and legible; 95% bootstrap CI; mixed-logit contrast.
| System | Coverage | 95% CI | Δ vs Tap8 | 95% CI | Result |
|---|---|---|---|---|---|
| Fable 5 | 0.790 | [0.717, 0.860] | +0.094 | [−0.183, +0.370] | TIE |
| Tap8 | 0.773 | [0.699, 0.843] | — | ||
| Remotion | 0.759 | [0.668, 0.845] | −0.166 | [−0.417, +0.085] | TIE |
| Hyperframes | 0.745 | [0.642, 0.842] | −0.199 | [−0.448, +0.050] | TIE |
| Seedance | 0.611 | [0.506, 0.710] | −0.826 | [−1.052, −0.601] | WIN |
| Veo | 0.596 | [0.494, 0.697] | −0.867 | [−1.091, −0.642] | WIN |
| LTX | 0.541 | [0.433, 0.658] | −1.179 | [−1.400, −0.958] | WIN |
Information-dense scenes
| System | Coverage | 95% CI | Δ vs Tap8 | 95% CI | Result |
|---|---|---|---|---|---|
| Tap8 | 0.836 | [0.773, 0.890] | — | ||
| Fable 5 | 0.792 | [0.710, 0.863] | −0.291 | [−0.601, +0.020] | TIE |
| Remotion | 0.765 | [0.669, 0.857] | −0.520 | [−0.801, −0.239] | WIN |
| Hyperframes | 0.727 | [0.607, 0.835] | −0.681 | [−0.952, −0.409] | WIN |
| Seedance | 0.677 | [0.578, 0.770] | −0.922 | [−1.182, −0.662] | WIN |
| Veo | 0.601 | [0.497, 0.714] | −1.213 | [−1.464, −0.963] | WIN |
| LTX | 0.545 | [0.418, 0.671] | −1.553 | [−1.798, −1.307] | WIN |
Coverage ↔ delivery: Pearson r = 0.70 across 173 videos; pooled OLS gives P(comprehended | on screen) = 0.96 and a 0.36 off-screen floor; a mixed logit agrees (0.945 / 0.30). Coverage’s absolute level is rater-sensitive (see D); the between-system comparison is not.
All 25 scenes — raw comprehension correct-rate, mixed-logit predicted P(correct) at a typical scene/question, and contrast.
| System | Raw rate | Predicted P | Δ vs Tap8 | 95% CI | Result |
|---|---|---|---|---|---|
| Fable 5 | 0.821 | 0.868 | +0.243 | [−0.091, +0.576] | TIE |
| Remotion | 0.812 | 0.857 | +0.147 | [−0.167, +0.461] | TIE |
| Tap8 | 0.795 | 0.838 | — | ||
| Hyperframes | 0.789 | 0.834 | −0.028 | [−0.330, +0.274] | TIE |
| Veo | 0.785 | 0.830 | −0.052 | [−0.353, +0.249] | TIE |
| LTX | 0.715 | 0.756 | −0.509 | [−0.786, −0.233] | WIN |
| Seedance | 0.691 | 0.730 | −0.647 | [−0.918, −0.377] | WIN |
Information-dense scenes
| System | Raw rate | Δ vs Tap8 | 95% CI | Result |
|---|---|---|---|---|
| Fable 5 | 0.846 | +0.056 | [−0.342, +0.453] | TIE |
| Tap8 | 0.845 | — | ||
| Veo | 0.840 | +0.005 | [−0.368, +0.378] | TIE |
| Remotion | 0.832 | −0.067 | [−0.433, +0.299] | TIE |
| Hyperframes | 0.790 | −0.393 | [−0.731, −0.054] | WIN |
| Seedance | 0.786 | −0.423 | [−0.759, −0.087] | WIN |
| LTX | 0.752 | −0.650 | [−0.971, −0.329] | WIN |
Delivery is relative to this seven-system field (leave-one-out gain / field-baselined logit); administration is open-book, which compresses differences toward the 0.86 median ceiling. Three estimators (raw, gain D, mixed logit) agree on the ordering in both frames.
T2VTextBench Text Accuracy — four-tier grade on [0,1]; cluster-robust contrast vs Tap8. * = 95% CI excludes zero.
| System | All 25 | 95% CI | Info-dense | 95% CI | Δ all-25 |
|---|---|---|---|---|---|
| Fable 5 | 0.783 | [0.714, 0.850] | 0.750 | [0.681, 0.825] | +0.062 |
| Tap8 | 0.740 | [0.660, 0.820] | 0.769 | [0.675, 0.856] | — |
| Hyperframes | 0.700 | [0.615, 0.785] | 0.706 | [0.613, 0.794] | −0.009 |
| Remotion | 0.695 | [0.610, 0.785] | 0.688 | [0.594, 0.781] | −0.050 |
| Seedance | 0.640 | [0.535, 0.740] | 0.700 | [0.594, 0.806] | −0.087 |
| Veo | 0.630 | [0.530, 0.730] | 0.650 | [0.531, 0.762] | −0.090 |
| LTX | 0.580 | [0.500, 0.660] | 0.619 | [0.537, 0.694] | −0.145* |
T2VWorldBench, all 25 scenes — mean grade per dimension and the paper composite. * = significant vs Tap8 (cluster-robust OLS).
| System | Quality | Realism | Relevance | Consistency | Composite |
|---|---|---|---|---|---|
| Fable 5 | 0.900* | 0.909 | 0.891 | 0.909 | 0.902* |
| Tap8 | 0.860 | 0.880 | 0.844 | 0.872 | 0.864 |
| Remotion | 0.800 | 0.860 | 0.860 | 0.824 | 0.836 |
| Hyperframes | 0.812 | 0.852 | 0.816 | 0.844 | 0.831 |
| Veo | 0.800 | 0.804 | 0.808 | 0.824 | 0.809 |
| Seedance | 0.808 | 0.808 | 0.772* | 0.796* | 0.796* |
| LTX | 0.792 | 0.784* | 0.688* | 0.792* | 0.764* |
The composite is internally valid (Cronbach α = 0.91) but folds one content-facing axis: Relevance correlates with Delivery/Coverage (+0.77) more than with its sibling visual dimensions (0.61–0.66). Info-dense frame: Tap8 stays 2nd on every dimension; only the Quality and Composite gaps to Fable 5 remain significant.
Star topology: every comparison is Tap8 vs one competitor on the same brief (296 comparisons). Win-rate counts ties as half; non-loss counts ties for Tap8; both are reported. Verdicts from binomial mixed logits vs the 0.5 no-preference line.
All 25 scenes
| vs | Win-rate | 95% CI | Non-loss | 95% CI | Verdict |
|---|---|---|---|---|---|
| Fable 5 | 0.587 | [0.482, 0.683] | 0.696 | [0.577, 0.808] | Tap8 does-not-lose* |
| Hyperframes | 0.440 | [0.347, 0.531] | 0.560 | [0.458, 0.667] | toss-up |
| Remotion | 0.430 | [0.338, 0.519] | 0.540 | [0.433, 0.643] | toss-up |
| Veo | 0.380 | [0.297, 0.464] | 0.480 | [0.389, 0.571] | toss-up |
| LTX | 0.320 | [0.217, 0.417] | 0.380 | [0.267, 0.500] | Tap8 loses* |
| Seedance | 0.300 | [0.206, 0.393] | 0.400 | [0.286, 0.500] | Tap8 loses* |
| Overall | 0.407 | [0.35, 0.47] | 0.507 | [0.44, 0.57] | less preferred than the field* |
Information-dense scenes — overall win-rate 0.432 [0.37, 0.50], non-loss 0.530: parity. The significant non-loss deficits vanish; on strict decisive comparisons only Seedance remains significantly preferred.
| vs | Win-rate | 95% CI | Non-loss | 95% CI |
|---|---|---|---|---|
| Fable 5 | 0.569 | [0.450, 0.692] | 0.667 | [0.542, 0.800] |
| Remotion | 0.463 | [0.354, 0.568] | 0.575 | [0.458, 0.692] |
| Hyperframes | 0.463 | [0.361, 0.563] | 0.600 | [0.500, 0.714] |
| Veo | 0.400 | [0.304, 0.500] | 0.500 | [0.385, 0.607] |
| LTX | 0.375 | [0.250, 0.500] | 0.425 | [0.292, 0.545] |
| Seedance | 0.338 | [0.225, 0.450] | 0.425 | [0.286, 0.567] |
Preference is empirically orthogonal to the informative metrics: correlating Tap8’s per-video advantage on each metric with the arena outcome gives −0.22 to −0.01 (n = 148 cells) — all negative or near zero. Non-loss is the favourable reading (20% of comparisons are ties); both are shown for that reason.
Significant wins/losses vs Tap8 per use case (95% interval excludes zero). Cells list competitors; “—” = none.
| Use case | Coverage: Tap8 beats | loses to | Delivery: Tap8 beats | loses to |
|---|---|---|---|---|
| Professional | Fable 5, Hyperframes, Veo, LTX | — | Fable 5, LTX | Seedance |
| News | Remotion, Veo, LTX, Seedance | — | Remotion, LTX, Seedance | — |
| How-To | Remotion, Hyperframes, Veo, LTX | — | Veo | — |
| Explainer | LTX, Seedance | Fable 5 | Seedance | Fable 5 |
| Marketing | Seedance | Fable 5, Remotion, Hyperframes | Seedance | Fable 5, Remotion, Hyperframes |
No system wins everywhere: Seedance leads Professional delivery yet collapses on Marketing (P 0.30); Veo ties the News delivery lead yet is worst-in-field on How-To. Tap8’s arena win-rate by use case: Professional 0.450, News 0.433, Explainer 0.429, How-To 0.417, Marketing 0.308. Tap8’s composite Informative Quality by use case: News 87, How-To 89, Explainer 81, Professional 71, Marketing 63.