Skip to content

[Docs/Plan] t2v benchmark gaps - #523

Open
wu6u3tw wants to merge 23 commits into
mlcommons:mainfrom
wu6u3tw:docs/t2v-benchmark-gaps
Open

wu6u3tw wants to merge 23 commits into
mlcommons:mainfrom
wu6u3tw:docs/t2v-benchmark-gaps

Conversation

@wu6u3tw

@wu6u3tw wu6u3tw commented Oct 1, 2026

Copy link
Copy Markdown
Collaborator

What does this PR do?

This is the gap closing plan for T2V benchmarks.
It includes the change to endpoints.

Type of change

  • Bug fix
  • New feature
  • Documentation update
  • Refactor/cleanup

Related issues

Testing

  • Tests added/updated
  • All tests pass locally
  • Manual testing completed

Checklist

  • Code follows project style
  • Pre-commit hooks pass
  • Documentation updated (if needed)

Tin-Yin Lai and others added 23 commits September 28, 2026 12:16
The client can already drive video-generation endpoints, but every shipped
video config is a single measurement point. This records what is missing to
run one as a multi-point pareto sweep, verified against 4235a9c, with
requirements cited against endpoints_policies v1.0_rules_dev @ 6b0b1ef.
That branch is unmerged and moving, so citations are marked as-of.

Two things are already correct for a tokenless workload and need no change:
Report.qps is n_completed over the tracked duration, which for video is
videos/second, and a per-request latency percentile is already emitted with
P90 in the default grid. The per-user rate mirrors v1.0's latency-reciprocal
definition rather than dividing throughput by concurrency.

The blockers are that no minimum-duration floor exists in the stop
predicate; that min_issue_duration_ms is a poisson count-sizer rather than a
floor and is rejected for concurrency; that steady-state detection is gated
on a resolved tokenizer and streaming, which a video run lacks by
construction, while v1.0 4.4 makes that verdict the official reporting
basis; and that config-lock requires target_concurrency == 1 and so fails at
every point of a curve.

Includes a workaround that satisfies both the duration floor and 6.4's
whole-pass requirement by sizing the sample count to the duration, an
itemised task breakdown with dependencies and sizes, and an explicit
out-of-scope section.

Co-Authored-By: Claude Sonnet 4.6 <[email protected]>
Retitles and rescopes docs/videogen_pareto_support.md to
docs/pareto_sweep_support.md.

Five of the seven gaps are not specific to video at all. No minimum
duration floor exists in the stop predicate, min_issue_duration_ms is a
poisson count-sizer rejected for concurrency, config lock requires
target_concurrency == 1, there is no sweep driver, and the per-user rate
is not emitted. Each blocks a multi-point curve for any model, which is
why existing curves are built by running each concurrency separately and
stitching the results downstream.

Only two gaps are tokenless-specific: steady state cannot certify a window
without TPOT, and the video adapter routes a file path through the field
the OSL trigger tokenizes. The gap table now carries a Scope column
separating the two, and the task breakdown is grouped the same way.

Video generation remains the motivating case, since it is what surfaced
all of this, but it is no longer presented as the scope.

Co-Authored-By: Claude Sonnet 4.6 <[email protected]>
Co-Authored-By: Claude Sonnet 4.6 <[email protected]>
Regrouping the task tables lost 'register a video benchmark as a ruleset
model and dataset', and trimmed the temperature-field clause from the
config-lock gap. Both are restored, with registration in the tokenless
group since text models are already registered.

Co-Authored-By: Claude Sonnet 4.6 <[email protected]>
Renames docs/pareto_sweep_support.md to docs/t2v_benchmark_gaps.md and
leads with the workload rather than the mechanism, since the document
exists to get a text-to-video benchmark running.

The gap analysis is unchanged. The Scope column stays, because which
gaps also block other models is useful to a reader deciding what to fix
first: two gaps are specific to a tokenless workload, and the other five
block a multi-point curve for any model.

Co-Authored-By: Claude Sonnet 4.6 <[email protected]>
Opens with the seven categories and whether each is tokenless-specific or
blocks any model, so the shape of the work is visible before the detail.
Per-gap evidence stays in section 3 and the itemised tasks in section 5.

Co-Authored-By: Claude Sonnet 4.6 <[email protected]>
Naming, Missing metrics, Duration support, Multi-point support,
Validation support, Tokenless support, Example configs.

Co-Authored-By: Claude Sonnet 4.6 <[email protected]>
Which gaps also affect other models is still recorded per gap in the
section 3 table; the opening summary does not need it.

Co-Authored-By: Claude Sonnet 4.6 <[email protected]>
Section 0 did not match the order of anything that followed. Gaps and
tasks are now both ordered by the same seven categories, and each
category links to its task group.

Splits the old config-lock gap in two, since it conflated Naming
(_resolve_model raising for an unregistered model) with Validation (the
single-stream config lock). Eight gaps now, renumbered, with the
references in sections 1 and 4 updated to match.

Co-Authored-By: Claude Sonnet 4.6 <[email protected]>
Co-Authored-By: Claude Sonnet 4.6 <[email protected]>
Admissibility is settled; rules text will be added. The bullet framed it
as an open working-group question, which is no longer accurate.

Co-Authored-By: Claude Sonnet 4.6 <[email protected]>
'Nothing in the serving path is missing' could be read as covering the
model server. It means the client side of the request path.

Co-Authored-By: Claude Sonnet 4.6 <[email protected]>
Each item is either answered or already tracked as a task: calibration is
answered by measured per-video latency, the sweep-driver question is task
D1, and power normalisation is tracked outside this repository.

Co-Authored-By: Claude Sonnet 4.6 <[email protected]>
Four dense paragraphs become short lists: what blocks a submission, what
already exists, how the eight gaps split, and the scope.

Co-Authored-By: Claude Sonnet 4.6 <[email protected]>
Co-Authored-By: Claude Sonnet 4.6 <[email protected]>
Co-Authored-By: Claude Sonnet 4.6 <[email protected]>
The document listed 8 gaps, 5 of which do not block a T2V benchmark. They
surfaced while working out how to run a sweep, which is not the same as
blocking it: text curves already work around all 5, and the same
workaround serves video.

Leaves 3 gaps plus example configs. Only one is a hard blocker: steady
state cannot certify a window without tokens, and v1.0 4.4 makes that
verdict the official reporting basis.

The 5 become a short context note in 3.3, kept so the section 4 recipe
still makes sense. One of them was also wrong to list at all:
compliance/checker.py is scoped to Edge-Agentic submissions and reachable
only through a standalone script, so it is not the validator for an
Endpoints pareto submission and never was on this path.

Co-Authored-By: Claude Sonnet 4.6 <[email protected]>
Section 2 listed three pieces of generic client behaviour with no bearing
on a video workload: the ConcurrencyScheduler semaphore, performance-phase
isolation, and target_concurrency landing in the run config. The last was
also stale, since the per-user rate is now derived from latency rather
than from concurrency.

What remains in section 2 is there for a reason: throughput and latency
because they are the video metrics, the infinite sample order because the
section 4 recipe depends on it, TEST04 because the shipped video configs
run it, and graceful tokenizer absence because video has no tokenizer.

Co-Authored-By: Claude Sonnet 4.6 <[email protected]>
An independent review found the section 4 recipe sized against the wrong
constraint, and verification against the rules confirmed it.

Section 4.4 requires the steady window to span more than 4 super-passes
(MIN_TREND_N = 4). Sizing on the duration floor alone gives N = 1 at low
concurrency, which 4.4 classes insufficient_passes and does not count as
an official result. The rule is now N = max(5, ceil(duration x throughput
/ dataset_size)).

Section 4 also claimed to satisfy 6.2. It does not: 6.2 is measured over
the steady window's issue-time span, which section 3.1 says is
uncomputable for this workload. The recipe now distinguishes what it
satisfies outright (6.4's issued count) from what it only makes possible
(6.2 and 4.4, uncertifiable until the gating gap closes). Section 1's
claim that neither rule is expressible is corrected: 6.4 is, and section
4 is the proof.

Restores the stream_all_chunks conflict as gap 4, dropped when the list
was cut to three. Section 6.5 mandates it true for all performance runs
and a non-streaming workload cannot comply, so it needs a rules
amendment rather than a client change.

Adds two caveats the recipe was missing: termination counts issued while
6.4's minimum counts completed, so failures can end a point short; and
6.5 mandates WithReplacementSampleOrder, which the example now sets.

Fixes the per-user-rate citation, which pointed at qps rather than the
latency P90 it is actually the reciprocal of, and re-baselines the rules
citations from 6b0b1ef to cdb203c.

Co-Authored-By: Claude Sonnet 4.6 <[email protected]>
The sizing recipe set settings.runtime.sample_order, which RuntimeConfig
does not declare and rejects as an extra key. RuntimeSettings.from_config
always builds the without-replacement default, so no perf run can meet
the v1.0 section 6.5 requirement for WithReplacementSampleOrder.

Record it as gap 4 under a new Run settings category with task C1, drop
the invalid key from the recipe, and renumber the following sections.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Review against the client at 4235a9c and the v1.0 rules at cdb203c.

- Steady state: a tokenless point falls back to the whole-run total,
  which section 4.4 names the official fallback. It blocks a valid point
  only if the WG rules a not-found run invalid. CL-B1 is now conditional
  and sized L, with its real touch points listed.
- stream_all_chunks is a plain client flag a video config can set; the
  rules-conflict gap is removed.
- New gap: the example configs are not submission-shaped (144-prompt
  subset, 10-minute caps, 20-sample SingleStream, no submission_ref).
- Recipe: N = 1 per point, seeds bound through submission_ref, check
  n_samples_failed because failures count as completed. Super-pass
  wording fixed: window of at least 4, run of more than 4.
- Target shape follows the elected-Offline decision: 7 runs, 4 accuracy
  runs, C_min 1, C_max 72.
- CL-A1 re-scoped to a dataset entry and the shared slug; CL-B2 keeps
  the path VBench needs. Task IDs prefixed CL-.
- Adds a point timeline and an official-basis decision chart.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
A pending streaming-check change (d9e1ea6) would block every T2V point; a pending Rules TF patch (b36c220) makes the accuracy count exact without saying how an elected Offline point counts. Remove the point timeline and official-result charts.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
@wu6u3tw
wu6u3tw requested a review from a team October 1, 2026 23:48
@github-actions github-actions Bot added the size/normal PR Review Policy: <=500 non-test lines & <=20 files label Oct 1, 2026
@mlcommons mlcommons deleted a comment from github-actions Bot Oct 1, 2026
@codecov-commenter

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
⚠️ Please upload report for BASE (main@4235a9c). Learn more about missing BASE report.

Additional details and impacted files
@@           Coverage Diff           @@
##             main     #523   +/-   ##
=======================================
  Coverage        ?   81.35%           
=======================================
  Files           ?      157           
  Lines           ?    22495           
  Branches        ?        0           
=======================================
  Hits            ?    18301           
  Misses          ?     4194           
  Partials        ?        0           

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

size/normal PR Review Policy: <=500 non-test lines & <=20 files

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants