Conversation
The client can already drive video-generation endpoints, but every shipped video config is a single measurement point. This records what is missing to run one as a multi-point pareto sweep, verified against 4235a9c, with requirements cited against endpoints_policies v1.0_rules_dev @ 6b0b1ef. That branch is unmerged and moving, so citations are marked as-of. Two things are already correct for a tokenless workload and need no change: Report.qps is n_completed over the tracked duration, which for video is videos/second, and a per-request latency percentile is already emitted with P90 in the default grid. The per-user rate mirrors v1.0's latency-reciprocal definition rather than dividing throughput by concurrency. The blockers are that no minimum-duration floor exists in the stop predicate; that min_issue_duration_ms is a poisson count-sizer rather than a floor and is rejected for concurrency; that steady-state detection is gated on a resolved tokenizer and streaming, which a video run lacks by construction, while v1.0 4.4 makes that verdict the official reporting basis; and that config-lock requires target_concurrency == 1 and so fails at every point of a curve. Includes a workaround that satisfies both the duration floor and 6.4's whole-pass requirement by sizing the sample count to the duration, an itemised task breakdown with dependencies and sizes, and an explicit out-of-scope section. Co-Authored-By: Claude Sonnet 4.6 <[email protected]>
Retitles and rescopes docs/videogen_pareto_support.md to docs/pareto_sweep_support.md. Five of the seven gaps are not specific to video at all. No minimum duration floor exists in the stop predicate, min_issue_duration_ms is a poisson count-sizer rejected for concurrency, config lock requires target_concurrency == 1, there is no sweep driver, and the per-user rate is not emitted. Each blocks a multi-point curve for any model, which is why existing curves are built by running each concurrency separately and stitching the results downstream. Only two gaps are tokenless-specific: steady state cannot certify a window without TPOT, and the video adapter routes a file path through the field the OSL trigger tokenizes. The gap table now carries a Scope column separating the two, and the task breakdown is grouped the same way. Video generation remains the motivating case, since it is what surfaced all of this, but it is no longer presented as the scope. Co-Authored-By: Claude Sonnet 4.6 <[email protected]>
Co-Authored-By: Claude Sonnet 4.6 <[email protected]>
Regrouping the task tables lost 'register a video benchmark as a ruleset model and dataset', and trimmed the temperature-field clause from the config-lock gap. Both are restored, with registration in the tokenless group since text models are already registered. Co-Authored-By: Claude Sonnet 4.6 <[email protected]>
Renames docs/pareto_sweep_support.md to docs/t2v_benchmark_gaps.md and leads with the workload rather than the mechanism, since the document exists to get a text-to-video benchmark running. The gap analysis is unchanged. The Scope column stays, because which gaps also block other models is useful to a reader deciding what to fix first: two gaps are specific to a tokenless workload, and the other five block a multi-point curve for any model. Co-Authored-By: Claude Sonnet 4.6 <[email protected]>
Opens with the seven categories and whether each is tokenless-specific or blocks any model, so the shape of the work is visible before the detail. Per-gap evidence stays in section 3 and the itemised tasks in section 5. Co-Authored-By: Claude Sonnet 4.6 <[email protected]>
Naming, Missing metrics, Duration support, Multi-point support, Validation support, Tokenless support, Example configs. Co-Authored-By: Claude Sonnet 4.6 <[email protected]>
Which gaps also affect other models is still recorded per gap in the section 3 table; the opening summary does not need it. Co-Authored-By: Claude Sonnet 4.6 <[email protected]>
Section 0 did not match the order of anything that followed. Gaps and tasks are now both ordered by the same seven categories, and each category links to its task group. Splits the old config-lock gap in two, since it conflated Naming (_resolve_model raising for an unregistered model) with Validation (the single-stream config lock). Eight gaps now, renumbered, with the references in sections 1 and 4 updated to match. Co-Authored-By: Claude Sonnet 4.6 <[email protected]>
Co-Authored-By: Claude Sonnet 4.6 <[email protected]>
Admissibility is settled; rules text will be added. The bullet framed it as an open working-group question, which is no longer accurate. Co-Authored-By: Claude Sonnet 4.6 <[email protected]>
'Nothing in the serving path is missing' could be read as covering the model server. It means the client side of the request path. Co-Authored-By: Claude Sonnet 4.6 <[email protected]>
Each item is either answered or already tracked as a task: calibration is answered by measured per-video latency, the sweep-driver question is task D1, and power normalisation is tracked outside this repository. Co-Authored-By: Claude Sonnet 4.6 <[email protected]>
Four dense paragraphs become short lists: what blocks a submission, what already exists, how the eight gaps split, and the scope. Co-Authored-By: Claude Sonnet 4.6 <[email protected]>
Co-Authored-By: Claude Sonnet 4.6 <[email protected]>
Co-Authored-By: Claude Sonnet 4.6 <[email protected]>
The document listed 8 gaps, 5 of which do not block a T2V benchmark. They surfaced while working out how to run a sweep, which is not the same as blocking it: text curves already work around all 5, and the same workaround serves video. Leaves 3 gaps plus example configs. Only one is a hard blocker: steady state cannot certify a window without tokens, and v1.0 4.4 makes that verdict the official reporting basis. The 5 become a short context note in 3.3, kept so the section 4 recipe still makes sense. One of them was also wrong to list at all: compliance/checker.py is scoped to Edge-Agentic submissions and reachable only through a standalone script, so it is not the validator for an Endpoints pareto submission and never was on this path. Co-Authored-By: Claude Sonnet 4.6 <[email protected]>
Section 2 listed three pieces of generic client behaviour with no bearing on a video workload: the ConcurrencyScheduler semaphore, performance-phase isolation, and target_concurrency landing in the run config. The last was also stale, since the per-user rate is now derived from latency rather than from concurrency. What remains in section 2 is there for a reason: throughput and latency because they are the video metrics, the infinite sample order because the section 4 recipe depends on it, TEST04 because the shipped video configs run it, and graceful tokenizer absence because video has no tokenizer. Co-Authored-By: Claude Sonnet 4.6 <[email protected]>
Co-Authored-By: Claude Sonnet 4.6 <[email protected]>
An independent review found the section 4 recipe sized against the wrong constraint, and verification against the rules confirmed it. Section 4.4 requires the steady window to span more than 4 super-passes (MIN_TREND_N = 4). Sizing on the duration floor alone gives N = 1 at low concurrency, which 4.4 classes insufficient_passes and does not count as an official result. The rule is now N = max(5, ceil(duration x throughput / dataset_size)). Section 4 also claimed to satisfy 6.2. It does not: 6.2 is measured over the steady window's issue-time span, which section 3.1 says is uncomputable for this workload. The recipe now distinguishes what it satisfies outright (6.4's issued count) from what it only makes possible (6.2 and 4.4, uncertifiable until the gating gap closes). Section 1's claim that neither rule is expressible is corrected: 6.4 is, and section 4 is the proof. Restores the stream_all_chunks conflict as gap 4, dropped when the list was cut to three. Section 6.5 mandates it true for all performance runs and a non-streaming workload cannot comply, so it needs a rules amendment rather than a client change. Adds two caveats the recipe was missing: termination counts issued while 6.4's minimum counts completed, so failures can end a point short; and 6.5 mandates WithReplacementSampleOrder, which the example now sets. Fixes the per-user-rate citation, which pointed at qps rather than the latency P90 it is actually the reciprocal of, and re-baselines the rules citations from 6b0b1ef to cdb203c. Co-Authored-By: Claude Sonnet 4.6 <[email protected]>
The sizing recipe set settings.runtime.sample_order, which RuntimeConfig does not declare and rejects as an extra key. RuntimeSettings.from_config always builds the without-replacement default, so no perf run can meet the v1.0 section 6.5 requirement for WithReplacementSampleOrder. Record it as gap 4 under a new Run settings category with task C1, drop the invalid key from the recipe, and renumber the following sections. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Review against the client at 4235a9c and the v1.0 rules at cdb203c. - Steady state: a tokenless point falls back to the whole-run total, which section 4.4 names the official fallback. It blocks a valid point only if the WG rules a not-found run invalid. CL-B1 is now conditional and sized L, with its real touch points listed. - stream_all_chunks is a plain client flag a video config can set; the rules-conflict gap is removed. - New gap: the example configs are not submission-shaped (144-prompt subset, 10-minute caps, 20-sample SingleStream, no submission_ref). - Recipe: N = 1 per point, seeds bound through submission_ref, check n_samples_failed because failures count as completed. Super-pass wording fixed: window of at least 4, run of more than 4. - Target shape follows the elected-Offline decision: 7 runs, 4 accuracy runs, C_min 1, C_max 72. - CL-A1 re-scoped to a dataset entry and the shared slug; CL-B2 keeps the path VBench needs. Task IDs prefixed CL-. - Adds a point timeline and an official-basis decision chart. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
A pending streaming-check change (d9e1ea6) would block every T2V point; a pending Rules TF patch (b36c220) makes the accuracy count exact without saying how an elected Offline point counts. Remove the point timeline and official-result charts. Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
Codecov Report✅ All modified and coverable lines are covered by tests. Additional details and impacted files@@ Coverage Diff @@
## main #523 +/- ##
=======================================
Coverage ? 81.35%
=======================================
Files ? 157
Lines ? 22495
Branches ? 0
=======================================
Hits ? 18301
Misses ? 4194
Partials ? 0 ☔ View full report in Codecov by Harness. 🚀 New features to boost your workflow:
|
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What does this PR do?
This is the gap closing plan for T2V benchmarks.
It includes the change to endpoints.
Type of change
Related issues
Testing
Checklist