fix(metrics): score only sequences that carry a supervised label in seq_acc - #10247
Merged
tastelikefeet merged 1 commit intoSep 26, 2026
Merged
Conversation
…eq_acc `np.all()` of an empty selection is True, so a sequence whose labels are all ignored (-100) was appended to `seq_acc` as a correct prediction: a batch with no supervised label at all reported seq_acc as 1.0, and unsupervised rows also inflated the mean through the trainer's MeanMetric denominator. The same empty selection made `AccMetrics.compute_metrics` divide by zero for token_acc once nothing was supervised. Skip sequences without a supervised label and return no metric when there is nothing to score, which is what the existing float/shape guards already do.
tastelikefeet
approved these changes
Sep 26, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
PR type
PR information
Why
compute_acc(..., acc_strategy='seq')decides whether a sequence was answered correctly withwhere
m = labels[i] != -100. When a sequence carries no supervised label at all (mis all-False), the selection is empty andnp.all([])isTrue, so that sequence is scored as a correct prediction. Measured onc08110b3:and for a batch where nothing is supervised the reported
seq_accmean is1.0({'seq_acc': [True, True]}).That list goes straight into the trainer:
swift/trainers/mixin.py:1222-1225feeds it toMeanMetric.update(v), whoseupdatecountslen(state)— so every unsupervised row adds 1 to both the numerator and the denominator.eval/seq_accis therefore inflated by exactly the rows that contain no information, and the more aggressively a run truncates, the higher it reads.The same empty selection also breaks the
tokenpath:acc_liststays empty,compute_accstill returns{'token_acc': []}, andAccMetrics.compute_metrics(acc.py:55) doessum(v) / len(v)→(reproduced with
--eval_metric accon an all--100eval set).compute_accalready returns{}for "nothing to score" at:19and:26, so both of these are just two paths that were never connected to that existing contract.How a row ends up without a supervised label is documented behaviour, not a corner of the code. With a non-default
--truncation_strategy(Literal['delete', 'left', 'right', 'split'],swift/arguments/base_args/template_args.py:132, documented indocs/source_en/Instruction/Command-line-parameters.md):right:template/base.py:_encode_truncatedcalls_truncate, which protects the firstmax_length - n_protectedunprotected positions and drops the rest — the tail. Since the response is the tail, a row cut hard enough keeps no supervised label at all.split:_encode_truncatedsliceslabels[i:i + max_length]into chunks and returns every chunk, so a sample longer thanmax_lengthcontributes trailing chunks that are pure padding-only context, i.e. all--100rows that are still fed to the trainer as if they were supervised sequences.delete(the default) drops such rows entirely, so this needs an explicit flag — but withright/splitset, an all--100row is a normal batch outcome rather than a corrupt-input edge case.What changed
if not mask.any(): continueon the padding_free branch andif not m.any(): continueon the padded branch.{}like the two existing early returns, which keepsAccMetrics.compute_metricsandmixin.py:1222on the "no metric this step" path they already handle instead of dividing by zero.No change for any sequence that has at least one supervised label: the value is the same
np.all(...)over the same mask.Tests
tests/utils/test_acc_metrics.py(the file added by #10049) grows four cases, keeping its existingunittest.TestCasestyle and fixtures:test_padded_seq_acc_skips_a_sequence_without_supervised_labels— the batch above now reports[False], i.e. one scored sequence instead of a hit and a miss.test_padding_free_seq_acc_skips_a_sequence_without_supervised_labels— same throughcu_seqlens.test_seq_acc_reports_nothing_when_no_label_is_supervisedandtest_token_acc_reports_nothing_when_no_label_is_supervised—{}instead of the empty list that crashed the metric hook.Bug fix verification
Red on unmodified
c08110b3(4 of the 6 tests fail), green with the change:Lint with the versions pinned in
.pre-commit-config.yaml:I did not run a training job. What was executed here is
compute_accitself — the[False, True]inflation, the1.0mean on an all--100batch, theZeroDivisionErrorfrom theAccMetricsmean, and the six unit tests; a row that keeps at least one supervised label returns byte-identical values before and after (checked on[[-100,-100,-100,7]]→{'seq_acc': [False]}in both). The two--truncation_strategyroutes above are read from_truncate/_encode_truncated, not produced by an end-to-end encode.Disclosure: this PR was prepared, tested and submitted by an AI agent working on behalf of the account owner.