Skip to content

Fix WIS: use w_median * |obs - median| on the array backends - #141

Open
lucaluo925 wants to merge 1 commit into
frazane:mainfrom
lucaluo925:fix-wis-median-term
Open

lucaluo925 wants to merge 1 commit into
frazane:mainfrom
lucaluo925:fix-wis-median-term

Conversation

@lucaluo925

Copy link
Copy Markdown

Summary

Fixes #140.

In scoringrules/core/interval/_score.py, changed WIS += w_median * median to WIS += w_median * B.abs(obs - median). Per Bracher et al. (2021) Eq. (3), this term is the weight multiplied by the absolute deviation between observation and median, not the median itself. The numba gufunc implementation was already correct; this change aligns the shared array backend path with it. Since _score.py is shared across numpy, jax, and torch, this fixes all three backends at once.

Also updated documentation: the docstring in _interval.py stated the default weight is w_k = 2/alpha_k, while the actual code uses alpha/2. The code was correct (consistent with the paper); only the docstring is updated.

Tests

The original test_weighted_interval_score passed the same obs value as both observation and median, so |obs - median| = 0. Both formulas yielded identical results, and the CI for all four backends remained green, hiding the bug. The test is updated to obs = 0.4, median = 0 so this term contributes meaningfully, with a comment explaining why observation and median must differ.

Added test_weighted_interval_score_median_term, which asserts directly against the closed-form solution from Eq. (3) in the paper, without relying on the approximation that "WIS approximates CRPS".

Validation

After reverting the source change, the two tests fail under the numpy backend (0.458 vs expected 0.338), while the numba backend still passes — confirming the bug only affects non-numba paths. After restoring the fix, all 6 tests in tests/test_interval.py pass. Environment: Python 3.14, numpy 2.5.3, numba 0.67.0.

jax and torch are not installed in my environment, so I could only exercise the numpy and numba backends locally; CI will be the first run covering all four.

AI note

The code changes and tests were written with assistance from an AI assistant. I reproduced the bug locally, ran the tests, and verified that the tests fail when the fix is reverted.

Checklist

  • Tests pass across all backends — uv sync --all-extras --dev then uv run pytest tests/ (exercises every installed backend: numpy, numba, jax, torch)
  • New/changed numerics touch both the array-API and numba paths where applicable
  • Docs / docstrings updated if public behavior changed
  • A release-note label is applied (breaking, enhancement, bug, backend, documentation, or ci)

(I can't apply labels myself — bug looks like the right one here.)

@github-actions github-actions Bot added the core Core scoring API / numerics label Sep 29, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

core Core scoring API / numerics

Projects

None yet

Development

Successfully merging this pull request may close these issues.

weighted_interval_score adds w_median * median instead of w_median * |obs - median| on all non-numba backends

1 participant