Skip to content

[Klaud Cold] refactor: remove SWE-bench Lite eval suite - #3506

Draft
functionstackx wants to merge 2 commits into
mainfrom
klaud/remove-swebench-lite
Draft

functionstackx wants to merge 2 commits into
mainfrom
klaud/remove-swebench-lite

Conversation

@functionstackx

@functionstackx functionstackx commented Sep 26, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Remove the SWE-bench Lite eval suite and every reference to it across the codebase.
SWE-bench Lite is unused -- no master config selects it, and the eval framework has moved past it.

Deleted files (5)

  • infx/evals/swebench_lite.yaml -- task definition
  • infx/evals/swebench_score.py -- scorer
  • infx/evals/patches/patch_swebench_agent.py -- agent runtime patch
  • infx/evals/patches/patch_swebench_scoring.py -- scoring runtime patch
  • infx/tests/evals/test_swebench_eval.py -- dedicated test module

Edited files (28)

  • Workflows: benchmark-tmpl.yml, benchmark-multinode-tmpl.yml, e2e-tests.yml -- removed swebench inputs, env vars, secrets, artifact globs
  • Shell scripts: benchmark_lib.sh (removed ~280-line swebench function + dispatch case), runtime_settings.sh, srt_eval.sh
  • Launchers: launch_h100-cr.sh, launch_mi325x-tw.sh -- removed SWEBENCH_/MODAL_ docker env passthrough
  • Slurm infra: slurm_utils.sh, amd_utils/job.slurm, amd_utils/submit.sh, llm-d/job.slurm, llm-d/submit.sh
  • Tests: test_run_eval_dispatch.py (~470 lines of swebench/modal tests removed + unused signal import), test_eval_patches.py (dead _load_patch_module helper + unused imports), test_generate_sweep_configs.py (docstring)
  • Docs (EN+ZH): architecture.md, ci-procedures.md, eval-agentx-procedures.md, index.md, results-and-ingestion.md and their _zh.md counterparts
  • Eval docs: EVALS.md -- removed SWE-bench Lite section, patch entries, task file reference
  • Thresholds: thresholds.yaml -- removed swebench_lite: 0.50 entry

Deps removed

  • Modal credential plumbing (MODAL_TOKEN_ID / MODAL_TOKEN_SECRET) was exclusively consumed by swebench and is removed everywhere.
  • No pyproject.toml / lock changes needed (mini-swe-agent and swe-rex were never declared as deps).

Validation

  • bash -n passes on all touched .sh / .slurm files
  • uvx ruff check infx/ -- all checks passed
  • pytest infx/tests/evals/ infx/tests/matrix/ runners/ -- 607 passed
  • YAML workflow parsing -- all 4 files valid
  • git grep -n -i -E 'swe.?bench|mini.?swe|SWEBENCH' -- . ':!perf-changelog.yaml' -- zero hits

Test plan

  • CI e2e-tests workflow passes (no swebench env vars expected)
  • Benchmark template workflows validate (no swebench inputs/secrets referenced)
  • Existing non-swebench evals (lm-eval, BFCL, AgentX, vendor) unaffected
  • Multi-node slurm jobs launch without referencing removed env vars

Follow-up (41ea1b9)

  • .github/workflows/run-sweep.yml: stop passing MODAL_TOKEN_ID / MODAL_TOKEN_SECRET to benchmark-tmpl.yml / benchmark-multinode-tmpl.yml (12 call sites, 24 lines). The templates no longer declare those secrets, and GitHub rejects reusable-workflow calls that pass undeclared secrets, so every sweep would have failed to start. No MODAL references remain under .github/. The MODAL_TOKEN_* repo secrets themselves can be deleted separately.

🤖 Generated with Claude Code

SWE-bench Lite is unused: no master config selects it, and the eval
framework has moved past it.  Remove the task YAML, scorer, patches,
tests, and every reference in workflows, shell scripts, slurm jobs,
launcher env passthrough, docs (EN + ZH), and EVALS.md.  Drop Modal
credential plumbing (MODAL_TOKEN_ID / MODAL_TOKEN_SECRET) which was
exclusively consumed by swebench.

Deleted:
  infx/evals/swebench_lite.yaml
  infx/evals/swebench_score.py
  infx/evals/patches/patch_swebench_agent.py
  infx/evals/patches/patch_swebench_scoring.py
  infx/tests/evals/test_swebench_eval.py

Co-Authored-By: Claude Opus 4.6 <[email protected]>
@github-actions

Copy link
Copy Markdown
Contributor

Thanks for the contribution!

  • Review: If this PR changes files owned by someone other than a repository admin or @SemiAnalysisAI/core, ask one eligible CODEOWNER to complete the latest PR_REVIEW_CHECKLIST.md before contacting a core maintainer on Slack. Follow the template exactly, including As a PR reviewer and CODEOWNER, I have reviewed this and have, so sign-off verification triggers.
  • PR verification: Sweeps only run on labeled PRs. Add full-sweep-fail-fast (strongly recommended); use full-sweep-enabled only when matrix jobs should continue after a failure.
  • After merging: PR authors must ensure all GitHub Actions jobs pass. Transient failures often pass on rerun; see how to rerun failed jobs.
中文

感谢你的贡献!

  • **审阅:**如果 PR 修改的文件归属于仓库管理员及 @SemiAnalysisAI/core 之外的 CODEOWNER,请先联系一位有资格的 CODEOWNER 填写最新的 PR_REVIEW_CHECKLIST.md,再通过 Slack 联系核心维护者。必须严格遵循模板,并保留 As a PR reviewer and CODEOWNER, I have reviewed this and have,才能触发签核验证。
  • **PR 验证:**扫描仅在带有标签的 PR 上运行。强烈建议添加 full-sweep-fail-fast;仅当需要矩阵任务在失败后继续运行时才使用 full-sweep-enabled。
  • **合并后:**PR 作者必须确保所有 GitHub Actions 任务通过。临时性失败通常可以通过重新运行恢复;参见重新运行失败任务的说明。

The benchmark templates no longer declare these secrets, and GitHub rejects
reusable-workflow calls that pass undeclared secrets.

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
@adibarra

Copy link
Copy Markdown
Collaborator

Heads-up: #3576 (merged) replaced the bash launchers with a Python launcher, so this PR will conflict when you merge main, and the sweep won't start until that's resolved. Please merge main and move your launcher changes over to configs/runners.yaml / infx/launch/. Apologies for the churn, and thanks for your understanding as we wrap up the repo-wide refactoring push.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Status: No status

Development

Successfully merging this pull request may close these issues.

2 participants