Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
15 changes: 14 additions & 1 deletion AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -140,6 +140,19 @@ The aggregator is a separate process (`python -m inference_endpoint.async_utils.
- **Post-run service wait**: `drain_and_build_report` waits for service subprocesses **unbounded** (`wait_for_exit(None)`). `wait_for_exit` SIGKILLs on expiry, and the aggregator writes `final_snapshot.json` — the Report's primary source — as the last thing it does, so a parent-side deadline here trades a wedged-service hang for a lost snapshot. The abort path keeps its own ceiling (`interrupted_teardown_grace_s`, 30s), and the whole-run watchdog stays armed throughout.
- **Histogram bucket edges are dynamic per snapshot**: log-spaced over the observed `[min, max]`. Bucket count is fixed at construction; consumers MUST re-render from the snapshot's `(lo, hi, count)` triples each frame and MUST NOT track bucket-by-index across snapshots.

### SWE-bench Pyxis command transport

`evaluation/swebench_service/swebench_service/pyxis_persistent.py` owns one
long-lived Slurm step per agent environment. The container runs the packaged
`pyxis_command_worker.sh`, staged in a private control mount separate from tool
`/tmp`. Worker startup retries only confirmed prelaunch Slurm failures before any
request is published. A lock serializes callers through one atomically published
request directory and completion marker.
Commands run in fresh PID namespaces; accepted requests are never replayed after an
uncertain failure. `pyxis_environment.py` owns container creation, command result
mapping, and worker-before-container cleanup. All tool commands use the persistent
worker. Evaluation remains in `pyxis_worker.py`.

### CLI Modes

CLI is auto-generated from `config/schema.py` Pydantic models via cyclopts. Fields annotated with `cyclopts.Parameter(alias="--flag")` get flat shorthands; all other fields get auto-generated dotted flags (kebab-case).
Expand Down Expand Up @@ -283,7 +296,7 @@ src/inference_endpoint/
│ ├── types.py # Pydantic: VideoPathRequest, VideoPathResponse, VideoPayloadResponse
│ └── adapter.py # VideoGenAdapter (HttpRequestAdapter) + VideoGenAccumulator (no-op)
├── evaluation/ # Accuracy evaluation (extractor, scoring, livecodebench)
│ └── swebench_service/ # Isolated uv service for Docker-backed SWE-bench runs
│ └── swebench_service/ # Isolated uv service for Docker/Pyxis SWE-bench runs
├── compliance/ # Submission compliance checks (config-lock, accuracy gate, run validity)
│ ├── __init__.py
│ └── checker.py # check_submission() + Check/ComplianceReport (Edge-Agentic ruleset)
Expand Down
2 changes: 1 addition & 1 deletion examples/10_Agentic_Inference/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -260,7 +260,7 @@ Reference mean values are shown in parentheses.

| Metric | Kimi K3 | Qwen3.6-35B-A3B | DeepSeek-V4.1-Flash |
| ------------------ | ------------------------ | ------------------------ | ------------------------ |
| Inline accuracy | `>= 58.32%` (`58.9%`) | `>= 55.86%` (`56.43%`) | `>= 52.36%` (`53.16%`) |
| Inline accuracy | `>= 58.32%` (`58.9%`) | `>= 55.86%` (`56.43%`) | `>= 52.36%` (`53.16%`) |
| OSL per-turn mean¹ | `425-520` tokens (`472`) | `344-422` tokens (`383`) | `793-970` tokens (`882`) |
| SWE-bench accuracy | `>= 93.5%` (`94.83%`) | `>= 69%` (`71.7%`) | `>= 96.4%` (`97.5%`) |

Expand Down
4 changes: 2 additions & 2 deletions pyproject.toml
Original file line number Diff line number Diff line change
Expand Up @@ -96,7 +96,7 @@ dependencies = [
"colorama==0.4.6",
# Fix pytz-2024 import warning
"pytz==2026.1.post1",
"urllib3==2.7.0",
"urllib3==2.8.0",
"pyyaml==6.0.3",
# anyio is pulled in transitively (openai -> httpx -> anyio); pinned here to
# force past CVE-2026-63374 / CVE-2026-64847 (both fixed in 4.14.2), which
Expand Down Expand Up @@ -130,7 +130,7 @@ dev = [
# patched here, pinned only inside the bfcl fork. Closes CVE-2025-68146 /
# CVE-2026-22701 (filelock) and CVE-2026-22702 (virtualenv).
"filelock>=3.20.3",
"virtualenv>=20.36.1",
"virtualenv==21.7.13",
# pip is pulled in transitively (pip-audit -> pip-api -> pip); force past
# PYSEC-2026-3721 (fixed in 26.2), which otherwise fails `uv run pip-audit`.
"pip==26.2",
Expand Down
30 changes: 26 additions & 4 deletions src/inference_endpoint/evaluation/swebench_service/README.md
Original file line number Diff line number Diff line change
Expand Up @@ -66,10 +66,32 @@ never forwarded.

During generation, the service still uses mini-swe-agent for the agent loop and
model requests, but replaces its Docker environment with `PyxisEnvironment`. Every
trajectory receives a named, writable Pyxis container. Each tool call becomes an
overlapping `srun` step in that container, preserving filesystem changes across
turns. Tool commands run in private PID namespaces so one trajectory cannot signal
processes belonging to another trajectory.
trajectory receives a named, writable Pyxis container and one long-lived overlapping
`srun` command worker. Tool calls use atomic request and response files in the
container's private `/tmp` mount, avoiding Slurm step creation and Enroot startup
on every turn. Each worker handles one request at a time; it publishes a completion
marker only after the command exits and its output is closed.
Filesystem changes persist, while each command runs in a fresh shell
and private PID namespace. Commands cannot signal the worker or other trajectories;
remaining child processes are removed when the command's PID namespace exits.
The service must stay on the allocated node, and node-local `TMPDIR` is recommended
for the request files.

The command worker is the packaged `swebench_service/pyxis_command_worker.sh`.
The Python service copies it into the private `/tmp` mount and starts it with Bash
inside the task container; no service Python installation is needed in task images.

Command failures preserve their exit status and merged stdout/stderr. A command
timeout terminates its process group with a five-second kill grace; loss of the
worker, invalid responses, and driver deadlines fail the run as infrastructure
errors. An accepted request is never automatically replayed because its execution
may already have changed the repository. The worker is stopped and reaped before
its container and temporary files are removed.

All tool commands use the persistent worker.
Container initialization, worker startup, evaluation, and cleanup still use Slurm
steps. This reduces per-command scheduler traffic; it does not bypass allocation
limits or guarantee any particular end-to-end evaluation time.

After generation, the Pyxis worker evaluates each prediction in a fresh `srun`
container step because the Docker-based SWE-bench evaluator cannot run on the
Expand Down
Original file line number Diff line number Diff line change
@@ -0,0 +1,49 @@
#!/usr/bin/env bash
# SPDX-FileCopyrightText: Copyright (c) 2026 NVIDIA CORPORATION & AFFILIATES. All rights reserved.
# SPDX-License-Identifier: Apache-2.0

set -uo pipefail
root=$1
shift
interpreter=("$@")
request="$root/running"
touch "$root/ready" || exit 70

while [ ! -e "$root/stop" ]; do
if [ ! -d "$root/request" ]; then
sleep 0.05
continue
fi
# The host serializes callers and removes the previous result before publishing.
[ ! -e "$request" ] || exit 70
mv -- "$root/request" "$request" || exit 70
timeout_s=$(cat "$request/timeout") || exit 70
case "$timeout_s" in ''|*[!0-9]*|0) exit 70 ;; esac
# A sentinel preserves trailing newlines in the command and working directory.
cwd=$(cat "$request/cwd" && printf x) || exit 70
# Keep tool text out of supervisor argv so pkill -f cannot match it there.
unshare --pid --fork --mount-proc \
timeout -k 5 "$timeout_s" bash -c '
status=$1; cwd=$2; shift 2
command=$(cat <&3 && printf x) || exit 70
exec 3<&-
if cd -- "$cwd"; then
"$@" "${command%x}"
rc=$?
else
rc=125
fi
printf "%s\n" "$rc" > "$status" || exit 70
exit "$rc"
' command-status "$request/command_status" "${cwd%x}" \
"${interpreter[@]}" 3< "$request/command" > "$request/output" 2>&1
returncode=$?
timed_out=0
# Explicit exits 124/137 are command results, not timeout notifications.
if [ "$(cat "$request/command_status" 2>/dev/null)" != "$returncode" ]; then
case "$returncode" in 124|137) timed_out=1 ;; *) exit 70 ;; esac
Comment thread
leopck marked this conversation as resolved.
fi
size=$(wc -c < "$request/output") || exit 70
printf '%s %s %s\n' "$returncode" "$timed_out" "$((size))" > "$request/complete.tmp" || exit 70
mv -- "$request/complete.tmp" "$request/complete" || exit 70
done
Original file line number Diff line number Diff line change
Expand Up @@ -6,18 +6,23 @@
import logging
import os
import platform
import random
import re
import shutil
import subprocess
import tempfile
import threading
import time
import uuid
from pathlib import Path
from typing import Any

from pydantic import AliasChoices, BaseModel, Field

from .pyxis_persistent import PersistentExecChannel
from .pyxis_slurm import SRUN_MAX_ATTEMPTS as _SRUN_MAX_ATTEMPTS
from .pyxis_slurm import (
is_retryable_prelaunch_failure as _is_retryable_prelaunch_failure,
)
from .pyxis_slurm import wait_for_prelaunch_retry
from .runner import RunnerError

logger = logging.getLogger(__name__)
Expand Down Expand Up @@ -45,18 +50,8 @@
"SLURM_CONF",
)
_STEP_STATUS = "/tmp/.mlperf_srun_status"
_SRUN_MAX_ATTEMPTS = 5
_RETRYABLE_PRELAUNCH_ERRORS = (
"spank_sybil: rpc request error",
"required plugin spank_sybil.so",
"failed to connect to any sack sockets",
"failed to create token",
"curl: (56) connect tunnel failed",
"unable to confirm allocation for job",
)
_IGNORABLE_SRUN_PREAMBLE_LINES = {
"srun: lua: Checking requeue policy with options:",
}
_PERSISTENT_ROOT = "/.mlperf_persistent_exec"
_COMMAND_WORKER = Path(__file__).with_name("pyxis_command_worker.sh")
_STEP_SCRIPT = r"""set +e
status_path=$1
timeout_s=$2
Expand All @@ -73,14 +68,6 @@ def safe_srun_env() -> dict[str, str]:
return {name: os.environ[name] for name in _SAFE_SRUN_ENV if name in os.environ}


def _is_retryable_prelaunch_failure(status: str, output: str) -> bool:
"""Return whether Slurm rejected the step before its command started."""
if status != "pending":
return False
lowered = output.lower()
return any(marker in lowered for marker in _RETRYABLE_PRELAUNCH_ERRORS)


def build_srun_command(
*,
argv: list[str],
Expand Down Expand Up @@ -187,15 +174,7 @@ def run_srun_step(
).strip()
retryable = _is_retryable_prelaunch_failure(status, output)
if retryable and attempt < _SRUN_MAX_ATTEMPTS:
backoff_s = min(2**attempt, 16)
delay_s = backoff_s + random.uniform(0.0, backoff_s)
logger.warning(
"Retrying Pyxis pre-launch failure in %.1fs (attempt %d/%d)",
delay_s,
attempt,
_SRUN_MAX_ATTEMPTS,
)
time.sleep(delay_s)
wait_for_prelaunch_retry(attempt)
continue

if failure_path is not None:
Expand Down Expand Up @@ -232,6 +211,7 @@ class PyxisEnvironmentConfig(BaseModel):
env: dict[str, str] = Field(default_factory=dict)
timeout_s: int = Field(
default=30,
gt=0,
validation_alias=AliasChoices("timeout_s", "timeout"),
serialization_alias="timeout",
)
Expand All @@ -246,70 +226,89 @@ def __init__(self, **kwargs: Any):
self.name = f"mswe_{safe_run_id}_{uuid.uuid4().hex[:8]}"
self._tmp = tempfile.TemporaryDirectory(prefix=f"pyxis_{self.name}_")
self._tmp_dir = Path(self._tmp.name)
self._tmp_dir.chmod(0o1777)
self._lock = threading.Lock()
self._cleaned = False
try:
(self._tmp_dir / "tmp").mkdir(mode=0o700)
shutil.copyfile(_COMMAND_WORKER, self._tmp_dir / _COMMAND_WORKER.name)
# A no-op initializes and validates the named persistent container.
run_srun_step(
image=self.config.image,
name=self.name,
mounts=[(self._tmp_dir, "/tmp")],
mounts=[(self._tmp_dir / "tmp", "/tmp")],
workdir=self.config.cwd,
argv=["true"],
status_path=self._tmp_dir / Path(_STEP_STATUS).name,
status_path=self._tmp_dir / "tmp" / Path(_STEP_STATUS).name,
timeout_s=self.config.timeout_s,
failure_path=self.config.infrastructure_failure_path,
)
except RunnerError as exc:
self._persistent_channel = PersistentExecChannel(
self._tmp_dir / "channel",
self._persistent_server_command(),
safe_srun_env(),
failure_path=self.config.infrastructure_failure_path,
launch_timeout_s=self.config.timeout_s + 30,
)
self._persistent_channel.start()
except (RunnerError, OSError) as exc:
self.cleanup()
raise RunnerError(
f"failed to start Pyxis container for {self.config.image}"
) from exc

def _persistent_server_command(self) -> list[str]:
return build_srun_command(
name=self.name,
# Tool cleanup of /tmp must not remove the worker or its protocol files.
mounts=[
(self._tmp_dir / "tmp", "/tmp"),
(self._tmp_dir, _PERSISTENT_ROOT),
],
workdir=self.config.cwd,
argv=[
"env",
*(f"{key}={value}" for key, value in self.config.env.items()),
"unshare",
"--pid",
"--fork",
"--mount-proc",
"--kill-child",
"bash",
f"{_PERSISTENT_ROOT}/{_COMMAND_WORKER.name}",
f"{_PERSISTENT_ROOT}/channel",
*self.config.interpreter,
],
)

def execute(
self, action: dict[str, Any], cwd: str = "", *, timeout: int | None = None
) -> dict[str, Any]:
command = action.get("command", "")
logger.debug("Executing Pyxis command: %s", command)
argv = ["env"]
argv.extend(f"{key}={value}" for key, value in self.config.env.items())
argv.extend([*self.config.interpreter, command])
result = run_srun_step(
argv=argv,
status_path=self._tmp_dir / Path(_STEP_STATUS).name,
timeout_s=timeout or self.config.timeout_s,
failure_path=self.config.infrastructure_failure_path,
name=self.name,
mounts=[(self._tmp_dir, "/tmp")],
workdir=cwd or self.config.cwd,
timeout_s = self.config.timeout_s if timeout is None else timeout
result = self._persistent_channel.execute(
command=command,
cwd=cwd or self.config.cwd,
timeout_s=timeout_s,
)
output: dict[str, Any]
if result.returncode == 124:
if result.timed_out:
output = {
"output": result.stdout,
"output": result.output,
"returncode": -1,
"exception_info": "The command timed out",
"extra": {
"exception_type": "TimeoutExpired",
"exception": (
f"command timed out after {timeout or self.config.timeout_s}s"
),
"exception": f"command timed out after {timeout_s}s",
},
}
else:
output = {
"output": result.stdout,
"output": result.output,
"returncode": result.returncode,
"exception_info": "",
}
lines = output.get("output", "").lstrip().splitlines(keepends=True)
# Some Slurm cli_filter plugins write informational messages to stderr.
# run_srun_step merges stderr into stdout so command errors remain visible,
# which can place this cluster-generated preamble before mini-swe-agent's
# otherwise first-line submission marker.
while lines and lines[0].strip() in _IGNORABLE_SRUN_PREAMBLE_LINES:
lines.pop(0)
if (
lines
and lines[0].strip() == "COMPLETE_TASK_AND_SUBMIT_FINAL_OUTPUT"
Expand Down Expand Up @@ -353,6 +352,14 @@ def cleanup(self) -> None:
return
self._cleaned = True
try:
channel = getattr(self, "_persistent_channel", None)
if channel is not None:
try:
channel.close()
except (OSError, subprocess.SubprocessError):
logger.warning(
"Could not stop Pyxis worker %s", self.name, exc_info=True
)
if os.environ.get("SLURM_JOB_ID", "").strip():
try:
subprocess.run(
Expand Down
Loading
Loading