Skip to content

[Klaud Cold] Add live-logs skill to stream a sweep job's per-node logs in the browser / [Klaud Cold] 新增 live-logs 技能:在浏览器中实时查看扫描任务的逐节点日志 - #3524

Merged
functionstackx merged 5 commits into
mainfrom
klaud/live-logs-skill
Sep 28, 2026
Merged

functionstackx merged 5 commits into
mainfrom
klaud/live-logs-skill

Conversation

@functionstackx

Copy link
Copy Markdown
Collaborator

Summary

Adds a live-logs repo skill (.agents/skills/live-logs/, with a .claude/skills/ symlink). Give it a PR number, PR URL, workflow run URL or job URL, and it opens a live, per-node log viewer in the browser for each running multi-node job.

python3 .agents/skills/live-logs/scripts/live_logs.py 3523
python3 .agents/skills/live-logs/scripts/live_logs.py https://github.com/SemiAnalysisAI/InferenceX/actions/runs/<run>/job/<job>
python3 .agents/skills/live-logs/scripts/live_logs.py --stop
  • Resolution: GitHub job → runner name (e.g. b300-dsxe_01) → Slurm job of the same name (squeue -n, with sacct as a fallback) → scontrol WorkDir → <WorkDir>/outputs/<job>/logs. You get one viewer per Slurm job, e.g. benchmark and eval.
  • Streaming: ssh tail -F feeds a local server that pushes lines to the page over Server-Sent Events. It binds to 127.0.0.1 only and uses just the Python standard library.
    • New files, such as aiperf output when the benchmark starts, are found every 30 s.
    • After an SSH drop, the files are re-read so no lines are lost.
  • Page layout:
    • Rows for workers (prefill/decode/agg), orchestration (srtctl, Dynamo frontend, Mooncake master, etcd) and benchmark (aiperf, benchmark.log).
    • A phase tracker per pane: weights → KV sized → EFA up → Mooncake registration → autotune → healthy.
    • Error highlighting and counts.
    • Per-node free host memory and CPU load in the header.
  • Controls:
    • Regex filter, "errors only", and a Mooncake-metrics noise filter.
    • show toggle: last 4000 lines (the default) or the entire log.
    • Resizable panes: drag the gutters, double-click a gutter to even out its panes, maximize a pane; sizes are saved.
    • Jump to start/end and copy buttons.
    • Download all logs: a zip of every full .out/.log/.txt under logs/, with no tachometer metrics.
  • No infra in the repo: cluster login hosts come from ~/.config/infx-live-logs/clusters.json or INFX_LIVE_LOGS_HOST_<CLUSTER>, following the debug-runs canvas rule.

Test plan

中文

新增 live-logs 仓库技能(.agents/skills/live-logs/,并在 .claude/skills/ 下建立符号链接)。给定 PR 编号、PR 链接、workflow run 链接或 job 链接,即可为每个正在运行的多节点任务在浏览器中打开实时的逐节点日志查看器。

  • 解析流程: GitHub job → runner 名称 → 同名 Slurm 作业 → WorkDir → outputs/<job>/logs。
  • 实时推送: 通过 ssh tail -F 读取日志,本地服务以 SSE 推送到页面。仅监听 127.0.0.1,只依赖 Python 标准库。
  • 页面: 按 worker、编排服务、基准测试分行显示;每个窗格显示当前阶段和错误数;顶部显示各节点空闲内存与 CPU 负载。
  • 控件:
    • 正则过滤、仅看错误、Mooncake 指标噪声过滤。
    • 显示最近 4000 行(默认)或完整日志。
    • 窗格可拖动调整大小,也可最大化。
    • 「下载全部日志」:打包 logs/ 下所有完整的 .out/.log/.txt 为 zip,不含 tachometer 指标。
  • 不在仓库中存放集群登录地址: 登录主机从本地配置或环境变量读取。

AI disclosure: authored with Claude Code (Claude Opus 5.5).

🤖 Generated with Claude Code

@claude claude Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Findings marked 🟡 are optional suggestions and need no follow-up push.

Comment on lines +88 to +101
cmd = SSH + [args.host, f"cd {args.logdir} && exec tail -n {args.tail_from} -F " + " ".join(rels)]
st.status["ssh"] = "connected"
st.publish({"t": "status", "s": st.status})
p = subprocess.Popen(cmd, stdout=subprocess.PIPE, stderr=subprocess.DEVNULL, text=True, errors="replace", bufsize=1)
cur = None
for line in p.stdout:
line = ANSI.sub("", line.rstrip("\n"))
m = re.match(r"^==> (.+) <==$", line)
if m:
cur = by_base.get(m.group(1), cur)
continue
if cur is None or not line:
continue
st.buffers[cur].append(line)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔴 When tail_batch is started for exactly one log file, no lines ever reach the viewer because GNU tail only prints the ==> file <== header "with more than one FILE" (per tail --help). cur (server.py:92) stays None forever, so if cur is None or not line: continue (server.py:99) drops every line silently. This hits the exact case the PR advertises: discover_loop finds new files every 30s (server.py:113-127), and a single newly-appeared file (e.g. aiperf.log alone, or a single-file job) is passed as a one-element rels list to tail_batch, permanently blackholing that pane with no error shown. Fix: pass -v/--verbose to the ssh tail -F command at server.py:88 so headers are always emitted regardless of file count, or seed cur when len(rels) == 1.

Why this was flagged

Trigger: discover_loop (server.py:113-127) spawns tail_batch with a single-file rels list whenever exactly one new matching file appears (routine, since aiperf/benchmark logs often show up alone) or when a job's log dir only ever contains one matching file. GNU tail (server.py:88, tail -n {tail_from} -F + one filename) omits the ==> name <== header for a single file. cur (server.py:92) is initialized to None and the header-matching branch (server.py:95-98) never fires, so cur never gets set. Every subsequent line hits if cur is None or not line: continue (server.py:99) and is discarded — the pane for that file stays empty forever with no error surfaced to the user, unlike the base (no such feature existed) or a correctly working multi-file case.

Verification: normal. server.py:88 runs tail -n {tail_from} -F + " ".join(rels) with no -v/--verbose flag. GNU tail prints the ==> file <== header only "with more than one FILE" (default), never for a single file. cur is initialized to None (server.py:92) and is set only inside the header-match branch (server.py:95-98). Line 99 if cur is None or not line: continue therefore drops EVERY line…

if self.path == "/":
with open(page_path) as f:
page = f.read()
body = page.replace("__TITLE__", args.title).replace("__JOB__", json.dumps(str(args.job))).encode()

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔴 Anyone who opens the viewer for a job whose GitHub Actions job name contains HTML gets arbitrary JS run in their local browser tab, unlike JOB on the same line which is safely JSON-escaped. page.replace("__TITLE__", args.title) splices args.title (from live_logs.py's j['name'], the raw GH Actions job display name, live_logs.py:169) straight into <h1>__TITLE__</h1> with no HTML escaping. A job/matrix name containing <, > or " (e.g. from a PR-supplied benchmark config value that ends up in the matrix-derived job name) breaks out of the h1 and executes, and since the page has no CSP it can fetch() the live /events SSE stream and exfiltrate cluster log contents to an external host. …

Why this was flagged

…Fix: HTML-escape args.title (e.g. html.escape) before substitution, matching the json.dumps escaping already used for JOB.

Trigger: live_logs.py:167 builds title from j['name'] (raw GitHub Actions job name, live_logs.py:169) with no sanitization, then server.py main() stores it as args.title (server.py:265). do_GET's / handler (server.py:188) does page.replace("__TITLE__", args.title) and writes the result as the HTML response body with no escaping, while the adjacent JOB substitution on the same line is escaped via json.dumps. If the job name contains <script> or a " breaking out of the h1/attribute context, that markup/script is served verbatim to whoever opens the viewer URL in their browser. The page sets no CSP, so injected script can call fetch() to reach an external origin. This is new: a malicious or unusual job/matrix name has no other place to be neutralized before reaching the browser.

Verification: normal; security-relevant reflected XSS introduced by this new-file change. index.html:54 places the token in an HTML element context: <h1>__TITLE__</h1>. server.py:188 splices args.title into it with no escaping: body = page.replace("__TITLE__", args.title).replace("__JOB__", json.dumps(str(args.job))).encode() — the adjacent JOB substitution is passed through json.dumps, but… | nit.…

"""Return [(run_id, job)] for in-progress, non-setup jobs of the target."""
m = re.search(r"/actions/runs/(\d+)/job/(\d+)", target)
if m:
return [(m.group(1), gh(f"repos/{REPO}/actions/jobs/{m.group(2)}"))]

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🔴 Users who paste a job URL for a queued, setup, or already-finished job get an unhandled crash instead of the graceful message other inputs get. resolve_jobs() at live_logs.py:42-44 returns the job unconditionally for a .../runs/<r>/job/<j> URL, skipping the status == "in_progress" and j.get("runner_name") guard applied to the run-URL/PR path (live_logs.py:62). In main() (line 156) runner = j["runner_name"] is then passed straight to runner.rsplit("_", 1) (line 157) with no None check, so a job whose runner_name is null (not yet scheduled, or a completed/setup job) raises AttributeError and traces back instead of the friendly "no in-progress jobs..." exit used elsewhere. …

Why this was flagged

…Fix: apply the same in_progress/runner_name check to the job-URL branch (or guard runner before use in main), covering both the direct job-URL path and any other caller of resolve_jobs.

Trigger: user runs live_logs.py <job URL> for a job URL copied from a workflow that is queued, still on "setup", or has already completed (SKILL.md explicitly documents job URLs as supported input). resolve_jobs() (live_logs.py:42-44) returns that job's raw API record with no status/runner_name filter, unlike the run-URL path (line 62) which requires status == "in_progress" and j.get("runner_name"). main() (line 156-157) does runner = j["runner_name"]; cluster = runner.rsplit("_", 1)[0] with no None check, so when runner_name is null the script raises AttributeError and exits with a Python traceback instead of the existing clean error path ("no in-progress jobs with a runner found", line 153) that other inputs get.

Verification: normal. The job-URL branch of resolve_jobs (live_logs.py:42-44) returns gh(f"repos/{REPO}/actions/jobs/{m.group(2)}") unconditionally, without the status == "in_progress" and j.get("runner_name") filter that the run/PR path applies at line 62. main() then does runner = j["runner_name"] (line 156) and cluster = runner.rsplit("_", 1)[0] (line 157) with no None check. For a queued…

Comment on lines +158 to +174
p = subprocess.Popen(SSH + [args.host, pick],
stdout=subprocess.PIPE, stderr=subprocess.DEVNULL)
fd, out = tempfile.mkstemp(prefix=name + "-", suffix=".zip")
os.close(fd)
with tarfile.open(fileobj=p.stdout, mode="r|gz") as tar, zipfile.ZipFile(out, "w", zipfile.ZIP_DEFLATED) as zf:
for m in tar:
if not m.isfile():
continue
src = tar.extractfile(m)
with zf.open(f"{name}-logs/{m.name}", "w") as dst:
while True:
chunk = src.read(1 << 20)
if not chunk:
break
dst.write(chunk)
p.wait()
return out

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 (optional) On the new /download.zip endpoint, an operator who triggers 'download all logs' during a flaky SSH connection leaks a temp file and a zombie ssh process every time. build_zip() creates the temp zip via tempfile.mkstemp() at server.py:160 before streaming, but if the remote tar/gzip stream breaks (e.g. ssh drop mid-transfer) tarfile.open()/tar iteration raises inside the with block at server.py:162-172 before p.wait() (173) or return out (174) run. The caller at server.py:195-200 only catches TarError/OSError/SubprocessError to send a 502; it never learns out's path, so the temp zip is never removed, and p is never waited on. …

Why this was flagged

…Fix: wrap the ssh Popen and tar/zip work in try/finally inside build_zip so the temp file is deleted and the subprocess is waited/killed on any exception, not just cleaned up on the success path.

Trigger: clicking 'download all logs' (index.html dl button -> GET /download.zip) while the cluster SSH session used by build_zip drops or the remote tar command fails (per PR description, this streams up to hundreds of MB, e.g. a 243MB prefill log, so mid-transfer drops are realistic). build_zip (server.py:153-174) creates the temp file via tempfile.mkstemp at line 160 then raises inside the tarfile/zipfile with block (162-172) on a truncated/invalid gzip stream. do_GET's except at server.py:198-200 only sends a 502 and returns; it has no reference to out to delete it, and p.wait() (173) never runs so the ssh child is left unreaped. Each failed download leaves an orphaned slurm--*.zip in the OS temp dir and an unreaped subprocess, with no cleanup path anywhere else in the code.

Verification: nit. Real resource leak on the new /download.zip error path, but low impact in a local 127.0.0.1 single-user tool. In build_zip (server.py:153-174), the temp zip is created at line 160 (fd, out = tempfile.mkstemp(...); os.close(fd)) and the ssh subprocess is spawned at 158-159. The tar/gzip work runs inside the with block at 162-172. On an SSH drop or truncated/invalid gzip stream (per PR…

Comment on lines +155 to +162
for rid, j in jobs:
runner = j["runner_name"]
cluster = runner.rsplit("_", 1)[0]
host = cluster_host(cluster)
if not host:
missing.add(cluster)
continue
job, workdir = slurm_job_for_runner(host, runner)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 (optional) Users watching several jobs from one run (e.g. benchmark + eval) lose viewers for every healthy job when only one cluster's ssh hangs, a failure the base doesn't have since this script is new. slurm_job_for_runner() calls remote() (live_logs.py:78-80) with subprocess.run(timeout=90); if the remote squeue/sacct probe hangs past that, subprocess.TimeoutExpired propagates uncaught out of the call at live_logs.py:162 inside main()'s for rid, j in jobs: loop, killing the whole invocation before later jobs open their viewer. Fix: catch subprocess.SubprocessError/OSError per job in that loop (as server.py's discover_loop/status_loop already do for the same remote() pattern) so one stuck cluster doesn't abort the rest of jobs.

Why this was flagged

Trigger: a multi-job target (PR/run URL resolving to >1 in-progress job, e.g. benchmark + eval) where one job's cluster host is unreachable or slow enough that the remote squeue/sacct pipeline in slurm_job_for_runner (live_logs.py:83-93) doesn't return within subprocess.run's timeout=90 at live_logs.py:79. The changed code has no try/except around the call at live_logs.py:162, unlike server.py's remote() callers (discover_loop, status_loop) which explicitly catch subprocess.SubprocessError/OSError/ValueError around the same remote() pattern. The uncaught subprocess.TimeoutExpired unwinds main()'s for-loop entirely, so viewers for jobs processed after the stuck one in jobs never start and no URL is printed for them, whereas per-job isolation would let the healthy jobs still get a viewer.

Verification: normal. New file, so the base has no equivalent. remote() (live_logs.py:78-79) runs subprocess.run(SSH + [host, script], ..., timeout=90, check=False); check=False suppresses CalledProcessError but NOT subprocess.TimeoutExpired, which is raised whenever the 90s timeout fires. slurm_job_for_runner() (live_logs.py:83-93) wraps two remote() calls with no try/except, and main()'s loop `for rid, j…

Comment on lines +113 to +127
def discover_loop(args, st):
while True:
try:
found = remote(args.host, f"cd {args.logdir} 2>/dev/null && find . -maxdepth 5 -type f {FIND_EXPR} | sed 's#^./##' | sort")
new = [r for r in found if r not in st.buffers]
if new:
with st.lock:
for r in new:
st.buffers[r] = collections.deque(maxlen=args.max_lines)
st.files.append([r, label_of(r), group_of(r)])
st.publish({"t": "files", "f": st.files})
threading.Thread(target=tail_batch, args=(args, st, new), daemon=True).start()
except (OSError, subprocess.SubprocessError, ValueError) as e: # flaky network: keep going
st.status["ssh"] = f"discovery error: {e}"
time.sleep(30)

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🟡 (optional) Operators running a long multi-concurrency sweep get an unbounded, ever-growing pile of persistent SSH sessions and threads for the viewer's whole lifetime. discover_loop's 30s poll (server.py:113-127) starts a brand-new daemon thread per newly-seen file (line 124), and each such thread runs tail_batch's while True reconnect loop (server.py:87-110) that never exits and just re-ssh's every 5s on drop (line 110). As a sweep progresses through many conc_N steps, each new aiperf.log/benchmark.log spawns its own separate forever-running ssh tail session against the same login host, with nothing merging new files into an existing session or capping concurrent sessions per host. …

Why this was flagged

…Fix: when new files are discovered, fold them into the running tail_batch session(s) for that host (restart with the union of files) or cap/reuse concurrent tail sessions, instead of opening one permanent ssh connection per discovery batch.

Trigger: a long sweep with several concurrency steps causes discover_loop (server.py:116-127) to see new conc_N/aiperf.log or benchmark.log files at different 30s polls; each poll's new list (line 117) gets its own threading.Thread(target=tail_batch, ...) (line 124). tail_batch (line 82-110) has an infinite while True (line 87) that only reconnects on drop (time.sleep(5) at line 110) and never terminates, so each such thread and its ssh subprocess (line 91) live for the entire viewer process lifetime. Nothing consolidates these into the initial batch's session or bounds total concurrent sessions per host. Consequence: dozens of long-lived SSH connections accumulate from a single viewer against one cluster login node, which can exhaust sshd's MaxSessions/MaxStartups, stalling further discover_loop/status_loop remote() calls and…

Verification: nit. The mechanism the candidate describes is real. discover_loop (server.py:113-127) runs a while True poll every 30s (line 127). On each poll it computes new = [r for r in found if r not in st.buffers] (line 117), and if any new files appear it spawns threading.Thread(target=tail_batch, args=(args, st, new), daemon=True).start() (line 124). tail_batch (server.py:82-110) has an infinite…

functionstackx and others added 5 commits September 27, 2026 21:40
…to the browser

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
…file separately

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
…ck, srt-single log layout

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
…r is a checkout)

Co-Authored-By: Claude Opus 5.5 (1M context) <[email protected]>
@functionstackx
functionstackx merged commit 73c32b0 into main Sep 28, 2026
3 checks passed
@functionstackx
functionstackx deleted the klaud/live-logs-skill branch September 28, 2026 02:07
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

Status: Done

Development

Successfully merging this pull request may close these issues.

1 participant