Skip to content

Fix/prod log findings - #101

Merged
cypress-exe merged 12 commits into
mainfrom
fix/prod-log-findings
Sep 27, 2026
Merged

cypress-exe merged 12 commits into
mainfrom
fix/prod-log-findings

Conversation

@cypress-exe

Copy link
Copy Markdown
Owner

Description

Fixes issues seen in prod logs and just general prod usage.

Type of Change

  • Bug fix (non-breaking change which fixes an issue)
  • New feature (non-breaking change which adds functionality)
  • Breaking change (fix or feature that would cause existing functionality to not work as expected)
  • Documentation update
  • Performance / refactor (no functional change)

How Has This Been Tested?

All tests have been run and pass.

cypress-exe and others added 12 commits September 27, 2026 14:06
…rrors

on_ready fires again whenever a shard re-identifies, and each call ran
start_scheduler(). AsyncIOScheduler.shutdown() is deferred to the next
loop iteration, so the restart raised SchedulerAlreadyRunningError from
configure() and the queued shutdown then stopped the scheduler. Every
second on_ready left it dead; prod lost ~7 of 12 days of birthdays,
daily maintenance, DB maintenance and log rotation to this.

- start_scheduler() now leaves an already-running scheduler alone.
- Add a bot-level on_error so exceptions in event listeners reach the
  log file instead of nextcord's stderr-only default (which hid this
  bug).

Heavily edited by cypress.exe. Originially written by Claude Opus 5.5.
Every delete log sent a new timeout=None ShowMoreButton, which nextcord
stored against the message id forever. Prod's view store grew
boundlessly over a few days alongside an doubling of RSS.

The instance registered in view_manager.init_views already handles
clicks on every log message (the callback reads its toggle state from
the message), so construct the view with prevent_update=False and
nextcord no longer stores the per-message copies on send or on edit.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
✓ Human reviewed and edited by cypress.exe
The profanity DM decided "timed out vs. strike" from
message.author.communication_disabled_until. That object is a snapshot
that doesn't see the timeout just applied, so a strike that triggered a
timeout was treated as a plain strike and the DM looked up the strike
row the timeout had already deleted (KeyError; 6 in prod, each a lost
DM). A member with an old, expired timeout was also wrongly reported as
timed out.

- Add grant_and_punish_strike_with_result(), returning
  (action_successful, timed_out). grant_and_punish_strike() wraps it
  with its old bool contract.
- The profanity path uses the reported timed_out for the DM and admin
  log.
- The DM's strike lookup tolerates a row cleared by a concurrent
  message's timeout.

Human review note: I'm not sure if this is entirely true, since this
control flow has worked in the past. However, I agree that this is a
good change, and there was data in prod logs that supports this having
been an issue — at least to some degree.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
✓ Human reviewed and edited by cypress.exe
…vable

A fresh audit-log entry can carry entry.user = None (seen in prod logs).
log_role_change already handled that for its dedup key but then
dereferenced entry.user.mention for the description, raising
AttributeError and dropping the log. Fall back to "Someone", as the
nickname as timeout logs already do.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
✓ Human reviewed and edited by cypress.exe
After sending the changed embeds as replies, trigger_edit_log rewrote
the "Embed(s)" field with one link per embed and no length cap. Edits
touching enough embeds (or embeds with long titles) overflowed Discord's
1024-char field limit; the edit was rejected with a 400 and the log
stayed on "Please Wait..." forever.

fill_embed_links_field() lists as many links as fit and summarizes the
rest as "…and N more"; those embeds are still posted as replies to the
log.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
✓ Human reviewed and edited by cypress.exe
… History

The edit log's content/embed follow-ups and the delete log's file
uploads are sent as replies to the log message. Discord rejects replies
when the bot lacks Read Message History (403, code 160002), so those
logs failed part-way (edit logs stuck on "Please Wait...").

reply_reference() returns the log message only when the permission is
present; otherwise the follow-up goes out as a plain message. The
overflow summary from the previous commit now says "posted below" since
the follow-ups may not be replies.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
✓ Human reviewed and edited by cypress.exe
on_raw_message_edit fetched the channel, the edited message, and (for
non-Member authors) the member for every edit in every guild. Prod
logged >16k message 404s, >26k member 404s and ~400 429 retries from
this, with one webhook-heavy guild alone producing >6k of the member
404s.

- Return before any fetch unless logging, profanity moderation, spam
  moderation or leveling is on for the guild. Those are the only
  consumers of the work below (the message cache is read only by spam
  and leveling).
- Don't resolve webhook authors as members; they never are, so the fetch
  only ever 404s. (The profanity check already skips non-members.)
- Reuse the gate's Server object for the profanity check.

Human review note: This is a good change, since it also reduces the
number of bad requests InfiniBot makes to Discord, which I'm sure
Discord didn't really appreciate.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
✓ Human reviewed and edited by cypress.exe
psutil.cpu_percent() without an interval measures since its previous
call. The scheduler called it every 25 guilds (milliseconds apart) and
again for the progress/summary logs, which reset the baseline, so
readings came out as 0/50/66.7/100. Prod logged 10,546 "CPU 100.0%"
critical throttles, i.e. ~20-25s of pointless sleep per run, with
throttling that couldn't respond to real load.

CpuSampler takes a reading only once at least a second has passed since
the last one, and the throttle acts only on fresh readings. The progress
and summary logs report the last reading instead of calling psutil
(which reset the baseline).

Co-Authored-By: Claude Opus 5.5 <[email protected]>
✓ Human reviewed and edited by cypress.exe
send_error_message_to_server_owner suppressed repeats for only 30
minutes, so a misconfiguration that recurs (a missing join/leave
channel, no View Audit Log permission) DMed the owner every time the
event came back more than 30 minutes later. Prod logs revealed that this
happened several times.

Raise the cooldown to 24 hours per distinct warning. The owner is still
told about every distinct problem, and reminded daily while it persists.

Also fixed a duplicate import.
I've kept forgetting to change this, so I've had to mask these log lines
whenever I send a prod log to AI for the purposes of bug hunting. This
change protects privacy, and I'm no longer concerned about the profanity
detection system misclassifying words (it seems robust enough now).
The member-removal log rendered its description as f"{user} kicked
{member}.", which printed "None kicked X" when the audit-log actor was
unresolvable (e.g. a deleted account). The raw usernames also weren't
markdown-escaped, so names with underscores or asterisks got mangled.

- Known actor: "@mod kicked **name**." (actor as a mention, like the
other
  logs). Unknown actor: "**name** was kicked." instead of "None kicked".
- The removed user is shown by escaped name rather than a mention, since
  mentions of users no longer in the server can render as @unknown-user.
- Add the removed user's avatar as the embed author and their ID in the
  footer, matching the delete log and profanity admin embeds.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
✓ Human reviewed and edited by cypress.exe
…rmatting fixes

Updated log embeds to use most up-to-date timestamp, since logs go out a
while after the event.

Also fixed some formatting in action_logging.py, since it was a little
too long per line.

Changes initially drafted by Claude Opus 5.5. I implemented this commit
myself, but I copied some of the AI's docstrings.
@cypress-exe
cypress-exe merged commit ca4a08b into main Sep 27, 2026
8 checks passed
@cypress-exe
cypress-exe deleted the fix/prod-log-findings branch September 27, 2026 20:53
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant