Fix/prod log findings - #101
Merged
Merged
Conversation
…rrors on_ready fires again whenever a shard re-identifies, and each call ran start_scheduler(). AsyncIOScheduler.shutdown() is deferred to the next loop iteration, so the restart raised SchedulerAlreadyRunningError from configure() and the queued shutdown then stopped the scheduler. Every second on_ready left it dead; prod lost ~7 of 12 days of birthdays, daily maintenance, DB maintenance and log rotation to this. - start_scheduler() now leaves an already-running scheduler alone. - Add a bot-level on_error so exceptions in event listeners reach the log file instead of nextcord's stderr-only default (which hid this bug). Heavily edited by cypress.exe. Originially written by Claude Opus 5.5.
Every delete log sent a new timeout=None ShowMoreButton, which nextcord stored against the message id forever. Prod's view store grew boundlessly over a few days alongside an doubling of RSS. The instance registered in view_manager.init_views already handles clicks on every log message (the callback reads its toggle state from the message), so construct the view with prevent_update=False and nextcord no longer stores the per-message copies on send or on edit. Co-Authored-By: Claude Opus 5.5 <[email protected]> ✓ Human reviewed and edited by cypress.exe
The profanity DM decided "timed out vs. strike" from message.author.communication_disabled_until. That object is a snapshot that doesn't see the timeout just applied, so a strike that triggered a timeout was treated as a plain strike and the DM looked up the strike row the timeout had already deleted (KeyError; 6 in prod, each a lost DM). A member with an old, expired timeout was also wrongly reported as timed out. - Add grant_and_punish_strike_with_result(), returning (action_successful, timed_out). grant_and_punish_strike() wraps it with its old bool contract. - The profanity path uses the reported timed_out for the DM and admin log. - The DM's strike lookup tolerates a row cleared by a concurrent message's timeout. Human review note: I'm not sure if this is entirely true, since this control flow has worked in the past. However, I agree that this is a good change, and there was data in prod logs that supports this having been an issue — at least to some degree. Co-Authored-By: Claude Opus 5.5 <[email protected]> ✓ Human reviewed and edited by cypress.exe
…vable A fresh audit-log entry can carry entry.user = None (seen in prod logs). log_role_change already handled that for its dedup key but then dereferenced entry.user.mention for the description, raising AttributeError and dropping the log. Fall back to "Someone", as the nickname as timeout logs already do. Co-Authored-By: Claude Opus 5.5 <[email protected]> ✓ Human reviewed and edited by cypress.exe
After sending the changed embeds as replies, trigger_edit_log rewrote the "Embed(s)" field with one link per embed and no length cap. Edits touching enough embeds (or embeds with long titles) overflowed Discord's 1024-char field limit; the edit was rejected with a 400 and the log stayed on "Please Wait..." forever. fill_embed_links_field() lists as many links as fit and summarizes the rest as "…and N more"; those embeds are still posted as replies to the log. Co-Authored-By: Claude Opus 5.5 <[email protected]> ✓ Human reviewed and edited by cypress.exe
… History The edit log's content/embed follow-ups and the delete log's file uploads are sent as replies to the log message. Discord rejects replies when the bot lacks Read Message History (403, code 160002), so those logs failed part-way (edit logs stuck on "Please Wait..."). reply_reference() returns the log message only when the permission is present; otherwise the follow-up goes out as a plain message. The overflow summary from the previous commit now says "posted below" since the follow-ups may not be replies. Co-Authored-By: Claude Opus 5.5 <[email protected]> ✓ Human reviewed and edited by cypress.exe
on_raw_message_edit fetched the channel, the edited message, and (for non-Member authors) the member for every edit in every guild. Prod logged >16k message 404s, >26k member 404s and ~400 429 retries from this, with one webhook-heavy guild alone producing >6k of the member 404s. - Return before any fetch unless logging, profanity moderation, spam moderation or leveling is on for the guild. Those are the only consumers of the work below (the message cache is read only by spam and leveling). - Don't resolve webhook authors as members; they never are, so the fetch only ever 404s. (The profanity check already skips non-members.) - Reuse the gate's Server object for the profanity check. Human review note: This is a good change, since it also reduces the number of bad requests InfiniBot makes to Discord, which I'm sure Discord didn't really appreciate. Co-Authored-By: Claude Opus 5.5 <[email protected]> ✓ Human reviewed and edited by cypress.exe
psutil.cpu_percent() without an interval measures since its previous call. The scheduler called it every 25 guilds (milliseconds apart) and again for the progress/summary logs, which reset the baseline, so readings came out as 0/50/66.7/100. Prod logged 10,546 "CPU 100.0%" critical throttles, i.e. ~20-25s of pointless sleep per run, with throttling that couldn't respond to real load. CpuSampler takes a reading only once at least a second has passed since the last one, and the throttle acts only on fresh readings. The progress and summary logs report the last reading instead of calling psutil (which reset the baseline). Co-Authored-By: Claude Opus 5.5 <[email protected]> ✓ Human reviewed and edited by cypress.exe
send_error_message_to_server_owner suppressed repeats for only 30 minutes, so a misconfiguration that recurs (a missing join/leave channel, no View Audit Log permission) DMed the owner every time the event came back more than 30 minutes later. Prod logs revealed that this happened several times. Raise the cooldown to 24 hours per distinct warning. The owner is still told about every distinct problem, and reminded daily while it persists. Also fixed a duplicate import.
I've kept forgetting to change this, so I've had to mask these log lines whenever I send a prod log to AI for the purposes of bug hunting. This change protects privacy, and I'm no longer concerned about the profanity detection system misclassifying words (it seems robust enough now).
The member-removal log rendered its description as f"{user} kicked
{member}.", which printed "None kicked X" when the audit-log actor was
unresolvable (e.g. a deleted account). The raw usernames also weren't
markdown-escaped, so names with underscores or asterisks got mangled.
- Known actor: "@mod kicked **name**." (actor as a mention, like the
other
logs). Unknown actor: "**name** was kicked." instead of "None kicked".
- The removed user is shown by escaped name rather than a mention, since
mentions of users no longer in the server can render as @unknown-user.
- Add the removed user's avatar as the embed author and their ID in the
footer, matching the delete log and profanity admin embeds.
Co-Authored-By: Claude Opus 5.5 <[email protected]>
✓ Human reviewed and edited by cypress.exe
…rmatting fixes Updated log embeds to use most up-to-date timestamp, since logs go out a while after the event. Also fixed some formatting in action_logging.py, since it was a little too long per line. Changes initially drafted by Claude Opus 5.5. I implemented this commit myself, but I copied some of the AI's docstrings.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Description
Fixes issues seen in prod logs and just general prod usage.
Type of Change
How Has This Been Tested?
All tests have been run and pass.