Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
20 commits
Select commit Hold shift + click to select a range
51b79c3
fix(import): scope known-series recovery and preserve safety evidence
Sep 17, 2026
fd158e5
fix(scheduler): bound exclusive admission and retry busy maintenance
Sep 17, 2026
5454bc4
fix(airdcpp): tolerate transient queue telemetry overshoot
Sep 17, 2026
f6d53a2
fix(database): drain maintenance safely and bound health probes
Sep 17, 2026
97c32e7
fix(metadata): resume bounded sweeps and respect provider cooldowns
Sep 17, 2026
2b84de1
fix(indexers): share throttle cooldowns and serialize provider requests
Sep 17, 2026
0a281d8
fix(backup): release request transactions before maintenance
Sep 17, 2026
22023bb
docs: record background maintenance and provider retry contracts
Sep 17, 2026
267375f
fix(import): safely revalidate recovery sources and reviewed identities
Sep 17, 2026
c598869
fix(import): reconcile provisional targets before conflict grouping
Sep 17, 2026
4f0ad87
fix: preserve hyphenated issue suffixes across metadata and imports
Sep 17, 2026
736f8ad
fix(import): preserve control requests during worker handoff
Sep 17, 2026
c4131dd
fix(import): reconcile local volume and known annual evidence
Sep 17, 2026
3eca920
feat(import): recognize source volume folders and record Mylar proven…
Sep 17, 2026
42ba085
fix(import): recover exact catalog targets and guard mixed folders
Sep 18, 2026
51a4c77
fix(import): recover misplaced referenced files using catalog targets
Sep 18, 2026
628adb1
chore(security): scope approved zlib exception to DHI refresh
Sep 18, 2026
6b26e93
test(security): enforce the approved zlib package exception
Sep 18, 2026
ff6062e
fix(metadata): finish restore sweeps and stop keyless continuations
Sep 18, 2026
7e959bd
test: tolerate CI startup latency in job progress regression
Sep 18, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
6 changes: 6 additions & 0 deletions .grype.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -542,6 +542,12 @@ ignore:
name: zlib1g
version: 1:1.3.dfsg+really1.3.1-1+dhi3
type: deb
# Re-reviewed 2026-09-17 for the DHI package refresh; same exposure and deadline.
- vulnerability: CVE-2026-85091
package:
name: zlib1g
version: 1:1.3.dfsg+really1.3.1-1+dhi4
type: deb
- vulnerability: CVE-2026-85091
package:
name: zlib1g-dev
Expand Down
22 changes: 22 additions & 0 deletions docs/development/ARCHITECTURE_OVERVIEW.md
Original file line number Diff line number Diff line change
Expand Up @@ -553,6 +553,28 @@ task registry. The scheduler runs recurring jobs for search, downloads,
metadata refresh, health checks, imports, backups, dashboard refresh, blocklist
cleanup, and related operational work.

Exclusive scheduled maintenance reserves admission in memory, outside database
stats writes. It waits at most five seconds for other tasks, then releases its
reservation and schedules a retry after 60 seconds if work is still active.
Cancellation also releases the reservation.

Nightly issue and metadata sweeps checkpoint their last completed series and
initial upper bound in `SystemConfig`. Each batch handles at most 25 series and
checks a two-minute budget between series; an individual series has a separate
15-minute timeout. Pending batches resume through hidden continuations, including
after restart. Provider throttling pauses the sweep at its saved position instead
of repeatedly failing every remaining series. Series metadata writes are committed
before subsequent cover or issue-provider waits.
Removing the ComicVine key stops the active sweep and clears its continuation.
Post-restore aftercare keeps its recovery marker while continuation batches remain
active; it observes completion without retaining a database transaction between checks.
Cancellation leaves the marker and sweep checkpoint available for the next startup.

ComicVine and Newznab clients share process-local account cooldowns, so creating a
new client does not bypass a throttle response. Newznab also serializes request
pacing per provider/account; unrelated accounts remain independent. Sweep retry
deadlines are durable, while the shared client cooldown registry itself is not.

The event bus is intentionally small and in-process. It supports domain side
effects such as:

Expand Down
12 changes: 11 additions & 1 deletion docs/development/DATABASE_STANDARDS.md
Original file line number Diff line number Diff line change
Expand Up @@ -544,13 +544,23 @@ not suppress failures in those checks.

- Database maintenance windows coordinate app traffic through the shared
maintenance gate.
- Gate-aware sessions pause before database-touching operations.
- Gate-aware sessions pause new transactions while existing transactions have
up to five seconds to finish. If they cannot drain, maintenance yields with
a retryable busy result rather than holding the gate indefinitely.
- Manual backup, restore, and optimization requests release their own request
transaction before entering maintenance. Cancellation does not reopen the
gate until an already-running filesystem or SQLite worker has stopped.
- An exclusive nightly task runs SQLite `REINDEX` and
`PRAGMA optimize=0x10002` at 04:30, then verifies the database with
`PRAGMA quick_check`. The all-tables mask is intentional because the
maintenance connection has no prior query history.
- Full SQLite `VACUUM` compaction remains an explicit operator action because
it rewrites the database and can require substantial temporary disk space.
- Routine integrity health probes use a separate read-only SQLite connection
with a five-second execution budget. An interrupted or busy probe reports
incomplete verification, not confirmed database corruption.
- Search-log retention deletes at most 500 rows per transaction and releases
the writer between batches.

**Required standard**

Expand Down
127 changes: 120 additions & 7 deletions docs/development/IMPORT_REVIEW_RECOVERY.md
Original file line number Diff line number Diff line change
Expand Up @@ -62,6 +62,66 @@ The shared checks cover fresh Mylar discovery and saved-review reconciliation.
They do not modify Mylar, rename source files, or change completed-import
recovery and ownership rules. No metadata-provider requests are introduced.

Before copy/conflict grouping, matching reconciles automatic provisional issue
targets against exact identities discovered in the same import-series row.
This runs after all file pages, so an untagged CBR and a later tagged CBZ cannot
escape copy review merely because their identities were learned at different
times. Only an unambiguous target with the same exact issue designation and
issue type is reused; annuals, lettered issues, conflicting metadata, manual
decisions, skips, and safety blocks are not overridden. Reads are paged, source
signatures and files are unchanged, and no provider requests are added. The
shared finalization applies to Mylar and folder imports in either handling mode.

Provisional targets use each file's inspected issue type, not the type of the
first same-number file in a matching page. This preserves numbered collected
volumes when deferred ComicInfo changes an initial ordinary-issue classification
and keeps regular issues separate from same-number annuals.

An unrecorded annual can join an already identified annual series in the same
source folder when its filename and ComicInfo agree on the annual title and
exact issue designation, its publication year matches, and exactly one trusted
series is eligible. Conflicting IDs, dates, competing series identities, safety
blocks, and manual decisions remain untouched. This only changes the import
review grouping: it never moves source files, borrows a missing reference's
issue ID, or dismisses that reference. Normal issue matching and duplicate review
still run afterward. The rule applies to Mylar and folder imports, for both
copy and keep-in-place modes, without metadata-provider requests.

Exact issue designations retain letter suffixes and their hyphens, including
`13A`, `50-X`, and `50-O`. ComicVine metadata, local catalog reads, Mylar and
folder discovery, and release validation share this identity policy. A dashed
letter suffix is not a numeric range or a plain issue number; sibling letters
must not collapse into the same target. Refreshing live metadata can correct
an older zero-number fallback using its existing ComicVine issue ID without
discarding ownership. No schema migration or source-file rename is required.

## Source Volume Folders And Mylar Provenance

Automatic folder discovery recognizes `Series/v2017`, `Publisher/Series/v2017`,
`Series/v2`, and `Publisher/Series/v2`. The parent supplies a series-name hint;
`v2017` supplies a series-year hint, while `v2` is an ordinal volume, never an
issue number or a year. Filename identity with an issue designation or agreeing
local sidecar/ComicInfo evidence must corroborate the parent. Layout recognition
is not a confirmed ComicVine match. An ordinal folder without an explicit
series-year or series ID remains in the existing series-review workflow.

Step 1 remains filename-only and does not open archives or call providers.
The scan reuses its existing local metadata reads for corroboration. Weak or
contradictory evidence stays reviewable per file, while proven per-file identities
continue through existing mixed-folder recovery. A legitimate series named `V2`
is not replaced by its parent. Custom source layouts remain authoritative, and
parent sidecars are not inherited across releases. No source folder or file is
renamed, moved, or rewritten by this recognition, and it never selects a future
managed-library layout. Existing copy/in-place policies remain independent.

Mylar scans log `mylar3_source_provenance`, including optional
`mylar_info.DatabaseVersion`, series count, Story Arc/reading-list presence,
reading-list count, and supported configuration key names. Missing or malformed
version data becomes `unknown`; future numeric versions remain diagnostic data,
not admission gates. `config.ini` remains optional and no configuration values
or credentials are logged. Both paged and full snapshot readers remain read-only.
This introduces no database migration, monitoring change, or new import card.

## Import Follow-up

The Follow-up tab groups actionable work by import job rather than rendering
Expand Down Expand Up @@ -97,8 +157,10 @@ ineligible.
- **Skip unusable files** excludes empty, unsupported, and page-less files,
including those confirmed unusable by a later source recheck.
- **Allow oversized files once** retries only decompression-size blocks marked
overrideable. It does not change the global archive safety policy or approve
dangerous archive content.
overrideable, including failed files whose latest source recheck recorded an
overrideable size limit rather than a scan-time safety block. The original
recheck evidence is retained alongside the one-job exception. It does not
change the global archive safety policy or approve dangerous archive content.
- **Retry source inspection** rechecks files that were unreadable, changed, or
temporarily could not be inspected, then resumes only work that now passes.
- **Recognize already-owned issues** clears conflicts whose issue already has a
Expand All @@ -118,9 +180,14 @@ ineligible.
only files whose saved series and issue identities agree. Existing catalog
ownership, issue numbers, per-file conflicts, manual decisions, skips, and
safety blocks are checked before an actor-bound preview is issued. Ambiguous
duplicate candidates and a series containing an unpreviewed ready file are
excluded. The confirmed scope runs through normal background Step 4 source
validation and import rules. Successful files and source paths are untouched;
duplicate candidates remain excluded, including another ready file claiming
the same issue or a manually chosen keeper. Explicit file exclusions and
safety-review decisions remain protected. Eligible files are isolated into
new recovery groups with links to their original groups; an unresolved
sibling no longer blocks otherwise proven files. The actor-bound preview
authorizes only those groups, not other ready files or Story Arcs in the job.
The confirmed scope runs through normal background Step 4 source validation
and the original copy or keep-in-place rules. Successful files and source paths are untouched;
unresolved files remain in Follow-up. This is not a blanket repair of stale
IDs or a replacement for manual review when trusted evidence disagrees.
- **Recheck deferred files** checks the remaining unmatched files in a resumable
Expand All @@ -132,6 +199,28 @@ ineligible.
them again. A different file for an owned issue remains a review decision;
it is never automatically substituted for the owned copy.

The same recheck can repair already-imported, referenced comics in mixed folders,
even when their correct series or issue is not in Pullbox yet. It groups exact
title lookups against the local catalog and requires a unique issue number, type,
and agreeing per-file publication year (within one year for dated files).
Filename-only evidence without a publication year remains review-only. Trusted
embedded IDs must agree. Manual choices, safety decisions, managed artifacts,
duplicate candidates, and targets with another owned file are not overwritten.
Before changing an assignment, the worker rechecks the source path, size, and
timestamp inside an enabled reference-capable root. This also works with a
read-only root; no source bytes, filenames, or directories are changed.

Missing metadata targets are registered as partial catalogs, without creating a
series folder or running file import/conversion. The existing LibraryFile row is
retained and assigned to the verified issue; old and new ownership counters and
import Story Arc links are refreshed. Reading history remains unchanged.
Empty provisional issues left under the wrong series are kept for audit but
marked skipped rather than becoming new wanted downloads.
Each repair records its previous assignment and commits with its recovery
checkpoint. Resume preserves completed repairs, and repeated runs do not create
another file registration. This is logical library repair, not authorization to
reorganize the user's filesystem.

The deferred pass uses complete local catalogs first. Exact issue identity may
correct stale Mylar ownership only when the file's title, issue number, type,
and embedded identity agree with the target. Conflicting embedded IDs remain
Expand All @@ -151,7 +240,7 @@ are stored. Live provider progress does not rewrite the full durable recovery
snapshot; each completed catalog produces one durable checkpoint, so a worker
restart resumes after the last completed catalog without replaying it.

Recovered files run through normal Step 4 safety, current-source validation,
Previously unimported recovered files run through normal Step 4 safety, current-source validation,
ownership checks, and the original copy or keep-in-place settings. Only newly
prepared recovery groups execute, not unrelated ready files or Story Arcs.
Cancellation stops this pass without rolling back the original import or
Expand All @@ -174,12 +263,36 @@ status-only correction does not launch another import.
A completed source recheck reports a file ready only after both archive safety
and saved target identity checks pass. Missing, empty, or otherwise blocked
sources are counted as blocked even when the archive-level inspection itself
completed successfully.
completed successfully. A `source_identity_changed` result means the current
series or issue identity disagrees with the reviewed match, not that the file
was necessarily modified on disk. Its message directs the user to Follow-up;
it is neither automatically retryable nor overrideable.
Completed-import source rechecks inspect one bounded page before writing its
refreshed evidence, then commit that page before reading more archives. This
keeps slow archive I/O outside SQLite's single-writer window, bounds memory, and
leaves completed pages durable if a later source needs another attempt.

An identity-conflicted failed file remains assignable in Follow-up; archive
safety failures do not become assignable through that exception. A manual
issue assignment may supersede a disagreement between `series.json` and
`cvinfo` only when freshly inspected ComicInfo proves the assigned issue, its
series identity does not contradict the reviewed series, and issue-number
checks pass. The original folder conflict is retained as resolved evidence;
other metadata conflicts retain their specific IDs in the failure diagnostics.
Step 4 also verifies that the assigned issue belongs to the chosen local series.

Known-series, deferred, mixed-folder, and manually assigned recovery can
revalidate a device-number-only change after a container remount. The path,
inode, size, timestamp, and signature version must still agree. An archive
inspection under the approved source roots and an exact embedded issue-ID
check are required before refreshing the saved signature. The pre-inspection
and post-inspection identities must agree, and normal registration revalidates
the root and refreshed signature again. Existing approved size or single-page
exceptions remain scoped to that file; dangerous archive checks still run.
Ordinary signature validation, source files, Mylar metadata, and already-owned
library files are unchanged. These recovery rules apply to Mylar and folder
imports, in both copy and keep-in-place modes.

Recovery queries must not expand an entire library into SQL bind parameters.
Mixed-folder lookups join existing references and discard exact same-title
rows before loading archive diagnostics; the final shared identity rules still
Expand Down
3 changes: 2 additions & 1 deletion src/pullbox/api/v1/health.py
Original file line number Diff line number Diff line change
Expand Up @@ -25,6 +25,7 @@
import pullbox
from pullbox.api.deps import DbSession, InteractiveOperatorUser # noqa: TC001
from pullbox.core.exceptions import ValidationError
from pullbox.database import DatabaseMaintenanceBusyError
from pullbox.models.health import HealthCurrentStatus as HealthCurrentStatusModel
from pullbox.models.health import HealthStatus
from pullbox.models.import_job import ImportJob, ImportJobStatus
Expand Down Expand Up @@ -430,7 +431,7 @@ async def optimize_database(

try:
result = await DatabaseOptimizationRuntimeService(db_path).optimize()
except DatabaseOptimizationError as exc:
except (DatabaseOptimizationError, DatabaseMaintenanceBusyError) as exc:
raise ValidationError(str(exc)) from exc

await run_health_refresh(component="database")
Expand Down
20 changes: 18 additions & 2 deletions src/pullbox/api/v1/system.py
Original file line number Diff line number Diff line change
Expand Up @@ -49,6 +49,7 @@
from pullbox.core.config_resolver import load_system_config_values
from pullbox.core.shutdown import shutdown_manager
from pullbox.core.sqlite_lock import is_sqlite_locked_error
from pullbox.database import DatabaseMaintenanceBusyError
from pullbox.models.config import DEFAULT_SYSTEM_CONFIG, SystemConfig
from pullbox.services.backup_runtime_service import BackupRuntimeService
from pullbox.services.backup_service import BackupService
Expand Down Expand Up @@ -279,7 +280,14 @@ async def create_backup(
) -> BackupCreatedResponse:
"""Create a manual backup of the Pullbox database."""
svc = await _get_backup_runtime_service(session)
info = await svc.create_backup(backup_type="manual")
await session.commit()
await session.close()
try:
info = await svc.create_backup(backup_type="manual")
except DatabaseMaintenanceBusyError as exc:
raise HTTPException(
status_code=503, detail=str(exc), headers={"Retry-After": "60"}
) from exc
return BackupCreatedResponse(
message=f"Backup created: {info.filename}",
backup=BackupResponse(
Expand Down Expand Up @@ -383,7 +391,15 @@ async def restore_backup(

raise ValidationError(f"Invalid backup filename: {filename}")
svc = await _get_backup_runtime_service(session)
if not await svc.restore_backup(filename):
await session.commit()
await session.close()
try:
restored = await svc.restore_backup(filename)
except DatabaseMaintenanceBusyError as exc:
raise HTTPException(
status_code=503, detail=str(exc), headers={"Retry-After": "60"}
) from exc
if not restored:
from pullbox.core.exceptions import NotFoundError

raise NotFoundError("Backup", filename)
Expand Down
7 changes: 7 additions & 0 deletions src/pullbox/app.py
Original file line number Diff line number Diff line change
Expand Up @@ -669,6 +669,13 @@ def _cleanup_import_metadata_recovery_task(task: asyncio.Task[object]) -> None:
except Exception:
logger.warning("search_wanted_sweep_recovery_failed", exc_info=True)

try:
from pullbox.tasks.metadata_sweep_state import recover_metadata_sweep_schedules

await recover_metadata_sweep_schedules()
except Exception:
logger.warning("metadata_sweep_recovery_failed", exc_info=True)

search_on_add_recovery_task = asyncio.create_task(recover_recent_search_on_add_misses())
_startup_background_tasks.add(search_on_add_recovery_task)

Expand Down
Loading
Loading