soulsync/core/repair_jobs/dead_file_cleaner.py
Broque Thomas d8437c87c6 Fix Album Completeness Auto-Fill on Docker / shared-library setups (#476)
GitHub issue #476 (gabistek, Docker on Arch host): "Auto-Fill" / "Fix
Selected" on the Album Completeness findings page returned
"Could not determine album folder from existing tracks" for every album.
Reproduces on any setup where the media-server library lives outside the
SoulSync transfer/download folders — Docker is the headline case but
native installs that point Plex at a NAS via SMB hit it too.

Root cause: `core/repair_worker.py:_resolve_file_path` only probed the
transfer + download folders. Docker users have their Plex/Jellyfin
library bind-mounted at /music (or similar) — neither configured in
SoulSync. Every existing track got silently treated as missing, so
`album_folder` stayed None and the fix workflow bailed.

The same incomplete logic was duplicated four more times in the
repair_jobs/ modules, all with the same bug. Album Completeness was
just the most user-visible — the same setups were also producing false
"missing file" findings from Dead File Cleaner, silent skips in
MBID Mismatch Detector, etc.

The web server already had the correct logic at
`web_server.py:_resolve_library_file_path` (probes transfer + download
+ Plex-reported library locations + user-configured library.music_paths).
The repair workers had never been updated to match.

Fix:
- New `core/library/path_resolver.py` extracts the union logic into a
  single shared function `resolve_library_file_path()`. Probes (in
  order, deduped): explicit transfer/download kwargs, config-derived
  soulseek.transfer_path/download_path, Plex-reported library
  locations (when a plex_client is passed), user-configured
  library.music_paths. Each defensive: malformed config or a flaky
  Plex client degrades to the dirs that did succeed.
- `core/repair_worker.py:_resolve_file_path` becomes a delegating
  wrapper preserving the legacy signature, with a new `config_manager`
  kwarg. All 15 in-tree call sites updated to thread
  `self._config_manager` through.
- `core/repair_jobs/dead_file_cleaner.py`,
  `mbid_mismatch_detector.py`, and `lossy_converter.py` get the same
  treatment: duplicate function replaced with a thin wrapper, call
  sites pass `context.config_manager`.
- `core/repair_jobs/acoustid_scanner.py` and
  `unknown_artist_fixer.py` (which used to import from repair_worker)
  now call the shared resolver directly with `context.config_manager`.

Side benefit: every other repair job (Dead File Cleaner, MBID
Mismatch Detector, Lossy Converter, AcoustID Scanner, Unknown Artist
Fixer) also stops missing files in the media-server library mount.
Single fix unblocks five user-visible features.

Tests: `tests/library/test_path_resolver.py` — 20 cases covering all
four base-dir sources, suffix-walk algorithm, dedup, defensive paths
(None plex client, malformed config entries, raising config_manager.get,
broken plex attribute access), Docker path translation. Full suite
1677 passed locally.

WHATS_NEW entry under '2.4.2' dev cycle.
2026-05-03 10:11:06 -07:00

169 lines
7 KiB
Python

"""Dead File Cleaner Job — finds DB track entries where the file no longer exists."""
import os
from core.library.path_resolver import resolve_library_file_path
from core.repair_jobs import register_job
from core.repair_jobs.base import JobContext, JobResult, RepairJob
from utils.logging_config import get_logger
logger = get_logger("repair_job.dead_files")
def _resolve_file_path(file_path, transfer_folder, download_folder=None, config_manager=None):
"""Backwards-compat wrapper. Use ``resolve_library_file_path`` directly."""
return resolve_library_file_path(
file_path,
transfer_folder=transfer_folder,
download_folder=download_folder,
config_manager=config_manager,
)
@register_job
class DeadFileCleanerJob(RepairJob):
job_id = 'dead_file_cleaner'
display_name = 'Dead File Cleaner'
description = 'Finds database entries pointing to missing files'
help_text = (
'Checks every track in your database to verify the actual audio file still exists '
'on disk. If a file has been moved, renamed, or deleted outside of SoulSync, the '
'database entry becomes a "dead" reference.\n\n'
'Each dead reference is reported as a finding. You can then resolve it by re-downloading '
'the track or dismiss it to clean up the database entry.\n\n'
'This job only scans and reports — it never deletes database entries automatically.'
)
icon = 'repair-icon-deadfile'
default_enabled = True
default_interval_hours = 24
default_settings = {}
auto_fix = False
def scan(self, context: JobContext) -> JobResult:
result = JobResult()
# Safety: abort if transfer folder doesn't exist — prevents mass false positives
if not context.transfer_folder or not os.path.isdir(context.transfer_folder):
logger.error("Transfer folder not found: %s — aborting dead file scan to avoid false positives",
context.transfer_folder)
result.errors += 1
if context.report_progress:
context.report_progress(
phase='Aborted — transfer folder not found',
log_line=f'Transfer folder does not exist: {context.transfer_folder}',
log_type='error'
)
return result
# Fetch all tracks with file paths, joining to get artist/album names
tracks = []
conn = None
try:
conn = context.db._get_connection()
cursor = conn.cursor()
cursor.execute("""
SELECT t.id, t.title, ar.name, al.title, t.file_path,
al.thumb_url, ar.thumb_url
FROM tracks t
LEFT JOIN artists ar ON ar.id = t.artist_id
LEFT JOIN albums al ON al.id = t.album_id
WHERE t.file_path IS NOT NULL AND t.file_path != ''
""")
tracks = cursor.fetchall()
except Exception as e:
logger.error("Error fetching tracks from DB: %s", e, exc_info=True)
result.errors += 1
return result
finally:
if conn:
conn.close()
total = len(tracks)
if context.update_progress:
context.update_progress(0, total)
# Get download folder for path resolution fallback
download_folder = None
if context.config_manager:
download_folder = context.config_manager.get('soulseek.download_path', '')
if context.report_progress:
context.report_progress(phase=f'Checking {total} tracks...', total=total)
for i, row in enumerate(tracks):
if context.check_stop():
return result
if i % 200 == 0 and context.wait_if_paused():
return result
track_id, title, artist_name, album_title, file_path, album_thumb, artist_thumb = row
result.scanned += 1
if context.report_progress and i % 50 == 0:
context.report_progress(
scanned=i + 1, total=total,
phase=f'Checking {i + 1} / {total}',
log_line=f'Checking: {title or "Unknown"}{artist_name or "Unknown"}',
log_type='info'
)
# Use the same path resolution logic as library playback
resolved = _resolve_file_path(file_path, context.transfer_folder, download_folder,
config_manager=context.config_manager)
if resolved is None:
# File is truly missing — create finding
if context.report_progress:
context.report_progress(
log_line=f'Missing: {title or "Unknown"}{os.path.basename(file_path)}',
log_type='error'
)
if context.create_finding:
try:
context.create_finding(
job_id=self.job_id,
finding_type='dead_file',
severity='warning',
entity_type='track',
entity_id=str(track_id),
file_path=file_path,
title=f'Missing file: {title or "Unknown"}',
description=f'Track "{title}" by {artist_name or "Unknown"} points to a file that no longer exists',
details={
'track_id': track_id,
'title': title,
'artist': artist_name,
'album': album_title,
'original_path': file_path,
'album_thumb_url': album_thumb or None,
'artist_thumb_url': artist_thumb or None,
}
)
result.findings_created += 1
except Exception as e:
logger.debug("Error creating dead file finding for track %s: %s", track_id, e)
result.errors += 1
if context.update_progress and (i + 1) % 100 == 0:
context.update_progress(i + 1, total)
if context.update_progress:
context.update_progress(total, total)
logger.info("Dead file scan: %d tracks checked, %d missing files found",
result.scanned, result.findings_created)
return result
def estimate_scope(self, context: JobContext) -> int:
conn = None
try:
conn = context.db._get_connection()
cursor = conn.cursor()
cursor.execute("SELECT COUNT(*) FROM tracks WHERE file_path IS NOT NULL AND file_path != ''")
row = cursor.fetchone()
return row[0] if row else 0
except Exception:
return 0
finally:
if conn:
conn.close()