Issue #442 — files tagged with one spelling of an artist's name (Japanese kanji `澤野弘之`) get quarantined when SoulSync expects the romanized spelling (`Hiroyuki Sawano`). Raw similarity comparison scored 0% across scripts. MusicBrainz exposes alternate-spelling aliases on every artist record but the verifier never consulted them. This commit adds the pure helper that does the alias-aware comparison. No I/O, no DB access, no network. Caller supplies the aliases (looked up from library DB or live MB by later commits in this PR). Default threshold matches the verifier's existing `ARTIST_MATCH_THRESHOLD` (0.6) so wiring this in preserves current pass/fail semantics on the no-alias path. # API ``` artist_names_match(expected, actual, *, aliases=None, threshold=0.6, similarity=None) -> (matched: bool, best_score: float) ``` - Direct compare first (fast path + baseline score) - If below threshold, score each alias against `actual` - First alias to clear threshold → match - Returns the best score across all candidates so callers can log the score they made the decision on ``` best_alias_match(expected, actual, aliases=None, *, similarity=None) -> (winner: Optional[str], best_score: float) ``` Companion helper for callers that want to surface WHICH alias triggered the match (debug logs, UI explanations). No threshold — purely informative. # Architectural choices - **Pure function**: no I/O. Caller (verifier, future matching-engine consumers) owns alias lookup strategy + threshold tuning. - **Custom similarity callable**: lets the verifier pass its parenthetical-stripping normaliser without this module having to know about it. Defaults to lowercase + SequenceMatcher (matches the verifier's existing behaviour). - **Defensive coercion**: aliases input handles None entries, empty strings, non-string types, sets, tuples, lists — caller may feed raw MB response data without cleaning first. - **Backward compat**: `aliases=None` or empty → behaves identically to a plain similarity check. Paths not yet wired up to alias lookup see no behaviour change. # Tests (28) - Direct compare (no aliases): exact / case / whitespace / fuzzy / different - Cross-script with aliases: Japanese ↔ romanized (reporter's case 1), Cyrillic ↔ Latin (reporter's case 2), symmetric direction, no-match fallthrough so aliases don't mask genuine mismatches - Aliases input handling: None, empty, set, tuple, None-entries, non-string entries - Threshold: default matches verifier's 0.6, custom stricter, custom looser - Custom similarity: applies to both direct + alias compare - Best-alias-match introspection - Backward compat parametrised across 5 cases # What this commit does NOT do This is the helper module + tests only. Subsequent commits in this PR populate aliases (MB worker), provide live MB lookup with cache for un-enriched artists, and wire the helper into the AcoustID verifier where the quarantine decision actually fires.
175 lines
6.5 KiB
Python
175 lines
6.5 KiB
Python
"""Pure-function artist-name comparison with alias awareness.
|
|
|
|
Issue #442 — cross-script artist quarantines
|
|
-----------------------------------------------------
|
|
|
|
A file tagged with one spelling of an artist's name (e.g. the
|
|
Japanese kanji `澤野弘之`) was being quarantined when SoulSync's
|
|
expected-artist metadata used the romanized spelling
|
|
(`Hiroyuki Sawano`). Raw similarity comparison scores 0% across
|
|
scripts even though MusicBrainz already knows both names belong to
|
|
the same artist (its alias list).
|
|
|
|
This module is the shared resolution helper. Given an expected
|
|
artist name, an actual artist name, and an iterable of known
|
|
aliases, it returns whether they should be treated as the same
|
|
artist + the highest similarity score across the candidate set.
|
|
|
|
Pure function design:
|
|
- No I/O, no DB access, no network
|
|
- Caller supplies aliases (looked up from library DB or live MB)
|
|
- Caller supplies normalize + similarity functions to keep the
|
|
helper provider-neutral (the verifier and the matching engine
|
|
use slightly different normalizers — let each pass its own)
|
|
- Returns ``(matched: bool, score: float)`` so callers can log
|
|
the score they made the decision on
|
|
|
|
Backward compat: when ``aliases`` is empty (or the looking-up
|
|
caller hasn't been wired yet), the helper degrades to a plain
|
|
direct similarity comparison — identical to the pre-fix behaviour.
|
|
"""
|
|
|
|
from __future__ import annotations
|
|
|
|
from difflib import SequenceMatcher
|
|
from typing import Callable, Iterable, Optional, Tuple
|
|
|
|
|
|
# Default threshold matches the existing ARTIST_MATCH_THRESHOLD in
|
|
# core/acoustid_verification.py. Callers can override but the helper
|
|
# defaults are tuned to preserve current verifier behaviour.
|
|
DEFAULT_ARTIST_MATCH_THRESHOLD = 0.6
|
|
|
|
|
|
def _default_normalize(text: str) -> str:
|
|
"""Lowercase + strip whitespace. Minimal — caller's normaliser
|
|
almost always replaces this with something stricter (parenthetical
|
|
stripping, punctuation removal). Used only when the caller
|
|
doesn't pass a custom one."""
|
|
if not text:
|
|
return ''
|
|
return str(text).strip().lower()
|
|
|
|
|
|
def _default_similarity(a: str, b: str) -> float:
|
|
"""SequenceMatcher ratio after the default normaliser. Matches
|
|
the verifier's existing ``_similarity`` semantics for the no-
|
|
custom-callable path."""
|
|
na = _default_normalize(a)
|
|
nb = _default_normalize(b)
|
|
if not na or not nb:
|
|
return 0.0
|
|
if na == nb:
|
|
return 1.0
|
|
return SequenceMatcher(None, na, nb).ratio()
|
|
|
|
|
|
def _coerce_aliases(aliases: Optional[Iterable[str]]) -> Tuple[str, ...]:
|
|
"""Normalise the aliases input to a tuple of clean strings.
|
|
|
|
Accepts ``None``, empty iterables, lists, tuples, sets. Drops
|
|
None / empty / non-string entries silently — callers feeding us
|
|
raw MusicBrainz response dicts shouldn't have to clean first.
|
|
"""
|
|
if not aliases:
|
|
return ()
|
|
cleaned = []
|
|
for value in aliases:
|
|
if value is None:
|
|
continue
|
|
text = str(value).strip()
|
|
if text:
|
|
cleaned.append(text)
|
|
return tuple(cleaned)
|
|
|
|
|
|
def artist_names_match(
|
|
expected: str,
|
|
actual: str,
|
|
*,
|
|
aliases: Optional[Iterable[str]] = None,
|
|
threshold: float = DEFAULT_ARTIST_MATCH_THRESHOLD,
|
|
similarity: Optional[Callable[[str, str], float]] = None,
|
|
) -> Tuple[bool, float]:
|
|
"""Compare ``expected`` and ``actual`` artist names with alias
|
|
awareness.
|
|
|
|
Args:
|
|
expected: The artist name the caller expected (typically from
|
|
metadata-source data — Spotify / iTunes / Deezer track
|
|
payload).
|
|
actual: The artist name the caller observed (typically from
|
|
an AcoustID recording or a downloaded file's tag).
|
|
aliases: Iterable of known alternate spellings for ``expected``.
|
|
Each one gets compared against ``actual``; the best score
|
|
wins. Empty or omitted → plain direct comparison
|
|
(backward-compat with pre-fix behaviour).
|
|
threshold: Score at or above which we consider the names a
|
|
match. Defaults to 0.6 to match the verifier's existing
|
|
``ARTIST_MATCH_THRESHOLD``.
|
|
similarity: Optional caller-supplied similarity function
|
|
``(a, b) -> float in [0, 1]``. Lets the verifier pass its
|
|
stricter normaliser (parenthetical stripping etc.) without
|
|
this module having to know about it. Defaults to a
|
|
lowercase + SequenceMatcher comparison.
|
|
|
|
Returns:
|
|
``(matched, best_score)`` where ``matched`` is True iff the
|
|
best score across (actual, *aliases) ≥ threshold and
|
|
``best_score`` is that maximum. ``best_score`` is informative
|
|
for callers that want to log "matched at 0.83" or similar.
|
|
"""
|
|
sim = similarity or _default_similarity
|
|
|
|
# Direct compare first — both for the fast path and so the
|
|
# returned score reflects the actual-vs-expected baseline (callers
|
|
# may want it for logging even when an alias is the actual winner).
|
|
direct_score = sim(expected, actual)
|
|
best_score = direct_score
|
|
if direct_score >= threshold:
|
|
return True, direct_score
|
|
|
|
# Alias compare: each alias is a known alternate spelling of the
|
|
# EXPECTED artist; match it against the ACTUAL name we observed.
|
|
# Highest score wins.
|
|
for alias in _coerce_aliases(aliases):
|
|
score = sim(alias, actual)
|
|
if score > best_score:
|
|
best_score = score
|
|
if score >= threshold:
|
|
return True, score
|
|
|
|
return False, best_score
|
|
|
|
|
|
def best_alias_match(
|
|
expected: str,
|
|
actual: str,
|
|
aliases: Optional[Iterable[str]] = None,
|
|
*,
|
|
similarity: Optional[Callable[[str, str], float]] = None,
|
|
) -> Tuple[Optional[str], float]:
|
|
"""Return the alias that best matched ``actual`` (or None for the
|
|
direct expected-vs-actual comparison) and its score.
|
|
|
|
Companion to ``artist_names_match`` for callers that want to
|
|
surface which alias triggered the match (debug logging, UI
|
|
explanations). Doesn't apply a threshold — purely informative.
|
|
|
|
Returns:
|
|
``(winner, score)`` where ``winner`` is the alias string when
|
|
an alias outscored the direct comparison, ``None`` when the
|
|
direct comparison won (or both tied at zero).
|
|
"""
|
|
sim = similarity or _default_similarity
|
|
direct_score = sim(expected, actual)
|
|
winner: Optional[str] = None
|
|
best = direct_score
|
|
|
|
for alias in _coerce_aliases(aliases):
|
|
score = sim(alias, actual)
|
|
if score > best:
|
|
best = score
|
|
winner = alias
|
|
|
|
return winner, best
|