Three related fixes to make album/track results look like a real
artist discography instead of a firehose of fan-compiled bootlegs.
1. Drop 'compilation' from the release-group browse primary-type filter.
MB's OR filter (`type=album|ep|single|compilation`) silently breaks
when 'compilation' is included — Metallica drops from 1076 matches
to 82 because `compilation` is a SECONDARY type on MB, not a primary
type. The invalid value corrupts the filter for all types, not just
itself. Now we request `type=album|ep|single` which returns the full
1076; actual compilations (primary=Album + secondary=[Compilation])
are filtered out by the studio-preference logic below.
2. Filter release-groups with non-studio secondary-types
(Live/Compilation/Soundtrack/Remix/Demo/Mixtape/Interview/Audiobook/
Audio drama). For Metallica, the first 100 browse results are 12
studio albums + 83 live bootlegs + 5 compilations — without this
filter the Albums section was dominated by 2019-2021 broadcast
recordings. Falls back to the unfiltered list if filtering leaves
the result set empty (covers live-only niche artists).
3. Sort chronologically ASC by first-release-date. Wikipedia-style
discography ordering — debut album on top, then chronological.
Previous DESC sort put the most recent release on top which, for
prolific artists, meant 2020s material before their classics.
Track side of the same fix:
- Re-orders each recording's `releases` array to put studio releases
first before `_recording_to_track` picks up the first release for
album context. Without this, MB's arbitrary release order often
buried the canonical studio album under random live bootlegs.
- Filters out recordings that only exist on live/compilation release-
groups (keeps the ones with at least one studio release). Falls
back to the full set if the artist has no studio recordings at all.
- Sorts recordings by earliest studio-release year ASC so classic
tracks surface first.
Smoke test against live MB API confirmed:
- Artists: [Metallica score=100]
- Albums: Kill 'Em All (1983) → Ride the Lightning → Master of Puppets
→ ...And Justice for All → Metallica (Black Album) → Load → Reload
→ St. Anger → Death Magnetic → Lulu (2011)
- Tracks: real Metallica recordings (Killing Time, Nothing Else
Matters, Creeping Death, etc.) — a few remastered demos still leak
in where MB metadata quality is thin, but the bulk is correct.
- Total latency: 3.5 seconds.
4 new tests covering the studio filter, live-only fallback, preferred
release ordering, and live-only recording exclusion.
Credit: kettui flagged the poor MB results during PR #371 review.
The previous commit's `browse_artist_recordings` call passed
`inc=releases+artist-credits` — but MusicBrainz's recording browse
endpoint rejects `inc=releases` with HTTP 400. The adapter's error
handler returned an empty list, so the Tracks section stayed empty
even though the fix was supposed to populate it.
Browse without release info is useless for our search UI (tracks
would render with no album), so swap to the fielded Lucene search
`arid:<mbid>` on the `/recording` endpoint. That's the canonical MB
pattern for "find recordings by this artist WITH release context":
- arid: search accepts the artist MBID and returns recordings with
`releases` (release-group, date, media) embedded in each result.
- One API call per lookup, same as browse would have been.
Renamed the method to `search_recordings_by_artist_mbid` so the name
matches its behaviour — it's a search, not a browse. Adapter updated
to call the new name; tests updated to match.
Verified against the live API: Metallica's MBID returns 5 recordings
in ~1.8 seconds (vs the previous 400 error).
Cover Art Archive URLs are deterministic from the MBID: a GET either
307-redirects to the image or returns 404. The previous adapter fired
`requests.head(timeout=3)` per search result to probe for the image
first. 10 results × 3s worst-case = up to 30s of blocking HEAD calls
before a search returned.
The probe was defensive overhead — the frontend already handles 404 via
`<img onerror>` fallback. Building the URL deterministically and letting
the browser load it lazily collapses the tail latency to the real MB API
calls (artist-search + browse = ~3s at the 1-rps rate limit).
Also prefer release-group scope over per-release scope when both are
available — release-group covers every edition of an album, so the hit
rate is noticeably higher than pinning to a specific regional release.
Removes now-unused `self._art_cache` and the `requests` import.
Bare name queries (typing 'metallica') now resolve to an artist MBID via
the fuzzy search added in the previous commit, then BROWSE that artist's
release-groups and recordings instead of text-searching release/recording
titles. That's the only way to fix the core garbage-results issue: MB
indexes release/recording titles, not artist names, so 'recording:metallica'
matches random tracks literally titled 'Metallica' (all scoring 100).
Structure:
- `_split_structured_query` — detects 'Artist - Title' / 'Artist – Title' /
'Artist — Title' shapes. When present, text-search is correct (user
gave an explicit title to match).
- `_resolve_top_artist` — memoized per-instance lookup for the top-scoring
artist MBID. Backend fires artists/albums/tracks searches in parallel
against one shared client instance, and albums+tracks both need the
same artist lookup. Cache + lock means one HTTP call instead of three.
- `_release_group_to_album` / `_recording_to_track` — shared projection
helpers between the browse and text paths so both paths return the
same dataclass shape.
Search flow per kind:
- `search_albums('metallica')` → resolve top artist → browse release-groups
with `type=album|ep|single|compilation` → sort by type priority then
release date desc → Album dataclasses for top N.
- `search_tracks('metallica')` → resolve top artist → browse recordings
with `inc=releases+artist-credits` → dedupe by normalized title (MB
has many live/compilation variants of the same song) → sort by release
date desc → Track dataclasses for top N.
- `search_albums('foo - bar')` → structured query → text-search path
(unchanged behavior, now score-filtered to 80+).
- `search_tracks('foo - bar')` → same.
- Both text-search paths also dedupe through `_search_albums_text` /
`_search_tracks_text` helpers, which apply the 80-score filter that
the artist-first path gets free from the resolver's threshold.
Also dedupes text-path tracks through the new `_recording_to_track`
helper, replacing ~60 lines of inline projection code. Net change is
more lines overall (browse + helpers) but the text paths shrank and
the garbage-results issue is fixed.
Credit: kettui flagged the missing Artists section + unusable track
results during PR #371 review.
`MusicBrainzSearchClient.search_artists` has been a `return []` stub
since the feature landed, with a comment claiming the MB tab 'doesn't
show artists.' That's why kettui saw a missing Artists section on the
search page — not a missing render, a hardcoded empty list.
Re-enable it properly:
- New `strict=False` parameter on `MusicBrainzClient.search_artist`
sends a bare Lucene query instead of `artist:"..."`. MusicBrainz
matches bare queries against alias+artist+sortname indexes together,
which is the right behavior for user-facing fuzzy search (finds
typos, aliases, sortname variants). `strict=True` remains the
default for enrichment/AcoustID callers that want exact matches.
- Adapter filters results to `score >= 80`. MB assigns a 0-100 Lucene
score on every hit; the true artist + close variants score 100,
tribute bands and lookalikes typically land in the 40-65 range.
The cutoff keeps "Metallica" (100) and drops "Black Metallica
Tribute Band" (60) without hand-curated lists.
- Results returned as the same `Artist` dataclass used elsewhere in
the search-tab adapter layer. `popularity` carries the MB score
(0-100) so the frontend can sort/highlight top matches if desired.
MusicBrainz mandates a meaningful User-Agent with contact info, warning
that bare strings can trigger IP blocking under load. Our client was
sending `SoulSync/2.3` with no contact — and the search adapter passed
an app version hard-coded at "2.3" that's now stale (UI is at 2.40).
Fix: default contact to the project URL (`https://github.com/Nezreka/SoulSync`)
when no email is supplied, so every request lands as
`SoulSync/<version> ( https://github.com/Nezreka/SoulSync )`. Drop the
search-adapter version suffix to a generic "2" since the exact UI minor
version would add noise to every MB request without helping operators
track issues.
Reference: https://musicbrainz.org/doc/MusicBrainz_API — "it is
important that your application sets a proper User-Agent string."
New MusicBrainz tab in Enhanced and Global search — finds tracks and
albums on MusicBrainz's community database with Cover Art Archive
images. Covers obscure tracks that Spotify/Deezer/iTunes miss.
- core/musicbrainz_search.py: search adapter with Track/Artist/Album
dataclasses, Cover Art Archive integration, smart query parsing
- Albums deduplicated (keeps best version with date and art)
- No artist results shown (MusicBrainz has no artist images)
- Album detail with full tracklist for download modal
- Smart word-boundary splitting for queries without separators
- Global search results container widened from 620px to 920px
- UI version bumped to 2.32