Plans 013 and 014, the album page that prompted them, and the smaller fixes they turned up. Changelog, largest first. ## The local library is shaped like files, not like MusicBrainz `audio_files` carries its own tags and points at `albums` and `artists`; `file_genres` is the one real many-to-many. `recordings`, `release_group_recordings`, `artist_credit`, `artist_credit_artist`, `recording_genres`, `release_groups` and `release_to_rg` are gone from the local side, and with them a six-way join in every read, a `MIN(release_group_id)` subquery in eleven queries and a first-credited-artist subquery in nine. Measured on a real 25,966-file library, every many-to-many that model expressed was 1:1 in the data. - Ownership is a file. `GetFilePathsByRecordingMBIDs`, `LibraryMBIDIndex.CheckMBIDs`, `collectLibraryEntities` and `pruneStaleLocalCrossReferences` all join `audio_files`, so the 812 orphaned recordings, 216 release groups and 260 artists that library carried are now structurally impossible. - One projection: every track query selects from the `track_metadata` view, one row type, one mapper. Nine hand-rolled copies had drifted far enough to report different years on different screens. - `library_id = 0` means every library, so each list query exists once instead of scoped and unscoped with a branch at every call site. - No migration chain. `sql/schemas/` is the one description of the shape; `sql/migrations/`, `applyMigrations` and `schema_migrations` are squashed away, along with the drift between them that had sqlc generating against a stale schema. - `database.InsertTestTrack` is the one test seeder; twenty test files had been assembling the old FK chain each in its own order. ## The catalog stores its ids as bytes `explore_index`'s three 36-char MBID columns and its entity-type text are 16 raw bytes and a small integer. The table and its six indexes go 780 MB to 405 MB on a real 2,052,200-row catalog, which is why a fresh install is ~0.6 GB rather than ~1.0 GB. - `backend/explore/mbid.go` is the only place the encoding is known; everything above it speaks dashed strings. - `CHECK(length(mbid) = 16)` makes a stringly write fail at the insert rather than silently returning no rows, since SQLite does not coerce between TEXT and BLOB. - The importer asks the artifact what encoding it carries and converts on the way in, so the artifact already published keeps working and no format bump is needed. - `indexRowColumns`/`scanIndexRow` replace four copies of a 22-column list, and `TestStoredEncodingRoundTrips` sweeps every read path. ## An album page that says how much of the album is yours - One question, asked once: is there a file. `filePaths` is filled by a single batched lookup when the tracklist settles, and the badge, the Play count, the dimmed rows and every menu item read it — replacing four claims of decreasing confidence that could show a green tick on an album whose every action did nothing. - Play, Play 7 of 12, or no play button at all. - `total_tracks` on `explore_index` (~2 bytes over 400,677 release groups) and on `audio_files` from tags that have always carried it: a complete MBID-matched album now makes no catalog call at all, where it used to spend the most expensive request the app makes. - A merged cluster shows the running order the most releases agree on, and the version list marks the release you own rather than standing a synthetic entry in for it. - `AlbumReleasesFailed`: a slow fetch is no longer reported as a failed one by a 12-second timer. - Rows not in the library are dimmed in place (with `aria-disabled`) instead of the owned ones wearing a green tick and a legend. ## Caches and cover art get ceilings - Only the three tiers of a cover are stored; the full-resolution copy nothing rendered was 1,134 MB of a 1.4 GB covers directory. - One artist portrait is downloaded and the rest are remembered as URLs — 4.1 GB of a 5.3 GB cache was candidates no code path reads. - `browsedArtBudget` and `httpCacheBudget` bound what an age cannot: the same install held art for 5,770 artists in a 1,301-artist library. - `OrphanedArtistImagesJob` joined a bare MBID onto a sharded directory, so it deleted the rows that were the only record of the files it left behind. `explore.ArtistImageDir` is that layout's one definition now. ## The autotag queue asks whether there is work `tagging_items` was a row per album folder, not a queue, and no query read the `tag_status` column that held the answer. The four queue queries ask the files, which matters most where it is least visible: `startPrefetch` was scoring every album in a tagged library against MusicBrainz. ## Phantom playlist tracks resolve in place An M3U8 imported before its files leaves phantom rows; they now match by path and fall back to position, keep their place in the playlist when resolved, and pair best-first so two phantoms cannot claim the same file. ## Playing a track plays the list it is in Double-click, and Play on a single row's menu, queue the list as displayed with `startIndex` on that row — the album page and the track list used to queue one track and discard the album around it. A multi-row selection still plays exactly itself. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01AfVYUVExXsx1nSWrXN8mAh
384 lines
8.7 KiB
Go
384 lines
8.7 KiB
Go
// Package playlist provides playlist management functionality.
|
|
package playlist
|
|
|
|
import (
|
|
"math"
|
|
"path/filepath"
|
|
"regexp"
|
|
"strings"
|
|
"unicode/utf8"
|
|
)
|
|
|
|
// Scoring weights for candidate matching.
|
|
const (
|
|
weightFilename = 0.50
|
|
weightTitle = 0.30
|
|
weightDuration = 0.10
|
|
weightPathDirs = 0.10
|
|
autoMatchMinimum = 0.85
|
|
)
|
|
|
|
// maxCandidates is the default limit for search results.
|
|
const maxCandidates = 20
|
|
|
|
// maxLibrarySearchResults is the limit for manual library search.
|
|
const maxLibrarySearchResults = 50
|
|
|
|
// durationToleranceClose is the duration difference in seconds
|
|
// considered a near-exact match.
|
|
const durationToleranceClose = 1
|
|
|
|
// durationToleranceMedium is the medium tolerance threshold.
|
|
const durationToleranceMedium = 5
|
|
|
|
// durationToleranceFar is the maximum tolerance before scoring
|
|
// drops to zero.
|
|
const durationToleranceFar = 15
|
|
|
|
// separatorPattern splits file paths and names on common
|
|
// separators: slashes, hyphens, underscores, spaces, dots.
|
|
var separatorPattern = regexp.MustCompile(
|
|
`[/\\\-_. ]+`,
|
|
)
|
|
|
|
// trackNumberPattern matches leading track numbers like
|
|
// "01", "1", "01.", "01 -", etc.
|
|
var trackNumberPattern = regexp.MustCompile(
|
|
`^\d{1,3}[.\-\s]*$`,
|
|
)
|
|
|
|
// phantomProfile pre-computes all derived data for a phantom
|
|
// track so that scoring multiple candidates avoids redundant
|
|
// string processing.
|
|
type phantomProfile struct {
|
|
baseLower string // lowercase basename
|
|
baseStem string // basename without extension
|
|
baseWords []string // significant words from stem
|
|
dirWords []string // significant words from dir path
|
|
displayLow string // lowercase display title
|
|
parsedArt string // parsed artist from display title
|
|
parsedTitle string // parsed title from display title
|
|
titleWords []string // significant words from display title
|
|
durationSec int // phantom duration in seconds
|
|
}
|
|
|
|
// newPhantomProfile builds a phantomProfile from raw phantom
|
|
// data, performing all string splits and normalisation once.
|
|
func newPhantomProfile(
|
|
phantomPath string,
|
|
displayTitle string,
|
|
durationSec int,
|
|
) phantomProfile {
|
|
baseLower := strings.ToLower(
|
|
filepath.Base(phantomPath),
|
|
)
|
|
baseStem := stripExtension(baseLower)
|
|
displayLow := strings.ToLower(
|
|
strings.TrimSpace(displayTitle),
|
|
)
|
|
parsedArt, parsedTitle := parseDisplayTitle(displayLow)
|
|
|
|
return phantomProfile{
|
|
baseLower: baseLower,
|
|
baseStem: baseStem,
|
|
baseWords: significantWords(baseStem),
|
|
dirWords: pathDirWords(phantomPath),
|
|
displayLow: displayLow,
|
|
parsedArt: parsedArt,
|
|
parsedTitle: parsedTitle,
|
|
titleWords: significantWords(displayLow),
|
|
durationSec: durationSec,
|
|
}
|
|
}
|
|
|
|
// scoreCandidate computes a match confidence (0.0-1.0) between
|
|
// a phantom track and a candidate library track.
|
|
func scoreCandidate(
|
|
pp phantomProfile,
|
|
candidatePath string,
|
|
candidateTitle string,
|
|
candidateArtist string,
|
|
candidateDurationMs int64,
|
|
) float64 {
|
|
fnScore := scoreFilename(pp, candidatePath)
|
|
titleScore := scoreTitleArtist(
|
|
pp, candidateTitle, candidateArtist,
|
|
)
|
|
durScore := scoreDuration(
|
|
pp.durationSec, candidateDurationMs,
|
|
)
|
|
dirScore := scorePathDirs(pp, candidatePath)
|
|
|
|
// If duration is unknown, redistribute its weight to filename.
|
|
// Either side can be the one that does not know: an M3U8 written
|
|
// without EXTINF lines carries no duration, and neither does a
|
|
// library file whose length was never read. Scoring a candidate
|
|
// out of 0.9 for the *library's* gap made an otherwise exact
|
|
// filename match unable to reach the auto-match threshold.
|
|
fnWeight := weightFilename
|
|
durWeight := weightDuration
|
|
|
|
if pp.durationSec == 0 || candidateDurationMs == 0 {
|
|
fnWeight += durWeight
|
|
durWeight = 0
|
|
}
|
|
|
|
return fnScore*fnWeight +
|
|
titleScore*weightTitle +
|
|
durScore*durWeight +
|
|
dirScore*weightPathDirs
|
|
}
|
|
|
|
// scoreFilename compares the basenames of two file paths.
|
|
func scoreFilename(
|
|
pp phantomProfile, candidatePath string,
|
|
) float64 {
|
|
cBase := strings.ToLower(
|
|
filepath.Base(candidatePath),
|
|
)
|
|
|
|
// Exact basename match.
|
|
if pp.baseLower == cBase {
|
|
return 1.0
|
|
}
|
|
|
|
// Match ignoring extension.
|
|
cStem := stripExtension(cBase)
|
|
|
|
if pp.baseStem == cStem {
|
|
return 0.8
|
|
}
|
|
|
|
// Check if all significant words from phantom stem appear
|
|
// in candidate stem.
|
|
cWords := significantWords(cStem)
|
|
|
|
if len(pp.baseWords) == 0 {
|
|
return 0.0
|
|
}
|
|
|
|
return keywordOverlap(pp.baseWords, cWords)
|
|
}
|
|
|
|
// scoreTitleArtist compares the phantom's EXTINF display title
|
|
// against the candidate's DB title and artist fields.
|
|
func scoreTitleArtist(
|
|
pp phantomProfile,
|
|
candidateTitle, candidateArtist string,
|
|
) float64 {
|
|
if pp.displayLow == "" {
|
|
return 0.0
|
|
}
|
|
|
|
candidateTitle = strings.ToLower(
|
|
strings.TrimSpace(candidateTitle),
|
|
)
|
|
candidateArtist = strings.ToLower(
|
|
strings.TrimSpace(candidateArtist),
|
|
)
|
|
|
|
// Exact title match.
|
|
if pp.parsedTitle != "" &&
|
|
pp.parsedTitle == candidateTitle {
|
|
if pp.parsedArt != "" &&
|
|
pp.parsedArt == candidateArtist {
|
|
return 1.0
|
|
}
|
|
|
|
return 0.8
|
|
}
|
|
|
|
// Keyword overlap between display title and combined
|
|
// candidate metadata.
|
|
combined := candidateTitle + " " + candidateArtist
|
|
cWords := significantWords(combined)
|
|
|
|
if len(pp.titleWords) == 0 {
|
|
return 0.0
|
|
}
|
|
|
|
return keywordOverlap(pp.titleWords, cWords)
|
|
}
|
|
|
|
// scoreDuration computes a score based on duration proximity.
|
|
func scoreDuration(
|
|
phantomSec int, candidateMs int64,
|
|
) float64 {
|
|
if phantomSec == 0 || candidateMs == 0 {
|
|
return 0.0
|
|
}
|
|
|
|
diff := math.Abs(
|
|
float64(phantomSec) - float64(candidateMs)/1000.0,
|
|
)
|
|
|
|
switch {
|
|
case diff <= float64(durationToleranceClose):
|
|
return 1.0
|
|
case diff <= float64(durationToleranceMedium):
|
|
return 0.8
|
|
case diff <= float64(durationToleranceFar):
|
|
return 0.5
|
|
default:
|
|
return 0.0
|
|
}
|
|
}
|
|
|
|
// scorePathDirs compares the directory components of two paths.
|
|
func scorePathDirs(
|
|
pp phantomProfile, candidatePath string,
|
|
) float64 {
|
|
if len(pp.dirWords) == 0 {
|
|
return 0.0
|
|
}
|
|
|
|
cDirs := pathDirWords(candidatePath)
|
|
|
|
return keywordOverlap(pp.dirWords, cDirs)
|
|
}
|
|
|
|
// parseDisplayTitle splits an EXTINF display title on " - " into
|
|
// (artist, title). If no separator is found, returns ("", full).
|
|
func parseDisplayTitle(dt string) (artist, title string) {
|
|
idx := strings.Index(dt, " - ")
|
|
if idx < 0 {
|
|
return "", dt
|
|
}
|
|
|
|
return strings.TrimSpace(dt[:idx]),
|
|
strings.TrimSpace(dt[idx+3:])
|
|
}
|
|
|
|
// extractKeywords extracts meaningful search keywords from a file
|
|
// path by splitting on separators, removing track numbers, common
|
|
// noise words, and the file extension.
|
|
func extractKeywords(filePath string) []string {
|
|
// Remove extension.
|
|
stem := stripExtension(filePath)
|
|
|
|
// Split on separators.
|
|
parts := separatorPattern.Split(stem, -1)
|
|
|
|
var keywords []string
|
|
|
|
for _, p := range parts {
|
|
p = strings.TrimSpace(p)
|
|
if p == "" {
|
|
continue
|
|
}
|
|
|
|
// Skip pure track numbers.
|
|
if trackNumberPattern.MatchString(p) {
|
|
continue
|
|
}
|
|
|
|
// Skip very short tokens.
|
|
if len(p) < 2 {
|
|
continue
|
|
}
|
|
|
|
keywords = append(keywords, strings.ToLower(p))
|
|
}
|
|
|
|
return dedupStrings(keywords)
|
|
}
|
|
|
|
// significantWords extracts meaningful lowercase words from a
|
|
// string, filtering out noise.
|
|
func significantWords(s string) []string {
|
|
parts := separatorPattern.Split(s, -1)
|
|
|
|
var words []string
|
|
|
|
for _, p := range parts {
|
|
p = strings.TrimSpace(p)
|
|
if p == "" {
|
|
continue
|
|
}
|
|
|
|
// Skip pure track numbers.
|
|
if trackNumberPattern.MatchString(p) {
|
|
continue
|
|
}
|
|
|
|
// Skip single characters.
|
|
if countRunes(p) < 2 {
|
|
continue
|
|
}
|
|
|
|
words = append(words, strings.ToLower(p))
|
|
}
|
|
|
|
return words
|
|
}
|
|
|
|
// pathDirWords extracts lowercase words from the directory
|
|
// portion of a path (excluding the filename).
|
|
func pathDirWords(filePath string) []string {
|
|
dir := filepath.Dir(filePath)
|
|
if dir == "." || dir == "/" {
|
|
return nil
|
|
}
|
|
|
|
return significantWords(dir)
|
|
}
|
|
|
|
// keywordOverlap calculates the proportion of source words that
|
|
// appear in target words (Jaccard-like, asymmetric).
|
|
func keywordOverlap(source, target []string) float64 {
|
|
if len(source) == 0 {
|
|
return 0.0
|
|
}
|
|
|
|
targetSet := make(map[string]struct{}, len(target))
|
|
|
|
for _, w := range target {
|
|
targetSet[w] = struct{}{}
|
|
}
|
|
|
|
var matches int
|
|
|
|
for _, w := range source {
|
|
if _, ok := targetSet[w]; ok {
|
|
matches++
|
|
}
|
|
}
|
|
|
|
return float64(matches) / float64(len(source))
|
|
}
|
|
|
|
// stripExtension removes the file extension from a path or
|
|
// filename.
|
|
func stripExtension(s string) string {
|
|
ext := filepath.Ext(s)
|
|
if ext == "" {
|
|
return s
|
|
}
|
|
|
|
return s[:len(s)-len(ext)]
|
|
}
|
|
|
|
// dedupStrings removes duplicate strings, preserving order.
|
|
func dedupStrings(ss []string) []string {
|
|
seen := make(map[string]struct{}, len(ss))
|
|
|
|
var result []string
|
|
|
|
for _, s := range ss {
|
|
if _, ok := seen[s]; ok {
|
|
continue
|
|
}
|
|
|
|
seen[s] = struct{}{}
|
|
|
|
result = append(result, s)
|
|
}
|
|
|
|
return result
|
|
}
|
|
|
|
// countRunes returns the number of runes in a string.
|
|
func countRunes(s string) int {
|
|
return utf8.RuneCountInString(s)
|
|
}
|