Data model
This document describes how passive-music-dl persists state: the SQLite
database, the tags embedded in downloaded files, and the supported audio
formats.
SQLite database
State lives in a single SQLite database file (state.db by default), managed
by passive_music_dl/state/database.py. The schema is created on startup by
Database.initialize().
Track
One row per downloaded audio file.
| Column | Type | Constraints | Description |
|---|---|---|---|
Id |
INTEGER | PRIMARY KEY AUTOINCREMENT | Database ID. |
MusicBrainzId |
TEXT | UNIQUE | MusicBrainz recording ID. |
YouTubeMusicId |
TEXT | UNIQUE | YouTube video ID. |
SpotifyId |
TEXT | UNIQUE | Spotify track ID. |
FilePath |
TEXT | UNIQUE, NOT NULL | Absolute path of the downloaded file. |
Each provider ID column is nullable - a track has whichever IDs are known. The
FilePath column is the link back to the filesystem. An index exists on every
column to support lookups.
Playlist
One row per provider playlist seen during a sync.
| Column | Type | Constraints | Description |
|---|---|---|---|
PlaylistProvider |
TEXT | PRIMARY KEY | spotify or youtube_music. |
PlaylistId |
TEXT | PRIMARY KEY | Provider-specific playlist ID. |
Name |
TEXT | NOT NULL | Playlist name from the provider. |
The name is refreshed by an upsert on every sync cycle, so renames are picked up even when the playlist contents do not change.
PlaylistEntry
Associates a track with the playlist it came from.
| Column | Type | Constraints |
|---|---|---|
PlaylistProvider |
TEXT | spotify or youtube_music. |
PlaylistId |
TEXT | Provider-specific playlist ID. |
TrackId |
INTEGER | REFERENCES Track(Id) ON DELETE CASCADE. |
Primary key is (PlaylistProvider, PlaylistId, TrackId), so the same track can
belong to many playlists but only once per playlist. Each row references its
playlist in Playlist through the composite foreign key
(PlaylistProvider, PlaylistId) with ON DELETE CASCADE: deleting a playlist
removes its entries, and deleting a track cascades to its entries too.
Filesystem-based identity
The database is not the source of truth for what has been downloaded - the download directory is. On startup a full scan reconciles the two:
- a file on disk with no database row is inserted by reading its embedded IDs;
- a database row with no file on disk is deleted.
For this to work, every downloaded file carries its provider and MusicBrainz IDs in tag frames that mutagen reads back. These are the UFID (Unique File Identifier) frames, owned by namespace-style owner strings:
| ID | UFID owner | Stored value |
|---|---|---|
| MusicBrainz | http://musicbrainz.org |
Recording MBID |
| YouTube Music | https://music.youtube.com |
Video ID |
| Spotify | https://spotify.com |
Track ID |
UFID frames are read into UniqueTrackIdentifiers
(metadata/models.py) and written from TrackMetadata.
See ADR-0003: Hybrid Filesystem-Based State Management for the history behind this design.
Embedded metadata
Besides the UFID frames, each file gets its standard metadata written by the
per-format tagger (metadata/async_audio_file.py):
| Field | MP3 (ID3) | M4A (MP4) | OGG/FLAC/Opus (Vorbis) |
|---|---|---|---|
| Title | TIT2 |
©nam |
title |
| Artist | TPE1 |
©ART |
artist |
| Album | TALB |
©alb |
album |
| Year | TDRC |
©day |
date |
| Track number | TRCK |
trkn |
tracknumber |
| Genre | TCON |
©gen |
genre |
| Album art | APIC |
covr |
metadata_block_picture |
Not every field is always populated - the Spotify and YouTube Music providers
currently supply title, artist(s), album, and album art; MusicBrainz adds the
recording ID. TrackMetadata (metadata/models.py) carries all of the above
and any unused fields are simply not written.
Supported formats
The downloader converts every track to the configured codec, and the metadata
layer reads back any of the following extensions (see util/constants.py):
.mp3(ID3 tags).m4a(MP4 tags).ogg,.flac,.opus(Vorbis comments)
This lets the full scan and the taggers handle files regardless of where they came from, so long as they carry the UFID frames above.