ADR-0012: Adobe Glyph List Names for Verbalising Playlist File Names
Status
Accepted
Date
2026-10-07
Context
Exported playlists are written to a file named after the playlist's display
name (ADR-0009). That name comes from the provider, so it is whatever the user
called their playlist, and it has to be reduced to something safe to put on
disk. The original sanitise_file_name folded it to ASCII first: NFKD
normalise, drop everything non-ASCII, then delete anything that was not a word
character, a space, or a hyphen.
That worked for "Road Trip" and penalised everything else. "café" became
"cafe", "Ольга" became "", and a name made only of symbols — "!!!!" —
became "" as well. An empty stem is not merely unhelpful, it is dangerous:
joining an empty string to the output directory collapses the path back to the
directory itself, and .with_suffix() then appends to the parent, so a
playlist called "!!!" would try to write /srv/playlists.m3u8, one level
above the configured output directory.
ext4, APFS, and NTFS have accepted UTF-8 file names for decades, so there is
no reason left to fold names to ASCII: the cost lands on every non-English
name and the benefit is a restriction nobody asks for. What remains is a
narrower problem. A name with no letters at all — "!!!!", "…", "🎵", or
an empty Name from the database — still has to become a file, deterministically
and without ever escaping the output directory.
Decision
Keep what a file name can hold
sanitise_file_name NFKC-normalises the name, then keeps Unicode letters,
digits, combining marks, and underscores, plus whitespace and -/_ (runs
collapsed to one -, then trimmed). Everything a file name cannot hold —
/ : * ? " < > |, NUL, bidi overrides — is dropped.
Combining marks are kept deliberately. Python's \w is only str.isalnum()
or _, and isalnum() covers categories L* and N* but not M*, so the
obvious replacement re.sub(r"[^\w\s-]", "", ...) silently deletes 2,501 of
Unicode's combining marks. It mangles every Indic, Southeast Asian, and
Tibetan name — गानों की सूची would become गन-क-सच — and drops the
harakat and niqqud of vocalised Arabic and Hebrew. The mark-aware
comprehension replaces the regex at the same line count; over all 294,579
assigned codepoints the two agree on everything except those marks, so it adds
no unsafe character.
Spell out what is left
When sanitising empties the name, verbalised_sanitise_file_name replaces
each character with its Adobe Glyph List name, uppercased: ! becomes
EXCLAM, , becomes COMMA, / becomes SLASH, so "!!!" spells
EXCLAMEXCLAMEXCLAM. The table holds the 33 printable ASCII symbols a file
name cannot hold plus space; letters and digits have no entries because
sanitise_file_name keeps them. NFKC runs first, so !!! folds to !!!
and … to ... before lookup.
unicodedata.name() and code point escapes are not consulted. They would
make the mapping total, but two emoji-only playlists would both render
MUSICALNOTE and collide on one file, and U1F3B5 tells the user nothing.
Unmappable characters instead contribute nothing.
The identifier fallback
If the spelling is empty too — "🎵", "" — the file is named after the
provider and playlist ID: spotify-pl5.m3u8. This is a correctness rule, not
a taste one: an empty stem joins back to the output directory itself, so the
fallback is what keeps the write inside playlist_output_dir. Both fallbacks
share one warning: Playlist name held no filename characters, using fallback
name.
The byte budget
A file name component is capped at 255 bytes on ext4, APFS, and NTFS alike,
and the cap is in bytes, not characters: CJK costs 3 bytes per character and
emoji 4, so a stem that fits comfortably in characters can fail the write with
ENAMETOOLONG. Stems are therefore cut with _shorten_to_bytes to
MAX_FILENAME_BYTES - len(".m3u8") = 250 bytes, landing on a character
boundary rather than through one. Truncation happens in _playlist_filename,
the single place that names a file, so the rule has one owner.
The warning
WARNING Playlist file name truncated to fit the filesystem fires only when
a cut actually happens, and carries both playlist_name and the resulting
file_name. It cannot bite an ordinary Latin name — those sit nowhere near
250 bytes — but it can bite an 84-character CJK name or a 42-symbol joke name,
the cases a user is least likely to predict. The partial-token case is the
strongest reason: EXCLAM is 6 bytes, so 42 marks spell 252 bytes and are cut
to 250, which is not a multiple of 6. The written name ends ...EXCLAMEXCL
and would look like corruption next to the plain INFO Wrote playlist file
line, were the warning not there to say the filesystem caused it.
Consequences
Positive
- Names in any script survive to disk as written:
"Ольга"exports asольга.m3u8,"हालेर गान"asहालेर-गान.m3u8. - The AGL table, the provider-ID fallback, and the byte budget between them
guarantee a file lands inside
playlist_output_dirfor every possible name, at no more than 255 bytes. - The mapping is deterministic, total, and dependency-free: no table beyond the printable ASCII symbols, and the published glyph names are stable.
- Every fallback and every cut is logged, so an unexpected file name can be traced back to the playlist name that produced it.
Negative
- NFKC folds compatibility characters onto their ASCII equivalents
(
①→1,fi→fi), so two distinct names can resolve to the same stem; the later export overwrites the earlier. Rare, and NFKD is worse — it splits composed characters apart, which truncation can then cut through. - A name over 250 bytes is cut, never deduplicated, so two names sharing a 250-byte prefix resolve to the same file. The warning is what keeps this honest.
"!!!🎵"spellsEXCLAMEXCLAMEXCLAMwith the emoji silently dropped, since the non-empty result means neither fallback fires.
Neutral
- Windows reserved device names (
con,prn, …) are not handled; that gap is pre-existing and largely moot since Windows 10. - NFC/NFD display drift between macOS and Linux on a shared volume is possible; a single OS is unaffected.
- The symbol table gives letters and digits no entries, because
sanitise_file_namepreserves them and they are never replaced.
Alternatives Considered
The plain regex keep-filter
re.sub(r"[^\w\s-]", "", value) is the same line count as the comprehension
and reads more briefly. Rejected: it strips all 2,501 combining marks and
corrupts names across roughly a dozen scripts (see Context), while adding no
character the comprehension would not also keep.
Unicode character names for every symbol
unicodedata.name() would need no table and would cover emoji. Rejected: it
collides on emoji-only names (both render MUSICALNOTE), and its names are
often less recognisable than the AGL's — Unicode calls / SOLIDUS and _
LOW LINE.
Code points for everything
U1F3B5 is deterministic and compact but tells the user nothing. Rejected:
unrecognisable was the problem being solved.
Avoiding truncation entirely
Append a short hash or the playlist ID when a stem overruns, so no name is ever shortened and no warning is needed. Rejected: it adds a second naming scheme, collides with the verbalised output, and makes names less recognisable rather than more.