Skip to content

ADR-0012: Adobe Glyph List Names for Verbalising Playlist File Names

Status

Accepted

Date

2026-10-07

Context

Exported playlists are written to a file named after the playlist's display name (ADR-0009). That name comes from the provider, so it is whatever the user called their playlist, and it has to be reduced to something safe to put on disk. The original sanitise_file_name folded it to ASCII first: NFKD normalise, drop everything non-ASCII, then delete anything that was not a word character, a space, or a hyphen.

That worked for "Road Trip" and penalised everything else. "café" became "cafe", "Ольга" became "", and a name made only of symbols — "!!!!" — became "" as well. An empty stem is not merely unhelpful, it is dangerous: joining an empty string to the output directory collapses the path back to the directory itself, and .with_suffix() then appends to the parent, so a playlist called "!!!" would try to write /srv/playlists.m3u8, one level above the configured output directory.

ext4, APFS, and NTFS have accepted UTF-8 file names for decades, so there is no reason left to fold names to ASCII: the cost lands on every non-English name and the benefit is a restriction nobody asks for. What remains is a narrower problem. A name with no letters at all — "!!!!", "…", "🎵", or an empty Name from the database — still has to become a file, deterministically and without ever escaping the output directory.

Decision

Keep what a file name can hold

sanitise_file_name NFKC-normalises the name, then keeps Unicode letters, digits, combining marks, and underscores, plus whitespace and -/_ (runs collapsed to one -, then trimmed). Everything a file name cannot hold — / : * ? " < > |, NUL, bidi overrides — is dropped.

Combining marks are kept deliberately. Python's \w is only str.isalnum() or _, and isalnum() covers categories L* and N* but not M*, so the obvious replacement re.sub(r"[^\w\s-]", "", ...) silently deletes 2,501 of Unicode's combining marks. It mangles every Indic, Southeast Asian, and Tibetan name — गानों की सूची would become गन-क-सच — and drops the harakat and niqqud of vocalised Arabic and Hebrew. The mark-aware comprehension replaces the regex at the same line count; over all 294,579 assigned codepoints the two agree on everything except those marks, so it adds no unsafe character.

Spell out what is left

When sanitising empties the name, verbalised_sanitise_file_name replaces each character with its Adobe Glyph List name, uppercased: ! becomes EXCLAM, , becomes COMMA, / becomes SLASH, so "!!!" spells EXCLAMEXCLAMEXCLAM. The table holds the 33 printable ASCII symbols a file name cannot hold plus space; letters and digits have no entries because sanitise_file_name keeps them. NFKC runs first, so !!! folds to !!! and … to ... before lookup.

unicodedata.name() and code point escapes are not consulted. They would make the mapping total, but two emoji-only playlists would both render MUSICALNOTE and collide on one file, and U1F3B5 tells the user nothing. Unmappable characters instead contribute nothing.

The identifier fallback

If the spelling is empty too — "🎵", "" — the file is named after the provider and playlist ID: spotify-pl5.m3u8. This is a correctness rule, not a taste one: an empty stem joins back to the output directory itself, so the fallback is what keeps the write inside playlist_output_dir. Both fallbacks share one warning: Playlist name held no filename characters, using fallback name.

The byte budget

A file name component is capped at 255 bytes on ext4, APFS, and NTFS alike, and the cap is in bytes, not characters: CJK costs 3 bytes per character and emoji 4, so a stem that fits comfortably in characters can fail the write with ENAMETOOLONG. Stems are therefore cut with _shorten_to_bytes to MAX_FILENAME_BYTES - len(".m3u8") = 250 bytes, landing on a character boundary rather than through one. Truncation happens in _playlist_filename, the single place that names a file, so the rule has one owner.

The warning

WARNING Playlist file name truncated to fit the filesystem fires only when a cut actually happens, and carries both playlist_name and the resulting file_name. It cannot bite an ordinary Latin name — those sit nowhere near 250 bytes — but it can bite an 84-character CJK name or a 42-symbol joke name, the cases a user is least likely to predict. The partial-token case is the strongest reason: EXCLAM is 6 bytes, so 42 marks spell 252 bytes and are cut to 250, which is not a multiple of 6. The written name ends ...EXCLAMEXCL and would look like corruption next to the plain INFO Wrote playlist file line, were the warning not there to say the filesystem caused it.

Consequences

Positive

  • Names in any script survive to disk as written: "Ольга" exports as ольга.m3u8, "हालेर गान" as हालेर-गान.m3u8.
  • The AGL table, the provider-ID fallback, and the byte budget between them guarantee a file lands inside playlist_output_dir for every possible name, at no more than 255 bytes.
  • The mapping is deterministic, total, and dependency-free: no table beyond the printable ASCII symbols, and the published glyph names are stable.
  • Every fallback and every cut is logged, so an unexpected file name can be traced back to the playlist name that produced it.

Negative

  • NFKC folds compatibility characters onto their ASCII equivalents (① → 1, fi → fi), so two distinct names can resolve to the same stem; the later export overwrites the earlier. Rare, and NFKD is worse — it splits composed characters apart, which truncation can then cut through.
  • A name over 250 bytes is cut, never deduplicated, so two names sharing a 250-byte prefix resolve to the same file. The warning is what keeps this honest.
  • "!!!🎵" spells EXCLAMEXCLAMEXCLAM with the emoji silently dropped, since the non-empty result means neither fallback fires.

Neutral

  • Windows reserved device names (con, prn, …) are not handled; that gap is pre-existing and largely moot since Windows 10.
  • NFC/NFD display drift between macOS and Linux on a shared volume is possible; a single OS is unaffected.
  • The symbol table gives letters and digits no entries, because sanitise_file_name preserves them and they are never replaced.

Alternatives Considered

The plain regex keep-filter

re.sub(r"[^\w\s-]", "", value) is the same line count as the comprehension and reads more briefly. Rejected: it strips all 2,501 combining marks and corrupts names across roughly a dozen scripts (see Context), while adding no character the comprehension would not also keep.

Unicode character names for every symbol

unicodedata.name() would need no table and would cover emoji. Rejected: it collides on emoji-only names (both render MUSICALNOTE), and its names are often less recognisable than the AGL's — Unicode calls / SOLIDUS and _ LOW LINE.

Code points for everything

U1F3B5 is deterministic and compact but tells the user nothing. Rejected: unrecognisable was the problem being solved.

Avoiding truncation entirely

Append a short hash or the playlist ID when a stem overruns, so no name is ever shortened and no warning is needed. Rejected: it adds a second naming scheme, collides with the verbalised output, and makes names less recognisable rather than more.