Skip to content

initials() drops the accent from a decomposed (NFD) name: émile zola gives e. z. #585

Description

@derek73

An initial taken from a name typed in decomposed form (NFD, which macOS file names and some databases produce) loses its accent:

>>> import unicodedata
>>> parse(unicodedata.normalize("NFD", "émile zola")).initials()
'e. z.'
>>> parse("émile zola").initials()
'é. z.'

HumanName(...).initials() does the same. The parse itself is right: roles and tags for the NFD spelling equal the NFC ones.

Cause. nameparser/_render.py's initials() takes t.text[0], a token's first code point. In NFD that is the base letter e, and the combining accent (U+0301) that follows it is dropped.

Reach. Unlike capitalized(), initials() is a compared surface: the differential gate checks it under _initials (#484). The rules corpus has held two NFD names since #584, but both start with an unaccented letter (josé garcía), so the gate doesn't trip.

Options.

  1. Take the first letter plus the combining marks after it (the _past_marks helper Fix case repair splitting a decomposed (NFD) word at its accent #584 added to _render.py). The output keeps the input's form, matching how case repair now handles decomposed text (rules.md#R4, decisions.md#R4 2026-10-02).
  2. Compose the initial to NFC. Simpler, but the output's form would then differ from the input's. R3 states no "nothing but X" rule here, so this is a choice rather than a violation.

Either way, a rules.md#R3 example line in NFD should pin it. Found while fixing #542; recorded as out of scope in decisions.md#R4's #542 bullet.

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions