Skip to content

Fix case repair splitting a decomposed (NFD) word at its accent - #584

Merged
derek73 merged 5 commits into
masterfrom
fix/issue-542-nfd-case-repair
Oct 3, 2026
Merged

derek73 merged 5 commits into
masterfrom
fix/issue-542-nfd-case-repair

Conversation

@derek73

@derek73 derek73 commented Oct 2, 2026

Copy link
Copy Markdown
Owner

Closes #542.

A name typed decomposed (NFD: macOS file names, some databases) writes í as i followed by a combining accent (U+0301). Case repair's word regex (\w|\.)+ matches no combining mark, so the accent split the word and the letters after it were repaired as a separate word. josé garcía came out as José GarcíA. The parse was already right; only the capitalized() / capitalize() view was wrong.

The fix (nameparser/_render.py)

This is the issue's option 1. Option 2, composing the text to NFC before repair, was declined: it hands back text in a form the writer didn't use, which rules.md#R4 forbids and test_case_repair_changes_case_and_nothing_else would report. A decomposed word now repairs the way its composed spelling does, and the output keeps the form it was typed in.

  • Word splitting. _sub_words replaces _WORD.sub: each match runs on through the combining marks after it, tested with unicodedata.category(ch)[0] == "M". A hand-written list of combining-mark ranges would drift; it would already miss Adlam. A mark with no letter before it starts no word of its own.
  • Mask checks. The two places that ask about a letter's neighbours, _apply_mask's split-off-initial check and _letter_run_ge2, now look past marks via _beside. The issue said these needed no change. Measured, that was wrong: two mask shapes still disagreed once only the splitter was fixed.
  • Shape checks. The Mac/Mc check and the two initial checks (inside hyphenated names, and for raw text spliced into a field) now test the word's composed spelling.
    • macée was MacÉe composed against MacéE decomposed before this change.
    • A decomposed й. read as the conjunction й where its composed spelling is an initial: ivan petrov-й.-sidorov.
    • The docs review found the й. case.
  • Side effect. The documented ǰ limit is gone: a second forced pass now leaves J̌o unchanged instead of giving J̌O, because the mark stays with its letter.

Verification

The differential gate can't see this, because capitalized() is not a compared surface. Coverage instead:

  • Corpus walk. test_a_decomposed_name_repairs_as_its_composed_twin re-encodes every corpus and case-table text that has accents, then compares its repair with the composed text's on both surfaces, plain and forced.
    • Recorded control: 144 disagreements over 137 texts with the old _render.py. A live control test re-breaks the code on purpose to show the walk can fail.
    • It sets aside 47 texts whose decomposed form parses differently. All are unspaced Hangul, which is matched as written by a documented decision, and a test checks that only Hangul is ever set aside.
  • Unit tests. test_render.py covers the shapes no corpus name reaches: Mac, the two mask shapes, the hyphenated й., the spliced-in й., the ǰ repeat-pass case, and a leading mark. Each was confirmed to fail without the change; the mask rows were also checked against a version that ignores marks.
  • Rule examples. rules.md#R4 gets two NFD example lines. Their expected values spell the accent as ́, so it's visible that the output stayed decomposed, and an editor that silently converts the file to NFC would make them fail. corpus_rules.jsonl is regenerated to carry them.
  • Checks run: the full suite (11059 passed), mypy, ruff, the Sphinx and README doctests, and the differential gate at 1.4.0 / 2.0.0 / 2.1.0 / 2.2.0 / 2.3.0, all exit 0. The two new NFD corpus names produce no diff.
  • Design-docs review. Two rounds. The first found the й. gap, and before-fix outputs that had been copied from an intermediate state. The second found three wording errors. All are fixed.

Out of scope, recorded in decisions.md#R4

  • initials() takes a token's first character, so for a decomposed initial it drops the accent: émile zola NFD gives e. z.. That's a rules.md#R3 view, not case repair, so it gets its own issue.

🤖 Generated with Claude Code

derek73 and others added 4 commits October 2, 2026 15:10
…ase repair

A name typed decomposed (NFD) was split at each combining mark, since
_WORD's \w matches no category-M character: 'josé garcía' repaired to
'José GarcíA'. Every clause now reads a mark as part of its letter:

- _sub_words runs a _WORD match on through the marks after it
- the Mac/Mc clause is decided on the composed spelling and applied to
  the word as written ('macée' gave 'MacÉe' composed, 'Macée' NFD)
- _apply_mask's split-off-initial test and _letter_run_ge2 read a
  letter's neighbours past marks (_beside), two shapes the issue
  expected to need no change

The output keeps the form it was typed in. A mark with no letter
before it heads no word. As a side effect 'ǰ' (which upper-cases to
J + combining caron) is now a fixpoint under a second forced pass.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
rules.md#R4 gains the statement and two NFD example lines (the
expected values spell the mark as ́, so the form is visible and
an editor that NFC-normalizes the file fails the example);
corpus_rules.jsonl is regenerated to carry them. decisions.md#R4
records the choice, the declined NFC-compose option, the mask shapes
measured against the issue's claim, and retires the 2026-09-24
SPLITTING boundary. Release log and usage.rst's decomposed-text
section follow.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Docs review findings. The hyphen clause's _DOTTED_INITIAL and the
unclassified-text fallback's _INITIAL matched the raw word, so a
decomposed 'й.' (й being the one default conjunction that decomposes)
read as the conjunction where its composed twin reads as an initial:
'ivan petrov-й.-sidorov' repaired to 'Petrov-й.-Sidorov' NFD against
'Petrov-Й.-Sidorov'. Both now match the NFC spelling.

decisions.md#R4's #542 bullet quoted outputs from a splitter-only
intermediate as pre-fix behavior; it now gives both, says the 'é.x.'
mask row agreed before the fix, drops the false "no corpus name is
NFD", and records initials() dropping a decomposed accent as out of
reach. _lexicon's reason for storing mask values composed is updated.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Second docs review: the R4 bullet and capitalized()'s docstring said
every clause reads a mark as part of its letter, while the three
shape clauses ask about the composed spelling instead; both now name
the two mechanisms and the one case where they differ. The clause
count and the gate's compared surfaces are corrected.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@derek73 derek73 added this to the 2.4 milestone Oct 2, 2026
@derek73 derek73 added the bug label Oct 2, 2026
@derek73 derek73 self-assigned this Oct 2, 2026
@codecov

codecov Bot commented Oct 2, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 98.97%. Comparing base (97af1f0) to head (81b0481).

Additional details and impacted files
@@           Coverage Diff           @@
##           master     #584   +/-   ##
=======================================
  Coverage   98.97%   98.97%           
=======================================
  Files          45       45           
  Lines        4181     4210   +29     
=======================================
+ Hits         4138     4167   +29     
  Misses         43       43           

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@derek73
derek73 merged commit 10cf228 into master Oct 3, 2026
11 checks passed
@derek73
derek73 deleted the fix/issue-542-nfd-case-repair branch October 3, 2026 02:31
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

Projects

None yet

Development

Successfully merging this pull request may close these issues.

capitalize() breaks a decomposed (NFD) word at its combining mark: josé garcía typed NFD gives José GarcíA

1 participant