Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
1 change: 1 addition & 0 deletions docs/design/decisions.md
Original file line number Diff line number Diff line change
Expand Up @@ -1442,6 +1442,7 @@ Accepted costs, deferred to the rescoped #459 rather than relitigated here: the
- 2026-09-24 #478 — DECIDED (Derek, 2026-09-24): the hyphen clause keeps its reading in a name written wholly in one case, and the case where that disagrees with rules.md#P3 is a recorded boundary, not a change. The two collide on a marked letter. Spaced, a letter the vocabulary marks as reading both ways reads as an initial in a one-case name (P3), so `maria silva e sousa` repairs to `Maria Silva E Sousa`; hyphenated, the interior part reads as the connective whatever the name's case, so `maria silva-e-sousa` repairs to `Maria Silva-e-Sousa`. Derek's reasoning: hyphenating `Silva-e-Sousa` is the writer joining the surname on purpose, so the interior word is a connective by that act, whatever the one-case rule would say of the spaced form. THE COST, measured 2026-09-24: a one-case name whose hyphenated bare initials happen to spell a connective — `J-E-P DUPONT` gives `J-e-P Dupont`, where every release 1.4.0 through 2.3.0 and the parent 4d0680e6 gave `J-E-P Dupont`, and `JOHN A-Y-B SMITH` gives `John A-y-B Smith` (every release: `John A-Y-B Smith`). The period-marked spelling is unaffected (`j.-e.-p. dupont` keeps `J.-E.-P. Dupont`, the 2026-09-23 #478 bullet). `conjunctions_ambiguous`, P3's knob, does NOT reach a hyphenated word: under `Lexicon.default().add(conjunctions_ambiguous={"y"})` the spaced `JOSE ORTEGA Y GASSET` repairs to `Jose Ortega Y Gasset` (default lexicon: `Jose Ortega y Gasset`) while `JOSE ORTEGA-Y-GASSET` stays `Jose Ortega-y-Gasset` under both. rules.md#R4's hyphen sentence said the hyphens join the name "as the spaced connective would", which claimed an equality the one-case case breaks; it now says the hyphens are the writer's join and states the split, with an Accepted paragraph and the `J-E-P DUPONT` boundary line. Pinned by `tests/v2/test_render.py::test_the_hyphen_is_the_writers_join_even_in_a_one_case_name`.
- 2026-10-02 #541 — THE "LEFT FOR THE ORCHESTRATOR TO FILE" FINDING of the 2026-09-24 branch-review bullet (d) above is #541, and it was wider than that bullet saw: besides the two raises it names, an entry carrying whitespace v1 never matched SILENTLY ACTIVATED in every set field the shim copies through and in `capitalization_exceptions` keys — measured 2026-10-02 against 1.4.0 (run outside the worktree, `PYTHONSAFEPATH=1`): `titles ' dean '` gave title `dean` on `dean john smith` where 1.4.0 gave first `dean`, and likewise `prefixes`, `suffix_acronyms`, `suffix_not_acronyms`, `conjunctions`, `bound_first_names` and a `' zzc '` key; the raise (`suffix_not_acronyms 'ma '`) and the activation (`titles ' dean '`) both reproduce on the 2.0.0, 2.2.0 and 2.3.0 wheels. DECIDED (Derek): `Constants._snapshot()` drops every entry `_config_shim._v1_matchable` rejects — empty after `lc()`, or not equal to its own single-spaced re-join — from all nine set fields and the keys, BEFORE its set algebra, and names them in one `UserWarning` whose remedy runs (`tests/v2/test_config_shim.py::test_the_offered_remedy_runs_and_silences_the_warning`). The test is exact because a v1 piece comes from a whitespace split re-joined only with single spaces; it generalizes the filter `given_name_titles` already had, and `tests/v2/test_config_shim.py::_UNFILTERED_OUTCOME` records what each swept row does with it off. One exception, decided when the 1.4 pickle test caught it: every release from at least 0.5.8 through 1.4.0 SHIPPED two such entries in `TITLES` (`'actor '`, `'television '`), so a restored 1.4 pickle carries them and its user never wrote them; they are dropped WITHOUT the warning (`_V14_SHIPPED_UNMATCHABLE`, held to the pickle's own unmatchable set by `test_the_1_4_shipped_unmatchable_roster_is_exactly_the_pickles`), keeping that pickle warning-free. The exemption is keyed by entry, not by provenance, so a user who writes `'actor '` or `'television '` themselves is not warned either — accepted, since the two strings are 1.x's own typos. The drop is a reading change for that pickle: `Actor John Smith` gave title `Actor` on 2.0.0 through 2.3.0 and gives first `Actor` now, as 1.4.0 did. DECLINED: stripping whitespace in `SetManager` (activates every v1-inert entry above), and a silent drop for user entries (Derek: the entry is certainly a typo and certainly dead, so say so). ACCEPTED WIDENING, deliberately left: a key written with capitals or edge periods (`'McDonald'`, `'phd.'`) never matched on 1.4.0, whose lookup is `lc(word)` against keys stored as written, but has matched since 2.0 through `Lexicon`'s fold; dropping it would break a working 2.x config to restore an inertness nobody wanted. Also left: a set entry `Lexicon` folds to empty that `lc()` does not, a lone non-ASCII full stop (`'。'`), still raises at the first parse in every field `Lexicon` receives directly — `titles`, `prefixes`, `suffix_acronyms`, `suffix_not_acronyms`, `conjunctions`, `bound_first_names` and a `capitalization_exceptions` key; `first_name_titles` drops it quietly (`if t`), and `non_first_name_prefixes` and `suffix_acronyms_ambiguous` never pass it to `Lexicon` at all — v1-matchable, so outside this rule.
- 2026-10-02 #582 — SUPERSEDES the #541 bullet's "Also left" sentence above. DECIDED (Derek): an entry `Lexicon._normalize` folds to empty is dropped by `_config_shim._v1_matchable` with the #541 warning, in every set field and key. The case is a lone CJK full stop (`'。'`, `'.'`, `'。'`, or a run of them), which 2.3.0's `FULL_STOPS` fold (#322/#323) empties while v1's `lc()` strips only `.`. Measured 2026-10-02 on released wheels run outside the checkout with `PYTHONSAFEPATH=1`: 1.4.0 through 2.2.0 accept `c.titles.add('。')` and read `。 john smith` with title `。`, while 2.3.0 raises `ValueError` at the first parse, as did this tree before the change, in `titles`, `prefixes`, `suffix_acronyms`, `suffix_not_acronyms`, `conjunctions`, `bound_first_names` and a key. ACCEPTED DEPARTURE from 1.4.0, the only one this filter makes: such an entry used to act on a token that is nothing but that full stop. Since 2.3.0 `Lexicon` can neither hold the entry nor match the token, its lookup fold emptying both, so reproducing v1 is not on offer and the choice was drop or raise. The warning's wording moved with it, from "nameparser 1.x never matched" to "match no name word". `given_name_titles`' `if t` filter was deleted as unreachable: `_title_key` is empty only when every word folds away, and then the whole entry folds to empty and was already dropped. With no shim-reachable plain `ValueError` left from `_normpairs`, `test_an_unrelated_capitalization_exceptions_valueerror_has_no_v1_hint` turns the filter off to keep the except-by-type clause pinned. THE SAME FOLD, ONE CHARACTER IN: an entry with a CJK full stop at its EDGE (not only full stops) passed the filter, and the set algebra compared it raw while `Lexicon` folded it, so `c.suffix_not_acronyms.add('ma。')` (beside the ambiguous `ma`) raised the gate-bypass check and a bound `'zed。'` beside a never-given particle `zed` raised the contradiction check, both from 2.3.0 on; measured accepted on 1.4.0 and 2.2.0 by the review of this change. `_build_snapshot` now compares `_normalize`d spellings in exactly the two computations `Lexicon` re-checks after its own fold — dropping a `suffix_not_acronyms` entry whose fold is an ambiguous acronym's, and taking into `particles_ambiguous` every particle whose fold is a bound given name's — so such an entry reads exactly as its folded spelling (`test_an_edge_full_stop_collision_reads_as_its_folded_spelling`; the fold-off control is `_UNFOLDED_OUTCOME`). That reading is the `dean。` widening below, not v1's: 1.4.0 matched `ma。` only on a `ma。` token. DECLINED, after being written and reviewed: folding EVERY kept entry before the set algebra. `non_first_name_prefixes` and `suffix_acronyms_ambiguous` reach `Lexicon` only through raw set algebra, so folding them is a NEW match, not the 2.3.0 one: `prefixes zz` with `non_first_name_prefixes 'zz。'` moved `zz smith` from first `zz` to last `zz smith`, an NFD/NFC pair across those fields moved the same way, and `suffix_acronyms_ambiguous 'zq。'` beside `suffix_acronyms zq` moved `john zq` from suffix to last — the particle readings the same on 1.4.0, 2.2.0, 2.3.0 and the pre-fold tree, the `zq` one on 2.2.0, 2.3.0 and the pre-fold tree (1.4.0 gives last `zq` there by its own two-word reading, not by matching the entry). They are pinned at those readings by `test_an_edge_full_stop_entry_outside_the_two_checks_reads_as_before`. ACCEPTED WIDENING, left as it is (Derek): since 2.3.0 a `titles` entry `'dean。'` also matches a bare `dean`, where 1.4.0 matched only `dean。`. The same fold makes it, and dropping the entry would lose the `dean。` match too, so it sits beside the capitalized-key widening above.
- 2026-10-02 #542 — A DECOMPOSED WORD IS REPAIRED AS ITS COMPOSED TWIN, and the repaired text keeps the form it was typed in: no clause composes its OUTPUT. Two mechanisms do it. The word splitter and the mask's neighbour tests read a combining mark (Unicode category M) as part of the letter before it; the three clauses that ask a regex about a whole word's SHAPE — Mac/Mc and the two initial tests — ask it of the word's composed spelling. They differ only for a mark with no precomposed form, where composing leaves the word as written and so agrees with the parse: under a caller-added conjunction `q́`, `ivan q́. petrov` parses `q́.` as the conjunction, and the hyphen clause keeps `Petrov-q́.-Sidorov` lowercase as it does (measured by the docs review, 2026-10-02). A name typed NFD (macOS file names, some databases) had been split at each mark by `_WORD`'s `(\w|\.)+`, `\w` matching no mark, so `josé garcía` repaired to `José GarcíA` on every release from 1.4.0. Python's `re` has no `\p{M}`, and a hand-written class of the combining blocks would be a copy of Unicode data that drifts (it would already miss Adlam, a cased script whose marks sit outside them), so the test is `unicodedata.category(ch)[0] == "M"`, asked by `_render._sub_words` after each `_WORD` match and by `_render._beside` for a letter's neighbours. DECLINED: NFC-composing before repair, the issue's option 2. It fixes the split but hands back text in a form the writer did not use, which is a change beyond case under rules.md#R4 — `tests/v2/test_properties.py::test_case_repair_changes_case_and_nothing_else` compares with `casefold()`, which does not normalize, and would report it. THE ISSUE'S CLAIM THAT THE MASK NEEDED NO CHANGE ONCE THE WORD REACHED IT WHOLE WAS MEASURED FALSE, and the Mac/Mc clause needed one too; both were found by fixing the splitter alone and comparing again. `_apply_mask`'s split-off-initial override and `_letter_run_ge2` read a letter's neighbour by raw index, which in NFD is the mark rather than the full stop or the letter past it: with the splitter alone, under a caller's `('éx', 'éx')` mask `john smith é.x.` forced gave `É.X.` composed and `é.X.` decomposed — a shape that AGREED before the fix (`É.X.` both ways) only because the old splitter cut the word at its mark before the mask was looked up — and under `('éex', 'éeX')` `john smith ée.x` gave `ée.X` against `éE.X` (`ÉE.X` before the fix). Both read past marks now. The Mac/Mc clause's `\w{2,}` fails at the mark, so with the splitter alone `macée` gave `Macée` decomposed against `MacÉe` composed (`MacéE` before the fix, the split having made the last `e` a word of its own); the clause is now DECIDED on the composed spelling and applied to the word as written. The two initial-shape tests, `_DOTTED_INITIAL` in the hyphen clause and `_INITIAL` in the unclassified-text fallback, matched the raw word and now match its composed spelling, so a decomposed `й.` — `й` is the one default conjunction that decomposes — is the initial its composed twin is: `ivan petrov-й.-sidorov` repaired to `Ivan Petrov-й.-Sidorov` decomposed against `Ivan Petrov-Й.-Sidorov` composed, and a middle spliced in as decomposed `й.` kept `й.` against `Й.`, at 97af1f02 and at this change's first commit alike (found by the docs review, 2026-10-02). AMENDS the 2026-09-24 branch-review bullet (a): its SPLITTING half is retired. `parse("ǰo smith").capitalized(force=True)` gives `J̌o`, and a second forced pass now gives `J̌o` back rather than `J̌O`, the mark staying with its letter (measured 2026-10-02); the LENGTHENING half (`ß`, `ʼn`) stands. A mark with no letter before it heads no word: `str.capitalize()` on a word beginning with one would upper-case the mark and leave the letter after it lowercase. Out of reach and left there, two of them: `initials()`, a rules.md#R3 view rather than case repair, takes a token's first character, which for a decomposed initial letter is the base letter without its accent (`parse("émile zola")` initials `é. z.` composed and `e. z.` decomposed, measured 2026-10-02; open as #585); and unspaced hangul, whose NFD parse differs from its NFC parse because segmentation matches the census list as written (docs/usage.rst, "Decomposed text"), so repair follows a different parse; hangul jamo are letters, not marks, and nothing here touches them. VERIFICATION: the differential gate cannot see this — `capitalized()` is not a compared surface (this entry, 2026-08-29); the two NFD rows this change adds to the rules corpus reach the gate through the role fields, `_ambiguities` and `_initials` only. Standing in its place: rules.md#R4's two NFD example lines, `tests/v2/test_properties.py::test_a_decomposed_name_repairs_as_its_composed_twin`, which decomposes every corpus and case-table text that decomposes, compares its repair with the composed text's on both surfaces plain and forced, and records as its control 144 disagreements over 137 texts with 97af1f02's `_render.py` over this change's corpus (measured 2026-10-02), and `tests/v2/test_render.py::test_a_decomposed_word_repairs_as_its_composed_twin` and `test_a_spliced_decomposed_initial_reads_as_its_composed_twin` for the shapes no corpus name reaches, whose two mask rows were mutation-checked by restoring raw-index neighbours.

Excluded (CAPITALIZATION_EXCEPTIONS — meng, edd, lac, ded, left out of the 2026-09-24 masks, #459):

Expand Down
8 changes: 7 additions & 1 deletion docs/design/rules.md
Original file line number Diff line number Diff line change
Expand Up @@ -2574,7 +2574,11 @@ R4. Rationale: case repair is a display concern, applied only on
lowercase there like any other connective — the third part of a
comma form is the shape that puts one there.
Repair changes case and nothing else: the repaired word is the
word as written, recased. A vocabulary entry records casing as a
word as written, recased. A word is read whole however its letters
are encoded: a letter written as a base letter followed by its
combining accent is repaired as the same letter written whole,
and the repaired word keeps the encoding the writer used. A
vocabulary entry records casing as a
mask — its word's letters, each in the case it takes — and repair
lays it over the word as the writer punctuated it, so the one
entry for phd repairs phd to PhD and ph.d. to Ph.D.; the mask
Expand Down Expand Up @@ -2661,6 +2665,8 @@ R4. Rationale: case repair is a display concern, applied only on
"Smith, John, and" → capitalized_forced="John Smith and"
"Doe, Jane, and Jr." → capitalized_forced="Jane Doe and Jr."
"juan de la vega" → capitalized="Juan de la Vega" · boundary
"josé garcía" → capitalized="Jose\u0301 Garci\u0301a"
"JOSÉ GARCÍA" → capitalized="Jose\u0301 Garci\u0301a"
"jose ortega-y-gasset" → capitalized="Jose Ortega-y-Gasset"
"JOSE ORTEGA-Y-GASSET" → capitalized="Jose Ortega-y-Gasset"
"maria silva-e-sousa" → capitalized="Maria Silva-e-Sousa"
Expand Down
2 changes: 2 additions & 0 deletions docs/release_log.rst
Original file line number Diff line number Diff line change
Expand Up @@ -50,6 +50,8 @@ Release Log

- **Change case repair to write an unlisted dotted credential and a roman numeral past iv in capitals.** ``HumanName("john smith x.y.z.").capitalize()`` gives ``John Smith X.Y.Z.`` where every release gave ``John Smith X.y.z.``, the dotted word being a suffix now (the ``unlisted_dotted_suffixes`` change above) and repaired as a listed acronym is; and ``john smith vi`` gives ``John Smith VI`` where every release gave ``John Smith Vi``, with ``vii``, ``viii`` and ``ix`` alike. Both are keyed on the suffix role: ``Jack X.Y.Z.``, which keeps its surname, still repairs as a name word (``Jack X.y.z.`` under ``force=True``), and ``john smith xi`` still gives ``John Smith Xi``, the parser reading ``xi`` as the surname. An unlisted dotted credential written in mixed case is kept as written on the default path, by the suffix change above (``john smith B.Tech.`` gives ``John Smith B.Tech.``), and reads all capitals under ``force=True`` (``John Smith B.TECH.``, where every release gave ``John Smith B.tech.``), which a ``capitalization_exceptions`` mask such as ``{"btech": "BTech"}`` undoes. Over the 1340 names in the differential corpora at the commit before this change (2026-09-23), 1 moves on the default path and 14 under ``force=True``. The recipe is the ``R4`` entry's 2026-09-23 MEASURED bullet in ``docs/design/decisions.md`` (#459)

- **Fix case repair breaking a name typed with decomposed accents at each accent.** ``HumanName(unicodedata.normalize("NFD", "josé garcía")).capitalize()`` gives ``José García``, where every release from 1.4.0 through 2.3.0 gave ``José GarcíA``: decomposed text (NFD, which macOS file names and some databases hand back) writes ``í`` as ``i`` followed by a combining accent, and the letters after the accent were repaired as a separate word. A decomposed name now repairs the way its composed spelling does, Mac/Mc names and case-repair masks included, and the output keeps the form it was typed in. See the ``R4`` entry of ``docs/design/decisions.md`` (closes #542)

- **Change the parse pipeline to copy its state without dataclasses.replace.** Every stage returns a copy of its frozen state, and several also copy tokens one at a time; ``dataclasses.replace`` goes through ``fields()`` and ``__init__`` on every one of those copies. The stages now copy fields directly through a small helper that is limited to the pipeline's own three dataclasses and checks them at import. One parse of the benchmark's reference name makes 36 fewer calls on py3.11 and 3.12 and 54 fewer from 3.13 (on 3.11, ``parse`` 406 to 370 and ``HumanName`` 443 to 407), and the call-count baselines move with them. Recomputable with ``uv run python tools/perf/call_count.py --against e0f1a2f``; the counts for every interpreter are in the ``parse-cost`` entry of ``docs/design/decisions.md``. No user-visible behavior changes (#546)

- **Fix a long given part after a family comma costing quadratic time.** Since 2.3.0, parsing ``"Doe, Jane " + "Smith " * n`` took time growing with the square of the part's length: going from 1,600 to 6,400 words cost 8.4x the time, where 2.2.0 and the comma-less form cost 4x. A run of trailing titles in the same place (``Doe, Jane Smith Prof. Prof. ...``) cost 10x. Both cost 4x again (Python 3.11, measured 2026-09-28). No field moves (closes #553)
Expand Down
4 changes: 3 additions & 1 deletion docs/usage.rst
Original file line number Diff line number Diff line change
Expand Up @@ -334,7 +334,9 @@ decomposed name gets the same order rule as its composed twin.
Vocabulary lookup does the same before matching a word against
titles, honorifics and the rest, so a decomposed ``Señor`` or ``née``
— macOS-origin data again — is recognized as readily as its composed
spelling.
spelling. Case repair reads a decomposed accent as part of its letter,
so ``capitalized()`` gives a decomposed ``josé garcía`` the same
``José García`` as the composed spelling, still decomposed.

Splitting is the exception. An unspaced decomposed hangul name is
ordered correctly but not split, because surname matching runs against
Expand Down
Loading
Loading