From 6a2ab7c151b9a9fe2f17b39b7fb415df79bbf25f Mon Sep 17 00:00:00 2001 From: Derek Gulbranson Date: Sat, 3 Oct 2026 13:50:11 -0700 Subject: [PATCH 1/9] fix: read halfwidth katakana as katakana (#594) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The script table now classifies the halfwidth kana block U+FF65-U+FF9F as Script.KATAKANA, so 山田 タロウ takes the kana license and reads family 山田 as 山田 エミ does. Everything keyed on the table follows: the second-East-Asian-word count that keeps a segmenter from re-dividing a kanji name, the 间隔号's flank guard, and _NO_INITIALS, which closes the title misroute decisions.md#cjk-full-stops recorded (タナカ. John). A range rather than a fold: T1 forbids rewriting the text, and an NFKC fold at classification reaches far past the kana and maps a lone ゙ into the hiragana block. Wholly-katakana names, halfwidth included, keep the declared order; W4's rationale is amended to say the script cannot tell a Japanese reading from a transcription, not that the latter dominates. U+FF65 leaves _SANCTIONED_EXTRAS for the table, the ledgers' span copies widen to match, and fix(#594) rules classify the five movers at every baseline (no corpus line held halfwidth kana before the new case rows). Co-Authored-By: Claude Opus 5.5 --- AGENTS.md | 2 +- docs/design/decisions.md | 3 +- docs/design/rules.md | 11 ++-- nameparser/_pipeline/_vocab.py | 6 +- nameparser/_policy.py | 38 +++++++---- nameparser/locales/ja.py | 4 +- tests/v2/cases.py | 51 +++++++++++++++ tests/v2/pipeline/test_script_segment.py | 5 +- tests/v2/pipeline/test_vocab.py | 1 + tests/v2/test_ledger_guards.py | 59 +++++++++++++---- tools/differential/corpus_cjk.jsonl | 4 ++ tools/differential/corpus_cjk_tolerated.jsonl | 2 + tools/differential/corpus_rules.jsonl | 2 + tools/differential/expected_since_1.4.0.toml | 65 ++++++++++++++----- tools/differential/expected_since_2.0.0.toml | 39 ++++++++++- tools/differential/expected_since_2.1.0.toml | 38 +++++++++++ tools/differential/expected_since_2.2.0.toml | 38 +++++++++++ tools/differential/expected_since_2.3.0.toml | 40 ++++++++++++ 18 files changed, 355 insertions(+), 53 deletions(-) diff --git a/AGENTS.md b/AGENTS.md index e04b1aea..5cf3e3d6 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -263,7 +263,7 @@ logging.getLogger('HumanName').setLevel(logging.DEBUG) The library has two layers: `nameparser/config/` (data) and `nameparser/parser.py` (logic). -**Design philosophy — positional and language-agnostic.** The parser assigns parts by *position* plus small sets of words that join to neighbors; it never detects language. A name's language can't be reliably inferred from Latin-script transliteration ("Ali" is Arabic or Italian; "Van"/"Della"/"Bin" are first names in some cultures, particles in others), so language-specific rules belong in opt-in `Constants` config, never global defaults. Many "wrong for language X" reports (#133, #150, #130, #85, #103, #146, #83) are irreducible ambiguities — e.g. `de Mesnil` (want last name) vs `Van Johnson` (want first name) are the same `[prefix][word]` shape. Before adding a rule, confirm it doesn't break the opposite case (run the full suite — Portuguese and "Van Johnson" tests are the usual canaries). **The one scoped exception (2.1, #271/#272): script-conditional behavior is permitted exactly where the SCRIPT ITSELF — not statistics about it — determines the convention.** The never-detect-language rule above is about Latin *transliteration*, where the signal genuinely is destroyed; native script is a different question, and it is answered per behavior rather than per script. Five defaults fall out of it, plus a sixth that applies the same not-a-guess standard to specific WORDS rather than to a script (#308's honorific peel, below). Wholly-Han, wholly-Hangul and kana-licensed names read family-first (`Policy.script_orders`) — no language detection needed, because zh and ja both write family-first in native script, so order cannot be misread even though the language is unknowable. Unspaced hangul splits into surname + given name (`config/surnames.py` ships the Korean census list as DEFAULT vocabulary) — nothing but Korean is written in hangul and the surnames are a closed census set, and the vocabulary is self-selecting besides: a hangul entry can only ever match hangul text. Hiragana licenses Japanese (#272) — a name whose characters stay inside Han∪kana while carrying at least one kana cannot be Chinese (the kana rules it out) and is not a transcription (foreign names are transcribed in katakana ALONE, マイケル has no kanji), so 高橋みなみ and 山田 エミ read family-first too; mechanically they resolve to the HIRAGANA entry, the license's carrier key. PURE katakana is excluded and keeps the positional default: マイケル・ジャクソン is a transcribed foreign name in its source order. And the 间隔号 U+00B7 (#298) is the transcription marker for scripts that HAVE no transcription script: a name it divides (威廉·莎士比亚 — flanked by classified characters on both sides, so Catalan's Gal·la is untouched) keeps its source order and never segments — the orthography names the convention, exactly as pure katakana does, with the divider carrying the signal instead of the script. And a listed CJK honorific glued to the END of a name token is split off it (#308) — 田中さん is 田中 plus さん — on the same orthography-settles-it test, narrowed for the glued position: an entry peels only where it can never end a name, so 씨/님/さん/様/先生 peel while 양/군/氏/博士/殿 stay spaced-only (김지양 and 田中博士 are names, and ~90 Japanese surnames end in 殿) and 君 is in NEITHER set (王君 is a complete Chinese name), though its kana spelling くん peels. Like the nakaguro's tokenize-level separation described next, it is reached by neither policy opt-out — but for its own reason: the vocabulary carries the license itself rather than borrowing the script's, so `segment_scripts` has nothing to say about it. Since #312 it also crosses the 间隔号, which still stops the surname split standing right beside it: it answers where a name DIVIDES into surname and given, and the peel never asks that question. Whether it also crosses the FAMILY comma is tolerated rather than settled: the 2026-09-01 demotion (rules.md#W3) narrowed that half from contract to best-effort, since no CJK writing system's own convention puts a comma between family and given at all — so `김, 민준씨` reads today exactly as the spaced `김 민준씨` does (family 김, given 민준, suffix 씨) while the split stands down as before, but that reading is watched on the differential's radar tier rather than pinned as contract. Its site is accordingly the name-bearing segment runs — `segments[:2]` under a family comma, and `segments[0]` as before otherwise, the family comma being the one structure that splits the name itself across two runs, with the honorific as often glued to the given side as to the family. That is the whole reach and nothing past it (`김, 민준 지훈씨` peels; `김, 민준, 지훈씨` and `김,, 민준씨` do not, both landing in a third run), and whether `segments[1]` is name text at all is now ASKED rather than inferred from the structure — `segment` does not guarantee it, since a one-word part before the comma reads as FAMILY_COMMA even when the part after it is entirely suffix-shaped, and the peel walking into such a run took `V.` for its site, found no listed tail and abandoned (#319). The question is `segment`'s own suffix-comma predicate, lifted into `_vocab.is_wholly_suffix` so the two stages cannot drift: a wholly suffix-shaped second run is declined and the scan stays in `segments[0]`, so `田中さん, V.` and `田中さん, Ph. D.` give さん up as `田中さん, PhD` always did. The test is necessary but NOT sufficient, and the second condition is not decoration: every honorific tail is also a suffix word, so a glued honorific is itself part of what makes its run read as suffix-shaped, and declining a run that holds the ONLY site loses the peel outright. `segments[0]` must therefore offer a peel site of its own before the second run is declined — `이, J.씨` and `선생님, J.씨` pass the suffix test and are scanned anyway, keeping the pre-#319 reading, while `김민준씨, J.씨` has a site on both sides and peels the person's own 씨 rather than the junk one behind the comma. Uniform in the PEEL, that is — where the credential itself lands is `assign`'s question and still differs by spelling (`V.` → `given`, `PhD` and `Ph. D.` → `suffix`). Not `_is_post_nominal` pluralized: the run predicate says yes both to what the token predicate vetoes (`V.`, `V`, `I` — the class the defect was reported as) and to what the token predicate never sees at all, since `period_joined_vocab` and the delimiter routes are the run predicate's alone (`Msc.Ed.` and `J.씨` reach it that way, and `田中さん, Msc.Ed.` moves with the rest). `Policy(lenient_comma_suffixes=False)` drops this call to the strict token test too — so those three read as name text again and keep the pre-#319 answer, while `Ph. D.` peels under the knob regardless, its merged `phd` passing the strict test. `田中さん, 太郎` is unchanged, and not because of its comma — the honorific there is not at the end of the name, 太郎 is. The nakaguro belongs to the same doctrine but is decided a layer down: U+30FB and its halfwidth twin U+FF65 separate tokens like whitespace, unconditionally and in tokenize, so neither policy opt-out (`script_orders={}`, `segment_scripts=()`) reaches it — the codepoints are CJK-only and appear in no other script's names, which is what licenses a tokenize-level rule where U+00B7 (also the Catalan punt volat, interior to Gal·la) needs the flanked-by-classified-script guard `_tokenize_region` gives it (#298). Han segmentation stays OPT-IN (`locales.ZH` for Chinese, `locales.JA` for Japanese) — a zh surname list corrupts Japanese kanji names, since 高 is a common Chinese surname and 高橋一郎 would split 高+橋一郎 where the correct reading is 高橋+一郎; no surname list divides a kanji name at all, so `locales.JA` activates the stage and a pluggable `Parser(segmenter=...)` does the dividing. Latin-script input is never touched by any of this: "Kim Min-jun" is genuinely order-ambiguous and stays governed by `name_order` and opt-in packs. Before adding a script-conditional rule, work out which of the three it is — certain, certain for this one behavior only, or a statistical guess wearing a script's clothes. +**Design philosophy — positional and language-agnostic.** The parser assigns parts by *position* plus small sets of words that join to neighbors; it never detects language. A name's language can't be reliably inferred from Latin-script transliteration ("Ali" is Arabic or Italian; "Van"/"Della"/"Bin" are first names in some cultures, particles in others), so language-specific rules belong in opt-in `Constants` config, never global defaults. Many "wrong for language X" reports (#133, #150, #130, #85, #103, #146, #83) are irreducible ambiguities — e.g. `de Mesnil` (want last name) vs `Van Johnson` (want first name) are the same `[prefix][word]` shape. Before adding a rule, confirm it doesn't break the opposite case (run the full suite — Portuguese and "Van Johnson" tests are the usual canaries). **The one scoped exception (2.1, #271/#272): script-conditional behavior is permitted exactly where the SCRIPT ITSELF — not statistics about it — determines the convention.** The never-detect-language rule above is about Latin *transliteration*, where the signal genuinely is destroyed; native script is a different question, and it is answered per behavior rather than per script. Five defaults fall out of it, plus a sixth that applies the same not-a-guess standard to specific WORDS rather than to a script (#308's honorific peel, below). Wholly-Han, wholly-Hangul and kana-licensed names read family-first (`Policy.script_orders`) — no language detection needed, because zh and ja both write family-first in native script, so order cannot be misread even though the language is unknowable. Unspaced hangul splits into surname + given name (`config/surnames.py` ships the Korean census list as DEFAULT vocabulary) — nothing but Korean is written in hangul and the surnames are a closed census set, and the vocabulary is self-selecting besides: a hangul entry can only ever match hangul text. Hiragana licenses Japanese (#272) — a name whose characters stay inside Han∪kana while carrying at least one kana cannot be Chinese (the kana rules it out) and is not a transcription (foreign names are transcribed in katakana ALONE, マイケル has no kanji), so 高橋みなみ and 山田 エミ read family-first too; mechanically they resolve to the HIRAGANA entry, the license's carrier key. PURE katakana — full-width or halfwidth, the halfwidth block U+FF65–FF9F being katakana in the script table since #594 — is excluded and keeps the positional default: マイケル・ジャクソン is a transcribed foreign name in its source order, and a Japanese reading written in katakana (legacy halfwidth data writes every name that way) cannot be told from one, so the script settles nothing and a caller who knows opts in through `script_orders`. And the 间隔号 U+00B7 (#298) is the transcription marker for scripts that HAVE no transcription script: a name it divides (威廉·莎士比亚 — flanked by classified characters on both sides, so Catalan's Gal·la is untouched) keeps its source order and never segments — the orthography names the convention, exactly as pure katakana does, with the divider carrying the signal instead of the script. And a listed CJK honorific glued to the END of a name token is split off it (#308) — 田中さん is 田中 plus さん — on the same orthography-settles-it test, narrowed for the glued position: an entry peels only where it can never end a name, so 씨/님/さん/様/先生 peel while 양/군/氏/博士/殿 stay spaced-only (김지양 and 田中博士 are names, and ~90 Japanese surnames end in 殿) and 君 is in NEITHER set (王君 is a complete Chinese name), though its kana spelling くん peels. Like the nakaguro's tokenize-level separation described next, it is reached by neither policy opt-out — but for its own reason: the vocabulary carries the license itself rather than borrowing the script's, so `segment_scripts` has nothing to say about it. Since #312 it also crosses the 间隔号, which still stops the surname split standing right beside it: it answers where a name DIVIDES into surname and given, and the peel never asks that question. Whether it also crosses the FAMILY comma is tolerated rather than settled: the 2026-09-01 demotion (rules.md#W3) narrowed that half from contract to best-effort, since no CJK writing system's own convention puts a comma between family and given at all — so `김, 민준씨` reads today exactly as the spaced `김 민준씨` does (family 김, given 민준, suffix 씨) while the split stands down as before, but that reading is watched on the differential's radar tier rather than pinned as contract. Its site is accordingly the name-bearing segment runs — `segments[:2]` under a family comma, and `segments[0]` as before otherwise, the family comma being the one structure that splits the name itself across two runs, with the honorific as often glued to the given side as to the family. That is the whole reach and nothing past it (`김, 민준 지훈씨` peels; `김, 민준, 지훈씨` and `김,, 민준씨` do not, both landing in a third run), and whether `segments[1]` is name text at all is now ASKED rather than inferred from the structure — `segment` does not guarantee it, since a one-word part before the comma reads as FAMILY_COMMA even when the part after it is entirely suffix-shaped, and the peel walking into such a run took `V.` for its site, found no listed tail and abandoned (#319). The question is `segment`'s own suffix-comma predicate, lifted into `_vocab.is_wholly_suffix` so the two stages cannot drift: a wholly suffix-shaped second run is declined and the scan stays in `segments[0]`, so `田中さん, V.` and `田中さん, Ph. D.` give さん up as `田中さん, PhD` always did. The test is necessary but NOT sufficient, and the second condition is not decoration: every honorific tail is also a suffix word, so a glued honorific is itself part of what makes its run read as suffix-shaped, and declining a run that holds the ONLY site loses the peel outright. `segments[0]` must therefore offer a peel site of its own before the second run is declined — `이, J.씨` and `선생님, J.씨` pass the suffix test and are scanned anyway, keeping the pre-#319 reading, while `김민준씨, J.씨` has a site on both sides and peels the person's own 씨 rather than the junk one behind the comma. Uniform in the PEEL, that is — where the credential itself lands is `assign`'s question and still differs by spelling (`V.` → `given`, `PhD` and `Ph. D.` → `suffix`). Not `_is_post_nominal` pluralized: the run predicate says yes both to what the token predicate vetoes (`V.`, `V`, `I` — the class the defect was reported as) and to what the token predicate never sees at all, since `period_joined_vocab` and the delimiter routes are the run predicate's alone (`Msc.Ed.` and `J.씨` reach it that way, and `田中さん, Msc.Ed.` moves with the rest). `Policy(lenient_comma_suffixes=False)` drops this call to the strict token test too — so those three read as name text again and keep the pre-#319 answer, while `Ph. D.` peels under the knob regardless, its merged `phd` passing the strict test. `田中さん, 太郎` is unchanged, and not because of its comma — the honorific there is not at the end of the name, 太郎 is. The nakaguro belongs to the same doctrine but is decided a layer down: U+30FB and its halfwidth twin U+FF65 separate tokens like whitespace, unconditionally and in tokenize, so neither policy opt-out (`script_orders={}`, `segment_scripts=()`) reaches it — the codepoints are CJK-only and appear in no other script's names, which is what licenses a tokenize-level rule where U+00B7 (also the Catalan punt volat, interior to Gal·la) needs the flanked-by-classified-script guard `_tokenize_region` gives it (#298). Han segmentation stays OPT-IN (`locales.ZH` for Chinese, `locales.JA` for Japanese) — a zh surname list corrupts Japanese kanji names, since 高 is a common Chinese surname and 高橋一郎 would split 高+橋一郎 where the correct reading is 高橋+一郎; no surname list divides a kanji name at all, so `locales.JA` activates the stage and a pluggable `Parser(segmenter=...)` does the dividing. Latin-script input is never touched by any of this: "Kim Min-jun" is genuinely order-ambiguous and stays governed by `name_order` and opt-in packs. Before adding a script-conditional rule, work out which of the three it is — certain, certain for this one behavior only, or a statistical guess wearing a script's clothes. **A constant's membership is a question you may reopen.** Proposing that a word be ADDED, REMOVED or MOVED between vocabulary sets is ordinary design work — a shipped entry is not evidence that anyone judged it. `SUFFIX_ACRONYMS` arrived in `af5bdab` as a bulk Wikipedia import never reviewed against surname collisions: 572 of its 577 alphabetic entries leave `family` empty in `"John "` against five ambiguous-gated exceptions (recomputed 2026-09-07; the fifth is `ba`; 568 of 575 against seven, measured 2026-09-25 after #540), and `sa`, `se` and `om` are borne as surnames (measured 2026-08-23), as were `rai`, `cha`, `ba` and `mc` before `decisions.md#suffix-acronym-collisions` decided all four — `rai` and `cha` removed, `ba` marked ambiguous, `mc` left alone as no borne name at all. The same entry marked `meng` and `lac` ambiguous on 2026-09-25 (#540). When a fix starts to look like new machinery, check the vocabulary first. Criterion: `decisions.md#vocabulary-collisions`, with #360's positional qualifier. diff --git a/docs/design/decisions.md b/docs/design/decisions.md index e7793612..7843185f 100644 --- a/docs/design/decisions.md +++ b/docs/design/decisions.md @@ -468,6 +468,7 @@ Declined: - 2026-07-27 (script-scoped order amendment) — the family-first override is keyed to the SCRIPT of the written name, never to a guessed language: wholly-Han, wholly-Hangul, and kana-licensed Japanese read family-first because zh, ko and ja all write family-first in native script; wholly-katakana names are predominantly transcriptions and keep the declared order. Latin transliterations are never touched. - 2026-07-29 #272 — the kana license: Han∪kana with at least one kana cannot be Chinese and is not a transcription, so 高橋みなみ reads family-first though it is written in two scripts. +- 2026-10-03 #594 — HALFWIDTH KATAKANA IS KATAKANA. The script table classifies the halfwidth kana block U+FF65–FF9F as `Script.KATAKANA`, so `山田 タロウ` reads family 山田 as `山田 エミ` does; every release through 2.3.0 left the block unclassified and read it given 山田, family タロウ, the positional reading. Legacy data written for JIS X 0201 systems, which had no kanji, spells names this way — bank, payroll and CSV exports. DECIDED by widening the table rather than by folding: the T rules forbid rewriting the text (decisions.md#T1), so normalization at tokenize was never available, and an NFKC fold at classification — the other place the issue offered — reaches far past the kana (`JOHN` to `JOHN`, `㈱` to `(株)`) and maps a lone voicing mark `゙` (U+FF9E) to the combining U+3099, which sits in the HIRAGANA block, so the per-character interpunct flank guard would read it as hiragana. A block range is how the table already decides (U+30FB and the voicing marks are in by block), and it takes the voicing marks U+FF9E/U+FF9F in because they are spacing characters following their base. The span starts at U+FF65, the halfwidth nakaguro, which tokenize already treats as a separator (rules.md#T2), and stops short of U+FF61–FF64, whose U+FF61 is a full stop. Everything keyed on the table follows: the kana license, the second-East-Asian-word count that keeps a segmenter from being asked (`高橋一郎 タロウ`), the 间隔号's flank guard (`タロウ·ヤマダ` divides), and `_NO_INITIALS`, which closes the limit #cjk-full-stops recorded (`タナカ. John` read title `タナカ.`). THE RATIONALE FOR WHOLLY-KATAKANA NAMES IS AMENDED, NOT THE BEHAVIOR: a wholly-halfwidth name keeps the declared order exactly as a wholly-katakana one does (`ヤマダ タロウ` reads given ヤマダ), but the 2026-07-27 reason — that such names are predominantly transcriptions — is weaker here, since a system with no kanji wrote Japanese names in katakana too, family-first. What stands is the narrower claim: the script cannot tell a Japanese reading from a transcription, so it settles no order, and a caller who knows their katakana is Japanese adds `(Script.KATAKANA, FAMILY_FIRST)` to `script_orders` (docs/usage.rst's Boundaries section carries the recipe). `name_order=FAMILY_FIRST` works too, at the cost of reversing Latin names as well. MEASURED 2026-10-03: the differential gate saw nothing before this change because no corpus line held halfwidth kana; the six case rows #594 added are its whole population, three of them contract movers (`山田 タロウ`, `山田 ダイスケ`, `タロウ·ヤマダ`), two tolerated radar movers (`タナカ. John`, `マイケル.`) and one parity row (`ヤマダ タロウ`). RECOMPUTE: parse the six on this tree and on 2.3.0 from PyPI and read `title` beside the name fields. ### D1 — the segmenterless-activation warning @@ -1514,7 +1515,7 @@ ONE GAP SEEN FROM TWO SIDES, which is why the two issues were answered together. - **#323 turned out broader than it was filed.** The issue named the honorific peel. Making the classification fold read through an edge stop moved three readers of the `None` it used to return, and only one of them was the peel's neighbour test: the SURNAME SITE stepped past the family name onto the given name (`양. 지훈` cut 지훈 in half — given `양.`, middle `지`, family `훈` — and now reads given `지훈`, family `양.`), the ORDER RULE fell back to positional (`양 지훈.` lost family-first and now keeps it; `田中 太郎.` likewise), and the segmenter's neighbour precondition missed a writer-drawn boundary (`山田太郎 田中.` consulted a pluggable segmenter on `山田太郎` as if it stood alone, and is blocked now). The issue weighed two candidate fixes and chose neither: teaching the surname site to consult `is_suffix_strict` beside `effective_script`, or making the classification tolerant of a trailing period, which it called the broader one and expected to interact with #322. The broader one shipped, and the narrow one would have reached the filed name alone — `김.` is no suffix vocabulary, so `김. 민준` sits outside it entirely. - **The surname-site amendment was found by measurement during execution, not designed.** With the classification fold alone, `김. 민준` read given `김`, middle `.`, family `민준`: the token classified as hangul, the site matched `김` against its own head, and what was left over — the stop — BECAME the remainder. So the site matches on `text.rstrip(FULL_STOPS)` and cuts the core, a head being a prefix, which makes the offset that cuts the core cut the text; the stop rides with the remainder (`김민준.` divides as `김` + `민준.`, `김. 민준` as family `김.` plus given `민준`). `rstrip`, not `strip`: a LEADING stop would break the prefix argument. The reading is recorded in `ebf64db`'s comment at the site. - **Two readings JOINED the bundle, approved in session on 2026-09-09.** Neither was filed. FIRST, the glued-period honorific (rules.md#W3): a stop on the honorific's own word stood between the listed tail and the token's end, so nothing peeled and `田中さん.` and `김민준씨.` read as titles. The peel now matches the tail on the core and cuts BEFORE it, so both divide where their stop-less spellings do and the stop rides with the honorific. The move was free because W3 had recorded that reading as measured on 2026-09-05 and pinned by nothing, its own words being that neither string is a case row or a corpus line, so no row pinned those two readings; both are case rows now, tolerated. SECOND, H2's opening-abbreviation shape on a CJK word: `田中.` read title and now reads family `田中.`, `田中. 太郎` reads family `田中.`, given `太郎`. Han is what WITNESSES this one, and hangul cannot: hangul segmentation runs first, so `김민준.` never reaches the shape test and divides as `김` + `민준.` by the surname site instead. Latin and Cyrillic are untouched (`Smith. John` still reads title `Smith.`, `Проф. Иванов` title `Проф.`), and the veto is contains-any, so `Kim김. Smith` is refused a title too — deliberately: a word carrying a script with no abbreviations is not wearing an abbreviation's period. -- **Recorded limits, all five measured 2026-09-10 on this tree — and the segmenter one is a limit the review round CLOSED, kept in this bullet because this is where the claim stood.** HALFWIDTH KANA sits outside the script table, so it reads by the LATIN shapes at both period sites and not just at the initial one: `_vocab.is_initial("ラ.")` is `True` where `is_initial("ラ.")` is `False`, and H2's opening-abbreviation shape MISROUTES a longer halfwidth opener — `parse("タナカ. John")` gives title `タナカ.` where `parse("タナカ. John")` gives given `タナカ.`. Both because `_policy._SCRIPT_RANGES` excludes the halfwidth kana block U+FF65–FF9F on purpose — legacy bank and CSV data writes it, and the table's comment calls that a separate normalization problem, #272's separator handling touching only the halfwidth DOT U+FF65 and none of the kana behind it — and `_NO_INITIALS`, being a tuple of Script members, silently inherits the exclusion. That is #322's second mechanism, restated here because it is the one the bundle did NOT close. An earlier wording of this bullet said "nothing misroutes today, every kana honorific being multi-character", which had the direction backwards: the misroute is at the SHAPE site rather than the peel, and being multi-character is what carries a halfwidth opener INTO H2's shape rather than what keeps it out. NOTHING PINS EITHER READING — no case row, no corpus line, for the halfwidth block or for `タナカ. John` — so both can move without the suite or the differential saying so. It stays a recorded limit rather than a fix: what to do with halfwidth kana is a normalization question (widen the script table, or fold widths at classification) and not a full-stop one, and the bundle answered the full-stop one. RECOMPUTE: parse `タナカ. John` and `タナカ. John` and read `title` beside the name fields; `is_initial` alone reports only the shorter half. A PLUGGABLE SEGMENTER receives the CORE, the token with its trailing stops removed — the same string the vocabulary match reads, the surname site handing it `core` rather than `text` (corrected 2026-09-10, in the review round after this entry was first written: the site handed over the raw token, which this bullet recorded as a limit, and the fix costs nothing because a head is a prefix — every offset reported against the core cuts `text` in the same place, so the stops ride with the last piece and no answer of the segmenter's can make a stop a piece of its own). The shipped `ja` pack's `_wholly_japanese` therefore ACCEPTS `山田太郎.` now, where the raw token failed its repertoire test: measured 2026-09-10 with `namedivider-python` installed (`uv run --extra ja`), `parser_for(locales.JA, segmenter=locales.ja_segmenter())` reads `山田太郎.` as family 山田, given `太郎.` with a SEGMENTATION report scoring 0.44, where `d37b8ec` read title `山田太郎.` and the un-corrected branch read family `山田太郎.`; the already-divided `山田 太郎.` reads family 山田, given `太郎.` on this tree and given 山, middle 田, family `太郎.` at `d37b8ec`. Pinned by `test_a_consulted_segmenter_receives_the_core_without_the_stop`. A LEADING stop keeps the token out of the SURNAME site only, because the classification fold that gates it rstrips as it does: `.김민준` classifies as no script, never becomes a surname site and stays one whole word, reading given `.김민준` exactly as `d37b8ec` reads it (corrected 2026-09-10 in the same round: the fold stripped both edges when this entry was written, which moved the ROLE to family and — with a segmenter configured — offered the token to it raw, an answer of offset 1 dividing it into the stop and the name; the fold now matches the sites, and hiding such a token is the no-split rules.md#W1 already accepts). It does NOT keep the token out of the HONORIFIC PEEL, which consults no script at all — the tail alone is its license (rules.md#W2's Background) — so a leading stop does not defeat it: `.김민준씨` peels to `.김민준` and `씨` exactly as `김민준씨` does, pinned by `test_the_peel_reads_the_trailing_stop_only`. A BRACKETED CJK CREDENTIAL degrades: rules.md#S1's escape in `_extract` calls a clause suffix-shaped when it ends in an ASCII period, so `(김민준.) John Smith` is unwrapped and the H2 veto then leaves `김민준.` as name text the surname site divides — given 김, middle `민준. John`, family Smith, where 2.2.0 and `d37b8ec` read title `김민준.` — pinned as the tolerated `ko_name_with_a_period_in_a_bracketed_credential`. And a caller-authored NFD surname entry is composed at ingest, so raw NFD input then finds no entry and the name goes unsplit; pinned by `test_an_nfd_authored_entry_is_stored_composed`. +- **Recorded limits, all five measured 2026-09-10 on this tree — and the segmenter one is a limit the review round CLOSED, kept in this bullet because this is where the claim stood.** HALFWIDTH KANA sits outside the script table, so it reads by the LATIN shapes at both period sites and not just at the initial one: `_vocab.is_initial("ラ.")` is `True` where `is_initial("ラ.")` is `False`, and H2's opening-abbreviation shape MISROUTES a longer halfwidth opener — `parse("タナカ. John")` gives title `タナカ.` where `parse("タナカ. John")` gives given `タナカ.`. Both because `_policy._SCRIPT_RANGES` excludes the halfwidth kana block U+FF65–FF9F on purpose — legacy bank and CSV data writes it, and the table's comment calls that a separate normalization problem, #272's separator handling touching only the halfwidth DOT U+FF65 and none of the kana behind it — and `_NO_INITIALS`, being a tuple of Script members, silently inherits the exclusion. That is #322's second mechanism, restated here because it is the one the bundle did NOT close. An earlier wording of this bullet said "nothing misroutes today, every kana honorific being multi-character", which had the direction backwards: the misroute is at the SHAPE site rather than the peel, and being multi-character is what carries a halfwidth opener INTO H2's shape rather than what keeps it out. NOTHING PINS EITHER READING — no case row, no corpus line, for the halfwidth block or for `タナカ. John` — so both can move without the suite or the differential saying so. It stays a recorded limit rather than a fix: what to do with halfwidth kana is a normalization question (widen the script table, or fold widths at classification) and not a full-stop one, and the bundle answered the full-stop one. CLOSED 2026-10-03 by #594, which widened the table (decisions.md#W4 carries why not the fold): `タナカ. John` reads given `タナカ.` and `is_initial("ラ.")` is `False`, and both readings are pinned now — the case rows `ja_halfwidth_katakana_opener_with_a_period_is_a_name` and `ja_halfwidth_katakana_lone_name_with_a_period_is_not_a_title`, and `test_is_initial_script_repertoire`. RECOMPUTE: parse `タナカ. John` and `タナカ. John` and read `title` beside the name fields; `is_initial` alone reports only the shorter half. A PLUGGABLE SEGMENTER receives the CORE, the token with its trailing stops removed — the same string the vocabulary match reads, the surname site handing it `core` rather than `text` (corrected 2026-09-10, in the review round after this entry was first written: the site handed over the raw token, which this bullet recorded as a limit, and the fix costs nothing because a head is a prefix — every offset reported against the core cuts `text` in the same place, so the stops ride with the last piece and no answer of the segmenter's can make a stop a piece of its own). The shipped `ja` pack's `_wholly_japanese` therefore ACCEPTS `山田太郎.` now, where the raw token failed its repertoire test: measured 2026-09-10 with `namedivider-python` installed (`uv run --extra ja`), `parser_for(locales.JA, segmenter=locales.ja_segmenter())` reads `山田太郎.` as family 山田, given `太郎.` with a SEGMENTATION report scoring 0.44, where `d37b8ec` read title `山田太郎.` and the un-corrected branch read family `山田太郎.`; the already-divided `山田 太郎.` reads family 山田, given `太郎.` on this tree and given 山, middle 田, family `太郎.` at `d37b8ec`. Pinned by `test_a_consulted_segmenter_receives_the_core_without_the_stop`. A LEADING stop keeps the token out of the SURNAME site only, because the classification fold that gates it rstrips as it does: `.김민준` classifies as no script, never becomes a surname site and stays one whole word, reading given `.김민준` exactly as `d37b8ec` reads it (corrected 2026-09-10 in the same round: the fold stripped both edges when this entry was written, which moved the ROLE to family and — with a segmenter configured — offered the token to it raw, an answer of offset 1 dividing it into the stop and the name; the fold now matches the sites, and hiding such a token is the no-split rules.md#W1 already accepts). It does NOT keep the token out of the HONORIFIC PEEL, which consults no script at all — the tail alone is its license (rules.md#W2's Background) — so a leading stop does not defeat it: `.김민준씨` peels to `.김민준` and `씨` exactly as `김민준씨` does, pinned by `test_the_peel_reads_the_trailing_stop_only`. A BRACKETED CJK CREDENTIAL degrades: rules.md#S1's escape in `_extract` calls a clause suffix-shaped when it ends in an ASCII period, so `(김민준.) John Smith` is unwrapped and the H2 veto then leaves `김민준.` as name text the surname site divides — given 김, middle `민준. John`, family Smith, where 2.2.0 and `d37b8ec` read title `김민준.` — pinned as the tolerated `ko_name_with_a_period_in_a_bracketed_credential`. And a caller-authored NFD surname entry is composed at ingest, so raw NFD input then finds no entry and the name goes unsplit; pinned by `test_an_nfd_authored_entry_is_stored_composed`. - **Pickle compatibility, decided 2026-09-10 (Derek): a release-note line, not a forgiving load.** `Lexicon.__setstate__` re-runs `_normalize` over the loaded state and rejects any entry the fold changes, so a `Lexicon` pickled by 2.1.x or 2.2.x whose caller-added entries the widened fold now touches (a non-ASCII entry written with a wide stop such as `씨。`, or authored in NFD) raises `ValueError: incompatible Lexicon pickle: entries are not normalized (titles: 씨。); this state was not written by this version of nameparser` — a message that is wrong about provenance for exactly these pickles, which a released version did write. Measured 2026-09-10 (recompute: take `Lexicon.default().add(titles={"씨"}).__getstate__()`, add `씨。` to its `titles`, and load it into `Lexicon.__new__(Lexicon)`). The shipped vocabulary is unaffected (zero of 1735 strings change under NFC, and no shipped entry carries a wide stop). Two remedies were weighed: a release-note line with the fix a caller applies (rebuild the `Lexicon` from its source rather than unpickle it), or a load path that accepts state whose only drift is one fold pass away, re-folds it and warns naming the entries. The note was chosen: the load guard was designed to refuse rewritten data (the guarded-raise design #3-0-reevaluations records as right regardless of the in-a-minor friction) rather than become a fourth place caller data is corrected without a word, the population is caller-added CJK entries pickled across a minor, and a rebuild is the documented way to carry a `Lexicon` across versions. The 2.3.0 release note carries the line; `tests/v2/test_lexicon.py::test_unpickling_rejects_unnormalized_entries` is the pin of the refusal itself. - **CORRECTIONS this entry makes to entries above it.** #indic-honorifics' "Abbreviation marks and periods" bullet said `_normalize` "lowercases and strips edge whitespace and ASCII periods only — no NFC, no NFKC, no casefold anywhere in that module"; it now composes NFC for a non-ASCII word and strips all four stops. The VISARGA ARGUMENT still holds and holds for the same reason: ঃ U+0983 has no canonical decomposition, so NFC leaves it exactly where it was — recompute with that bullet's own one-liner, `Lexicon.default().add(titles={"মোঃ"}).titles & {"মোঃ"}`, still non-empty. #W1's 2026-07-29 §1a bullet said "segmentation MATCHING stays raw" and called the classification fold "the one deliberate exception"; matching is now folded too, so what stays raw is SEGMENTATION — the surname site's membership test and the peel's tail slice, which index the token's own text — and the folds are two, not one. #cjk-comma-demotion's 2026-09-05 five-parse bullet is amended in place: `田中さん 太郎.` no longer flips the order (the period is invisible to the classification, so it reads family `田中さん`, given `太郎.` as its no-period twin does), and `田中さん.` and `김민준씨.` no longer read as titles, so that bullet's summary — that the ONLY one of the five where a period changes nothing is the one where it sits on a word the split-off steps past — is superseded: the period changes nothing in four of the five now, and `田中さん 様.` is joined rather than alone. #initials-repertoire's principle bullet restated the veto phonologically ("morphemes or syllables"), which `_policy._NO_INITIALS`' own comment forbids in as many words — Devanagari is an abugida and Arabic an abjad, neither has letters in that sense and both abbreviate — so the bullet now states the CLDR criterion the constant actually uses. - **What moved, measured 2026-09-10 and re-measured the same day after the review round.** SEVENTEEN names in `tools/differential/corpus*.jsonl` read differently against `d37b8ec`, and all seventeen are the bundle's own new rows, every one of them on the RADAR tier in `corpus_cjk_tolerated.jsonl`. Twelve are the CJK movers this bundle's landing commits wrote; the thirteenth, the wholly-katakana `マイケル.`, joined in the whole-branch review that produced this entry, reading given `マイケル.` where `d37b8ec` reads title `マイケル.` — the katakana arm of the same H2 veto, pinned as `ja_katakana_lone_name_with_a_period_is_not_a_title` in `tests/v2/cases.py`. The last FOUR joined in the review round after that, as pins on readings nothing held: three stop-bearing FAMILY_COMMA spellings (`김민준씨., J.씨`, `田中さん., V.`, `이, J.씨.`) and the bracketed `(김민준.) John Smith`, the recorded degradation in the limits bullet above. A leading-stop shape moves the same way but sits in no corpus at all: `.김민준씨.` — a stop no script ever writes, glued before a honorific-bearing word — read one whole given `.김민준씨.` at `d37b8ec` and now reads given `.김민준`, suffix `씨.`, pinned only by the stage tests for its two edges taken separately (`test_the_peel_reads_the_trailing_stop_only` for the leading stop, `test_peels_a_listed_tail_through_a_trailing_full_stop` for the trailing one) rather than by any corpus row or a combined case of its own. The three comma rows agree with their stop-less twins in every field once the riding stop is removed from the suffix (`씨., J.씨` vs `씨, J.씨`; `さん.` vs `さん`; `씨.` vs `씨`) — each differs from its twin in exactly that one field, by exactly the stop that rides — which is the peel's own "one name, two spellings" argument holding as far as a raw field comparison can show it; the bracketed row does not agree with anything and is not meant to. `田中. 太郎` stays there with the rest, though it is what WITNESSES rules.md#H2's Accepted clause added here: `tools/differential/compare.py`'s tier comment records the 2026-09-05 precedent for exactly this collision — where a demoted file holds a text a rules.md example line names, the EXAMPLE LINE moves into the tolerated rule and the row is not promoted, because marking the row alone leaves the name enforced and documented as demoted. So the clause's CJK reading is carried in W3's tolerated example block, H2's own block keeps the Latin control `Smith. John`, and all four #323 shape rows stay `tolerated`. THE POPULATION THAT COULD HAVE MOVED AND DID NOT is two different sets depending on which criterion is read, and both are worth having in front of you. Under the SHAPE this bundle is about — an edge full stop glued to a word carrying a classified character — the corpora hold 22 names: the seventeen movers and FIVE non-movers (`田中さん 様.`, `田中さん, 様.`, `김민준 씨.`, `김민준 양.`, `김민준, 씨.`), all five already in `corpus_cjk_tolerated.jsonl` before this bundle. Two of the twenty-two need the detector to be written carefully, which is why the recipe is spelled out below: the stop on `김민준씨., J.씨` and `田中さん., V.` sits before a COMMA, so a detector splitting on whitespace alone sees `김민준씨.,` and finds no edge stop (an 18 recorded earlier in the day came from such a detector, over a corpus four names smaller), and the stop on `(김민준.)` is inside a bracketed clause until rules.md#S1's escape unwraps it. Under the looser reading of an ASCII PERIOD ANYWHERE in a CJK-bearing name, the corpora held 20 distinct names on 21 rows before this bundle (19 rows in the tolerated file and 2 in `corpus_issues.jsonl`, `Dr 김민준씨, Jr.` being the one name on two of them) — unmoved by the katakana row, which is a name this bundle's own work wrote, not a pre-bundle count. The looser set is wider because it sweeps in periods sitting on LATIN tokens inside a CJK name — `毛 泽东 Dr.` and `田中さん, V.` — which is precisely what the veto is scoped not to touch, it reading the word and not the name. NEITHER the five nor the twenty moves — the seventeen movers are the whole of what did, and every one of them is a row this bundle wrote. An earlier wording of this bullet said "the twenty-one CJK names carrying an ASCII period": that counted rows as names and quoted the looser criterion beside the shape's argument. RECOMPUTE: parse the union of the corpus files on both trees and diff the seven fields and the ambiguity kinds; for the SHAPE population, take each name's tokens (`_pipeline._tokenize`, not a whitespace split) and ask whether any token's `strip(FULL_STOPS)` is shorter and still carries a classified character, then add the bracketed clause by hand; for the looser one, count distinct names and rows separately. The gate exits 0 at all four baselines with `radar unclassified: 0`; one ledger rule carries them, `fix(#322/#323)`, ELEVEN members at 1.4.0 and 2.0.0 and SEVENTEEN at 2.1.0 and 2.2.0 — the three hangul names the older ledgers hand to their native-script CJK rule are one part of the difference, 2.1.0 being the release that shipped that behavior, and the three FAMILY_COMMA rows are the other: at 2.0.0 none of the three diffs at all, and at 1.4.0 the two that do (`田中さん., V.`, `이, J.씨.`) are already claimed by the broad `fix(cjk-comma-compound)` and `fix(cjk-comma-honorific-peel)` rules that ledger carries, so adding them to this rule would only take a name off a rule that describes it — the same narrow-first reasoning that keeps the three hangul names out. diff --git a/docs/design/rules.md b/docs/design/rules.md index e7b70590..69054647 100644 --- a/docs/design/rules.md +++ b/docs/design/rules.md @@ -2188,7 +2188,7 @@ O5. Rationale: O4 reads a name by comparing where its words stand, ## Scripts & writing systems (W) -Background: script-conditional behavior is permitted exactly where the writing system itself — not statistics about it — settles the convention; a language can never be inferred from Latin-script text, because transliteration destroys the signal. The facts this section builds on: Chinese and Japanese both write the family name first in native script, so the script settles the order without knowing the language. Hangul is written by exactly one language and Korean family names are a small closed census set. Han text does not identify its language — a Chinese surname list would divide Japanese 高橋一郎 as 高 + 橋一郎 — which is why Han division is opt-in and there is no Korean pack to opt into. Hiragana never transcribes a foreign name (transcriptions are katakana alone), so kanji-plus-kana is a Japanese name in Japanese order, while wholly-katakana is predominantly a transcribed foreign name already in given-first order. Real Chinese text is unspaced (毛泽东); the spaced 毛 泽东 is an artifact. A fuller narrative lives in docs/usage.rst's East Asian section. One fact carries its own consequence: none of the three writing systems marks the family name with a comma — position in the written form is what identifies it, so a comma standing between the family name and the given name is a listing convention carried in from elsewhere rather than a form the script produces. That is why the rule reading one (W3) is tolerated rather than normative. CLDR's own locale data says the same where a contrary convention would have had to appear: across its ko, zh and ja personName patterns not one of the 126 pattern strings carries a comma of any width, the surname-first referring patterns separating surname from given by a single space, and the only comma in reach belongs to the locale-neutral root's sorting format — a list-ordering format, which ko and zh override comma-free and ja all but one inherited slot (decisions.md#cjk-comma-demotion carries the pull verbatim, with its URLs, its commit and its date). A full stop of any width — the ASCII period, the fullwidth ., the ideographic 。 and its halfwidth 。 — glued after a script-written word is punctuation and not part of the word: it is invisible to the script reading and to the vocabulary, and it stays in the text on the word it arrived with, because no East Asian script writes an initial or an abbreviation with a period (#322, #323; decisions.md#cjk-full-stops). A stop glued BEFORE the word is punctuation to the vocabulary lookup, which folds both edges away, so .씨 is still the honorific. The classification fold that feeds the two division sites reads the trailing edge only, so a word wearing a leading stop is given no script at all and never becomes a surname site: .김민준 stays one whole word, given, no script rule reaching it. The honorific peel (W2) is not gated by that fold — the tail alone is its license — so it reads such a token regardless of a leading stop, and its own trailing-edge fold is what decides there: .김민준씨 peels to .김민준 and 씨, and .김민준씨. peels to .김민준 and 씨. (tests/v2/pipeline/test_script_segment.py's test_the_peel_reads_the_trailing_stop_only). +Background: script-conditional behavior is permitted exactly where the writing system itself — not statistics about it — settles the convention; a language can never be inferred from Latin-script text, because transliteration destroys the signal. The facts this section builds on: Chinese and Japanese both write the family name first in native script, so the script settles the order without knowing the language. Hangul is written by exactly one language and Korean family names are a small closed census set. Han text does not identify its language — a Chinese surname list would divide Japanese 高橋一郎 as 高 + 橋一郎 — which is why Han division is opt-in and there is no Korean pack to opt into. Hiragana never transcribes a foreign name (transcriptions are katakana alone), so kanji-plus-kana is a Japanese name in Japanese order, while wholly-katakana may be a transcribed foreign name already in given-first order or a Japanese name's reading written family-first, and nothing in the script says which. Katakana has a halfwidth form (タロウ for タロウ), the only kana systems without kanji could store, and legacy data still carries it: it is katakana, read the same way. Real Chinese text is unspaced (毛泽东); the spaced 毛 泽东 is an artifact. A fuller narrative lives in docs/usage.rst's East Asian section. One fact carries its own consequence: none of the three writing systems marks the family name with a comma — position in the written form is what identifies it, so a comma standing between the family name and the given name is a listing convention carried in from elsewhere rather than a form the script produces. That is why the rule reading one (W3) is tolerated rather than normative. CLDR's own locale data says the same where a contrary convention would have had to appear: across its ko, zh and ja personName patterns not one of the 126 pattern strings carries a comma of any width, the surname-first referring patterns separating surname from given by a single space, and the only comma in reach belongs to the locale-neutral root's sorting format — a list-ordering format, which ko and zh override comma-free and ja all but one inherited slot (decisions.md#cjk-comma-demotion carries the pull verbatim, with its URLs, its commit and its date). A full stop of any width — the ASCII period, the fullwidth ., the ideographic 。 and its halfwidth 。 — glued after a script-written word is punctuation and not part of the word: it is invisible to the script reading and to the vocabulary, and it stays in the text on the word it arrived with, because no East Asian script writes an initial or an abbreviation with a period (#322, #323; decisions.md#cjk-full-stops). A stop glued BEFORE the word is punctuation to the vocabulary lookup, which folds both edges away, so .씨 is still the honorific. The classification fold that feeds the two division sites reads the trailing edge only, so a word wearing a leading stop is given no script at all and never becomes a surname site: .김민준 stays one whole word, given, no script rule reaching it. The honorific peel (W2) is not gated by that fold — the tail alone is its license — so it reads such a token regardless of a leading stop, and its own trailing-edge fold is what decides there: .김민준씨 peels to .김민준 and 씨, and .김민준씨. peels to .김민준 and 씨. (tests/v2/pipeline/test_script_segment.py's test_the_peel_reads_the_trailing_stop_only). W1. Rationale: hangul is monoglot Korean and its surnames are a closed census set, so an unspaced hangul name divides at a @@ -2298,9 +2298,10 @@ W3. Rationale: a family name declared by a comma is the writer's W4. Rationale: Chinese, Japanese and Korean all write the family name first in native script — the script settles the order - without knowing the language — while a wholly-katakana name is - predominantly a transcribed foreign name already in its source - order. + without knowing the language — while a wholly-katakana name may + be a transcribed foreign name already in its source order or a + Japanese reading written family-first, so its script settles + nothing. A name written wholly in one East Asian script, or in the kana-licensed Japanese repertoire, reads family-first whatever order the caller declared; a wholly-katakana name keeps the @@ -2308,7 +2309,9 @@ W4. Rationale: Chinese, Japanese and Korean all write the family "김 민준" → family="김" "山田 太郎" → family="山田" "高橋 みなみ" → family="高橋" + "山田 タロウ" → family="山田" "マイケル ジャクソン" → given="マイケル" · boundary + "ヤマダ タロウ" → given="ヤマダ" · boundary Accepted: a name the interpunct divides keeps its source order — the divider itself marks a transcription (T3) — so the override stands down there; the katakana middle dot (T2) carries no such diff --git a/nameparser/_pipeline/_vocab.py b/nameparser/_pipeline/_vocab.py index 70bf9520..d0f1dedc 100644 --- a/nameparser/_pipeline/_vocab.py +++ b/nameparser/_pipeline/_vocab.py @@ -1431,8 +1431,10 @@ def effective_script(text: str) -> Script | None: katakana-only: マイケル has no kanji, but さくらエミ -- hiragana plus katakana -- is kana-only AND licensed) -- and resolves to the HIRAGANA carrier entry. Pure-katakana stays KATAKANA - (single_script's answer): a lone katakana token is predominantly a - transcribed foreign name, so nothing defaults on it.""" + (single_script's answer): a lone katakana token may be a + transcribed foreign name or a Japanese reading, and the script + cannot say which, so nothing defaults on it. Halfwidth kana is + katakana by the table (#594), so 山田タロウ is licensed too.""" # None for both shapes _wholly_ja could never match anyway (empty # text, or all-ASCII text): real work, not a leftover "if text" # guard, since the ASCII case is one a bare emptiness check would diff --git a/nameparser/_policy.py b/nameparser/_policy.py index 057183ef..bbed93b3 100644 --- a/nameparser/_policy.py +++ b/nameparser/_policy.py @@ -73,10 +73,14 @@ class Script(StrEnum): #: kanji+kana token (高橋みなみ) is Japanese and resolves HERE -- #: this member is the carrier key in script_orders/segment_scripts. HIRAGANA = "hiragana" - #: Japanese katakana. A PURE-katakana token is predominantly a - #: transcribed foreign name in its original order (マイケル), so - #: no default behavior keys on this member; it exists so the - #: classifier can name what it deliberately declines. + #: Japanese katakana, full-width and halfwidth (タロウ) alike. A + #: PURE-katakana token may be a transcribed foreign name in its + #: original order (マイケル) or a Japanese name's reading written + #: family-first (ヤマダ タロウ, and legacy halfwidth data), and the + #: script cannot say which, so no default behavior keys on this + #: member; it exists so the classifier can name what it + #: deliberately declines. A caller who knows their katakana is + #: Japanese maps it to FAMILY_FIRST in script_orders. KATAKANA = "katakana" @@ -125,11 +129,19 @@ class Script(StrEnum): # supplementary Han, which real surnames genuinely need. The Katakana # Phonetic Extensions block (U+31F0-U+31FF, 16 small katakana for Ainu # transcription) is excluded for the same reason -- no modern Japanese -# personal name uses them. Halfwidth kana (U+FF65-U+FF9F, including -# the voiced/semi-voiced sound marks U+FF9E/U+FF9F) is likewise -# deliberately excluded -- legacy bank/CSV data uses it, but it is a -# separate normalization problem; #272 Task 2b's separator handling -# only touches the halfwidth DOT (U+FF65), not the rest of that block. +# personal name uses them. Halfwidth kana (U+FF65-U+FF9F, #594) IS +# katakana here: legacy bank, payroll and CSV data written for JIS X +# 0201 systems spells names in it, and 山田 タロウ must read as 山田 タロウ +# does. Classifying it is a range, not a fold -- T1 forbids rewriting +# the text, and an NFKC fold at classification would reach far past +# the kana (fullwidth Latin, ㈱) while mapping a lone voicing mark +# ゙ to the HIRAGANA block's combining U+3099. The span takes in the +# halfwidth nakaguro U+FF65, which tokenize turns into a separator, +# on the same direct-call grounds as U+30FB below, and the voicing +# marks U+FF9E/U+FF9F, which are spacing characters following their +# base and so need the block to be classified at all. It stops short +# of U+FF61-U+FF64, the halfwidth CJK punctuation, whose U+FF61 is a +# full stop (_lexicon.FULL_STOPS). # This table classifies by Unicode BLOCK, not the UAX #24 Script # property: U+30A0, U+30FB (the middle dot), and U+30FC (the # prolonged sound mark) all carry Script=Common under UAX #24, and the @@ -154,7 +166,7 @@ class Script(StrEnum): (0xF900, 0xFAFF), (0x20000, 0x323AF)), Script.HANGUL: ((0xAC00, 0xD7A3),), Script.HIRAGANA: ((0x3040, 0x309F),), - Script.KATAKANA: ((0x30A0, 0x30FF),), + Script.KATAKANA: ((0x30A0, 0x30FF), (0xFF65, 0xFF9F)), } #: The Japanese repertoire: the three scripts Japanese names draw on. @@ -302,8 +314,10 @@ def _order_repr(value: tuple[Role, ...]) -> str: #: (transcriptions are katakana-only), so it is Japanese, written #: family-first -- another default change in a minor, release-log- #: classified fix, #294's mechanism. KATAKANA is deliberately absent: -#: a PURE-katakana token is predominantly a transcribed foreign name -#: kept in its source (usually given-first) order, so nothing should +#: a PURE-katakana token may be a transcribed foreign name kept in its +#: source (usually given-first) order, or a Japanese reading written +#: family-first -- legacy halfwidth data (#594) writes every name in +#: katakana -- and the script cannot tell them apart, so nothing should #: default on it. Canonical form: sorted (Script, order) pairs, #: matching the field's storage. DEFAULT_SCRIPT_ORDERS: tuple[ diff --git a/nameparser/locales/ja.py b/nameparser/locales/ja.py index dfa0a59f..ad9cce84 100644 --- a/nameparser/locales/ja.py +++ b/nameparser/locales/ja.py @@ -21,7 +21,9 @@ Han names have read family-first since the 2026-07-27 amendment. Pure KATAKANA is deliberately outside both the order rule and the activation set above -- katakana is how Japanese writes FOREIGN names - (マイケル・ジャクソン), in their original given-first order. + (マイケル・ジャクソン), in their original given-first order, and a + Japanese name's katakana reading (ヤマダ タロウ, or the halfwidth ヤマダ + タロウ of legacy data, #594) cannot be told from one by its script. Data sources: none. The pack carries no vocabulary of its own, so there is no list here to cite, extend or keep current -- the data that diff --git a/tests/v2/cases.py b/tests/v2/cases.py index b8938e0d..ebcbc88b 100644 --- a/tests/v2/cases.py +++ b/tests/v2/cases.py @@ -9123,6 +9123,39 @@ def _check_cjk_shape_purity(self) -> None: "reading to O5's convention, which reports it (#449); " "a name whose script order decides it stays silent. " "Roles parity; the flag is #449's"), + Case("ja_halfwidth_kanji_katakana_pieces", "山田 タロウ", + {"family": "山田", "given": "タロウ"}, + classification="fix(#594)", + notes="halfwidth katakana (U+FF65-U+FF9F) is katakana: the " + "kana license reads it exactly as it reads 山田 エミ. " + "Legacy JIS X 0201 data -- bank, payroll and CSV " + "exports -- writes kana this way. 2.3.0 and every " + "earlier release left the block unclassified and read " + "given 山田, family タロウ"), + Case("ja_halfwidth_voicing_marks_are_katakana", "山田 ダイスケ", + {"family": "山田", "given": "ダイスケ"}, + classification="fix(#594)", + notes="the halfwidth voicing mark ゙ (U+FF9E) is a SPACING " + "character after its base, not a combining one, so " + "only the block span classifies it; a range ending at " + "U+FF9D would leave ダイスケ mixed-script and the name " + "positional"), + Case("ja_halfwidth_pure_katakana_positional", "ヤマダ タロウ", + {"given": "ヤマダ", "family": "タロウ"}, + notes="parity row guarding the license's boundary in " + "halfwidth: wholly-katakana keeps the declared order " + "(rules.md#W4). Legacy halfwidth data writes Japanese " + "names this way too, family-first, but the script " + "cannot tell a reading from a transcription, so the " + "caller opts in through script_orders (#594)"), + Case("ja_interpunct_b7_halfwidth_katakana", "タロウ·ヤマダ", + {"given": "タロウ", "family": "ヤマダ"}, + classification="fix(#594)", + notes="the 间隔号 divides between two classified characters " + "(rules.md#T3), and halfwidth kana is classified now, " + "so the dot divides here as it does in タロウ·ヤマダ " + "and the parts keep source order. 2.3.0 kept the whole " + "text one token"), Case("ja_iteration_mark_is_han", "佐々木 太郎", {"family": "佐々木", "given": "太郎"}, classification="fix(#272)", @@ -9659,6 +9692,24 @@ def _check_cjk_shape_purity(self) -> None: "-- the veto decides title-or-name, the order rule " "decides which name. 2.2.0 read title マイケル.", tolerated=True), + Case("ja_halfwidth_katakana_lone_name_with_a_period_is_not_a_title", + "マイケル.", + {"given": "マイケル."}, ambiguities=("given-or-family",), + classification="fix(#594)", + notes="the halfwidth twin of the row above: katakana has no " + "initials (_NO_INITIALS), and halfwidth kana is katakana " + "now, so H2's opening-abbreviation shape declines it. " + "2.3.0 read title マイケル., the halfwidth limit " + "decisions.md#cjk-full-stops recorded", + tolerated=True), + Case("ja_halfwidth_katakana_opener_with_a_period_is_a_name", + "タナカ. John", + {"given": "タナカ.", "family": "John"}, + classification="fix(#594)", + notes="decisions.md#cjk-full-stops' recorded misroute: a " + "period-marked halfwidth opener read as a TITLE, where " + "タナカ. John reads given. Now both read given", + tolerated=True), Case("latin_period_marked_opening_word_is_still_a_title", "Smith. John", {"title": "Smith.", "family": "John"}, diff --git a/tests/v2/pipeline/test_script_segment.py b/tests/v2/pipeline/test_script_segment.py index d3ec60a3..be0e48cf 100644 --- a/tests/v2/pipeline/test_script_segment.py +++ b/tests/v2/pipeline/test_script_segment.py @@ -411,8 +411,9 @@ def test_a_neighbour_in_an_UNACTIVATED_script_still_blocks_the_consult() -> None # so the name is already divided and the segmenter must not be # asked. Narrowing the check to `in scripts` would make both of # these split 山 + 田太郎 while the writer's own boundary sat one - # token to the right. - for name in ("山田太郎 マイケル", "山田太郎 김민준"): + # token to the right. The halfwidth katakana neighbour is the + # same boundary in legacy data, and counts since #594 classified it. + for name in ("山田太郎 マイケル", "山田太郎 マイケル", "山田太郎 김민준"): out = _run(name, policy=_JA, segmenter=_fake((1,))) assert _texts(out) == name.split(), name assert out.ambiguities == (), name diff --git a/tests/v2/pipeline/test_vocab.py b/tests/v2/pipeline/test_vocab.py index 5d814deb..8ba5d38f 100644 --- a/tests/v2/pipeline/test_vocab.py +++ b/tests/v2/pipeline/test_vocab.py @@ -50,6 +50,7 @@ def test_is_initial_script_repertoire() -> None: assert not is_initial("김.") assert not is_initial("さ.") assert not is_initial("ラ.") + assert not is_initial("ラ.") # halfwidth katakana is katakana (#594) # unchanged: a digit is ONE edge of \w's reach and '_' is another, # and the shape half still owns both -- only the repertoire narrowed assert is_initial("2.") diff --git a/tests/v2/test_ledger_guards.py b/tests/v2/test_ledger_guards.py index 267e55a2..b91ed016 100644 --- a/tests/v2/test_ledger_guards.py +++ b/tests/v2/test_ledger_guards.py @@ -58,16 +58,19 @@ class declares, which members an alternation offers. Those are exact _entry_name, _exclusions, _rules, _unclassified_names, load_tool) -# The one sanctioned divergence between the differential rules' -# character classes and _SCRIPT_RANGES: the halfwidth middle dot -# separates tokens without being classified (halfwidth kana stays out -# of the table on purpose). U+00B7 is deliberately NOT here -- its -# flank guard means every name it can change matches through a -# classified flanking character already. Single-sourced: read by the -# span sweep below, and by the membership guard that keeps "sanctioned" -# meaning something -- an extra that becomes classified belongs in the -# table, not in this list. -_SANCTIONED_EXTRAS = frozenset({(0xFF65, 0xFF65)}) +# Sanctioned divergences between the differential rules' character +# classes and _SCRIPT_RANGES: a span a rule must cover although the +# table does not classify it. Empty since #594. Its one member was the +# halfwidth middle dot U+FF65, which separates tokens without having +# been classified while halfwidth kana stayed out of the table; #594 +# classified the whole halfwidth kana block, U+FF65 with it, and the +# membership guard below moved it into the table as it exists to. +# U+00B7 is deliberately NOT here -- its flank guard means every name +# it can change matches through a classified flanking character +# already. Single-sourced: read by the span sweep below, and by the +# membership guard that keeps "sanctioned" meaning something -- an +# extra that becomes classified belongs in the table, not in this list. +_SANCTIONED_EXTRAS: frozenset[tuple[int, int]] = frozenset() @pytest.mark.parametrize("field", _PHRASE_FIELDS) @@ -296,7 +299,7 @@ def test_script_ranges_membership_is_decided() -> None: The second assert is what makes _SANCTIONED_EXTRAS mean something. That set is the ledgers' licence to be WIDER than the table -- see - its definition above for why U+FF65 is in it and U+00B7 is not -- + its definition above for why U+FF65 left it and U+00B7 was never in it -- and a licence nobody audits is just a hole. An extra that becomes classified belongs in the table, not in the exception list, and fails here until it moves. @@ -3638,8 +3641,15 @@ def _claim(rule: dict) -> _Claim: # 2026-10-02, #585: 138 -> 139, one new corpus name and not a # wider rule: the decomposed katakana R3 row 'マイケル ジャクソン' # (NFD) lies in its script span, and fix(#585) explains it. + # 2026-10-03, #594: 139 -> 145, and this time the regex DID + # widen: its class copies the halfwidth kana block U+FF65-U+FF9F + # whole, where it held U+FF65 alone. What it reaches beyond that + # is exactly the six halfwidth case rows #594 added -- no corpus + # line held halfwidth kana before them -- and it explains the + # three whose name parts move; 'ヤマダ タロウ' is parity, as + # 'マイケル ジャクソン' already was inside the same class. "fix(#271/#272/#298) native-script CJK: family-first order, hangul segmentation, the kana license and the dots": - _Claim(139, ('family', 'given', 'middle'), "0feb71190b77", None), + _Claim(145, ('family', 'given', 'middle'), "cfd0d19b46d5", None), # 2026-09-19, #533: 33 -> 68. The count grew with the CORPUS # rather than with the rule -- this change added 35 # maiden-clause names as rules.md example lines and @@ -4606,6 +4616,8 @@ def _claim(rule: dict) -> _Claim: _Claim(3, ('family', 'suffix'), "54ce0dda7114", ('DEFAULT',)), "fix(#585) a decomposed initial keeps its whole first letter": _Claim(2, ('_initials',), "c4f045134f0f", None), + "fix(#594) a period-marked halfwidth katakana word is not a title": + _Claim(2, ('given', 'title'), "8cdafcf56c45", None), }, "expected_since_2.0.0.toml": { # The ph removal (#459/#521): one literal name, the cases.py @@ -4785,8 +4797,15 @@ def _claim(rule: dict) -> _Claim: # 2026-10-02, #585: 138 -> 139, one new corpus name and not a # wider rule: the decomposed katakana R3 row 'マイケル ジャクソン' # (NFD) lies in its script span, and fix(#585) explains it. + # 2026-10-03, #594: 139 -> 145, and this time the regex DID + # widen: its class copies the halfwidth kana block U+FF65-U+FF9F + # whole, where it held U+FF65 alone. What it reaches beyond that + # is exactly the six halfwidth case rows #594 added -- no corpus + # line held halfwidth kana before them -- and it explains the + # three whose name parts move; 'ヤマダ タロウ' is parity, as + # 'マイケル ジャクソン' already was inside the same class. "fix(#271/#272/#298) native-script CJK: family-first order, hangul segmentation, the kana license and the dots": - _Claim(139, ('_ambiguities', 'family', 'given', 'middle'), "0feb71190b77", None), + _Claim(145, ('_ambiguities', 'family', 'given', 'middle'), "cfd0d19b46d5", None), # 37 -> 35 with the same 2026-09-05 narrowing as the 1.4 twin, # whose entry carries the reason. Here the one name that # changed hands, '김민준 박사님', goes to the spaced rule @@ -5363,6 +5382,8 @@ def _claim(rule: dict) -> _Claim: _Claim(3, ('_ambiguities',), "54ce0dda7114", ('DEFAULT',)), "fix(#585) a decomposed initial keeps its whole first letter": _Claim(2, ('_initials',), "c4f045134f0f", None), + "fix(#594) a period-marked halfwidth katakana word is not a title": + _Claim(2, ('_ambiguities', 'given', 'title'), "8cdafcf56c45", None), }, # The 2.3 cycle's first rule, and a facade-only render fix: every # role is identical, so `_initials` alone. Reach and digest as in @@ -5839,6 +5860,10 @@ def _claim(rule: dict) -> _Claim: _Claim(3, ('_ambiguities',), "54ce0dda7114", ('DEFAULT',)), "fix(#585) a decomposed initial keeps its whole first letter": _Claim(2, ('_initials',), "c4f045134f0f", None), + "fix(#594) halfwidth katakana takes the kana license and the 间隔号": + _Claim(3, ('family', 'given'), "fdaa80515e52", None), + "fix(#594) a period-marked halfwidth katakana word is not a title": + _Claim(2, ('_ambiguities', 'given', 'title'), "8cdafcf56c45", None), }, "expected_since_2.1.0.toml": { # The ph removal (#459/#521): one literal name, the cases.py @@ -6560,6 +6585,10 @@ def _claim(rule: dict) -> _Claim: _Claim(3, ('_ambiguities',), "54ce0dda7114", ('DEFAULT',)), "fix(#585) a decomposed initial keeps its whole first letter": _Claim(2, ('_initials',), "c4f045134f0f", None), + "fix(#594) halfwidth katakana takes the kana license and the 间隔号": + _Claim(3, ('family', 'given'), "fdaa80515e52", None), + "fix(#594) a period-marked halfwidth katakana word is not a title": + _Claim(2, ('_ambiguities', 'given', 'title'), "8cdafcf56c45", None), }, "expected_since_2.3.0.toml": { # The ph removal (#459/#521): one literal name, the cases.py @@ -6881,6 +6910,10 @@ def _claim(rule: dict) -> _Claim: _Claim(3, ('_ambiguities',), "54ce0dda7114", ('DEFAULT',)), "fix(#585) a decomposed initial keeps its whole first letter": _Claim(2, ('_initials',), "c4f045134f0f", None), + "fix(#594) halfwidth katakana takes the kana license and the 间隔号": + _Claim(3, ('_ambiguities', 'family', 'given'), "fdaa80515e52", None), + "fix(#594) a period-marked halfwidth katakana word is not a title": + _Claim(2, ('_ambiguities', 'given', 'title'), "8cdafcf56c45", None), }, } diff --git a/tools/differential/corpus_cjk.jsonl b/tools/differential/corpus_cjk.jsonl index e48b6482..e4cf05f5 100644 --- a/tools/differential/corpus_cjk.jsonl +++ b/tools/differential/corpus_cjk.jsonl @@ -20,6 +20,8 @@ "山田 太郎 (マイケル・ジャクソン)" "山田 花子 旧姓 佐藤" "山田 花子(旧姓 佐藤)" +"山田 タロウ" +"山田 ダイスケ" "山田「タロ」太郎" "山田太郎様" "山田花子 旧姓 佐藤" @@ -68,3 +70,5 @@ "씨" "양 미선" "양 지훈" +"タロウ·ヤマダ" +"ヤマダ タロウ" diff --git a/tools/differential/corpus_cjk_tolerated.jsonl b/tools/differential/corpus_cjk_tolerated.jsonl index da1f6590..bf1847f7 100644 --- a/tools/differential/corpus_cjk_tolerated.jsonl +++ b/tools/differential/corpus_cjk_tolerated.jsonl @@ -57,3 +57,5 @@ "이, J.씨" "이, J.씨." "지훈, 남궁민수" +"タナカ. John" +"マイケル." diff --git a/tools/differential/corpus_rules.jsonl b/tools/differential/corpus_rules.jsonl index 681b139e..3d201ba7 100644 --- a/tools/differential/corpus_rules.jsonl +++ b/tools/differential/corpus_rules.jsonl @@ -488,6 +488,7 @@ "安东尼·陈志明" "山田 太郎" "山田 花子 旧姓:佐藤" +"山田 タロウ" "毛·泽东" "毛泽东" "王君" @@ -505,3 +506,4 @@ "남궁민수" "남궁민수 지훈" "선생님" +"ヤマダ タロウ" diff --git a/tools/differential/expected_since_1.4.0.toml b/tools/differential/expected_since_1.4.0.toml index de230553..ad2caa1d 100644 --- a/tools/differential/expected_since_1.4.0.toml +++ b/tools/differential/expected_since_1.4.0.toml @@ -181,16 +181,18 @@ issue = "fix(#271/#272/#298) native-script CJK: family-first order, hangul segme # supplementary-plane kana is hentaigana and other archaic forms no # modern name uses, so _SCRIPT_RANGES does not list it either. # -# U+FF65 is the one span here that _SCRIPT_RANGES does NOT have, and -# it is deliberate. The halfwidth middle dot separates tokens like -# its fullwidth twin, so 'マイケル・ジャクソン' splits where 1.4 left one -# token, but halfwidth kana is excluded from CLASSIFICATION on -# purpose (a separate normalization problem). The rest of the -# halfwidth block (U+FF66-U+FF9F) is left out for the usual tightness -# reason, and it was measured rather than assumed: a dotless -# halfwidth name such as 'マイケル ジャクソン' is byte-identical on both -# sides, so covering the block would pre-excuse a future regression -# on a shape that is parity today. U+00B7, the context-sensitive +# The halfwidth kana span (U+FF65-U+FF9F) is the table's own since +# #594. Until then U+FF65 alone stood here, a span _SCRIPT_RANGES did +# not have: the halfwidth middle dot separated tokens like its +# fullwidth twin, so 'マイケル・ジャクソン' split where 1.4 left one token, +# while the rest of the halfwidth block was left out of both because +# a dotless halfwidth name was parity -- which it still is where it +# is wholly katakana ('ヤマダ タロウ' keeps the declared order, as +# 'マイケル ジャクソン' does). What #594 moved is the kana license and +# the 间隔号's flank guard reaching halfwidth kana: '山田 タロウ' reads +# family-first as '山田 エミ' does, and 'タロウ·ヤマダ' divides. Both are +# this rule's diffs, so the class now copies the block whole like +# every other span here. U+00B7, the context-sensitive # 间隔号 (#298), deliberately gets NO span even though it too changes # parses: its flank guard divides only between classified-script # characters, so every name it can change already matches this class @@ -216,7 +218,7 @@ issue = "fix(#271/#272/#298) native-script CJK: family-first order, hangul segme # radar, so an unmatched diff on one of its names reports instead of # failing), and a rule fires on names it is given regardless of what # an unmatched diff would then cost. -name_regex = "[\\u3005-\\u3006\\u3040-\\u309F\\u30A0-\\u30FF\\u3400-\\u4DBF\\u4E00-\\u9FFF\\uF900-\\uFAFF\\uAC00-\\uD7A3\\uFF65-\\uFF65]" +name_regex = "[\\u3005-\\u3006\\u3040-\\u309F\\u30A0-\\u30FF\\u3400-\\u4DBF\\u4E00-\\u9FFF\\uF900-\\uFAFF\\uAC00-\\uD7A3\\uFF65-\\uFF9F]" fields = ["given", "middle", "family"] [[change]] @@ -1536,7 +1538,7 @@ issue = "fix(cjk-delimited-nickname) delimiter recognition compounds with the CJ # strings, because a pin selected the canonical CJK rule by those # substrings and asserted uniqueness; that pin is gone (#333) and the # span sweep finds rules by their character class instead. -name_regex = "(?s)(?=.*[「」『』・・])(?=.*[\\u3005-\\u3006\\u3040-\\u309F\\u30A0-\\u30FF\\u3400-\\u4DBF\\u4E00-\\u9FFF\\uF900-\\uFAFF\\uAC00-\\uD7A3\\uFF65-\\uFF65])" +name_regex = "(?s)(?=.*[「」『』・・])(?=.*[\\u3005-\\u3006\\u3040-\\u309F\\u30A0-\\u30FF\\u3400-\\u4DBF\\u4E00-\\u9FFF\\uF900-\\uFAFF\\uAC00-\\uD7A3\\uFF65-\\uFF9F])" fields = ["given", "family", "nickname"] [[change]] @@ -1589,7 +1591,7 @@ issue = "fix(cjk-fullwidth-paren-nickname) fullwidth-parenthesis recognition com # appears again, this rule wakes and the harness says the dormant # reason has gone false. dormant = "its one name, '山田 花子(旧姓 佐藤)', diffs in `maiden` rather than `nickname` since rules.md#M3 and is claimed by fix(cjk-maiden-marker); no other corpus name pairs a fullwidth bracket with a CJK script" -name_regex = "(?s)(?=.*[()])(?=.*[\\u3005-\\u3006\\u3040-\\u309F\\u30A0-\\u30FF\\u3400-\\u4DBF\\u4E00-\\u9FFF\\uF900-\\uFAFF\\uAC00-\\uD7A3\\uFF65-\\uFF65])" +name_regex = "(?s)(?=.*[()])(?=.*[\\u3005-\\u3006\\u3040-\\u309F\\u30A0-\\u30FF\\u3400-\\u4DBF\\u4E00-\\u9FFF\\uF900-\\uFAFF\\uAC00-\\uD7A3\\uFF65-\\uFF9F])" fields = ["given", "middle", "family", "nickname"] [[change]] @@ -1631,7 +1633,7 @@ issue = "fix(cjk-comma-honorific-peel) glued honorific peels off a post-comma gi # compound rule carries, hand-copied from _SCRIPT_RANGES and pinned by # the same sync test -- a comma alone matches every Latin 'Smith, Jr.' # in the corpus. -name_regex = "(?s)(?=.*,)(?=.*[\\u3005-\\u3006\\u3040-\\u309F\\u30A0-\\u30FF\\u3400-\\u4DBF\\u4E00-\\u9FFF\\uF900-\\uFAFF\\uAC00-\\uD7A3\\uFF65-\\uFF65])" +name_regex = "(?s)(?=.*,)(?=.*[\\u3005-\\u3006\\u3040-\\u309F\\u30A0-\\u30FF\\u3400-\\u4DBF\\u4E00-\\u9FFF\\uF900-\\uFAFF\\uAC00-\\uD7A3\\uFF65-\\uFF9F])" fields = ["given", "suffix"] [[change]] @@ -1702,7 +1704,7 @@ issue = "fix(cjk-comma-compound) comma routing compounds with the CJK order flip # combination and was falsified by driving classify() over the reach -- # the method decisions.md#differential-ledger names as the only one # that answers this question. -name_regex = "(?s)(?=.*,)(?=.*[\\u3005-\\u3006\\u3040-\\u309F\\u30A0-\\u30FF\\u3400-\\u4DBF\\u4E00-\\u9FFF\\uF900-\\uFAFF\\uAC00-\\uD7A3\\uFF65-\\uFF65])" +name_regex = "(?s)(?=.*,)(?=.*[\\u3005-\\u3006\\u3040-\\u309F\\u30A0-\\u30FF\\u3400-\\u4DBF\\u4E00-\\u9FFF\\uF900-\\uFAFF\\uAC00-\\uD7A3\\uFF65-\\uFF9F])" fields = ["given", "family", "title", "suffix"] [[change.precedes_narrower]] @@ -4926,3 +4928,36 @@ issue = "fix(#585) a decomposed initial keeps its whole first letter" # decomposition data. name_regex = "^(e\u0301mile zola|マイケル シ\u3099ャクソン)$" fields = ["_initials"] + +# --------------------------------------------------------------- +# #594: HALFWIDTH KATAKANA IS KATAKANA. The script table classifies +# the halfwidth kana block U+FF65-U+FF9F as Script.KATAKANA, so what +# reaches full-width katakana reaches halfwidth now: the kana license +# (rules.md#W4) reads 山田 タロウ family-first as it reads 山田 エミ, the +# 间隔号 divides between two halfwidth kana (rules.md#T3), and H2's +# opening-abbreviation shape declines a period-marked katakana word +# because katakana has no initials (rules.md#H2's Accepted clause), +# closing the limit decisions.md#cjk-full-stops recorded. A WHOLLY +# halfwidth name keeps the declared order, as wholly full-width +# katakana does, so 'ヤマダ タロウ' is parity and in neither rule. +# +# Literal-anchored to the names: no corpus line held a halfwidth kana +# before this change, so the movers ARE the six case rows #594 added, +# less the parity one. A shape spelling -- any halfwidth kana -- would +# be a hand copy of the span the canonical CJK rule already carries. +# LAST in the file: each diff reported unexplained, or unclassified on +# the radar, on the run before these rules existed (2026-10-03). +# --------------------------------------------------------------- + +# The license and 间隔号 movers need no rule at this baseline: the +# canonical fix(#271/#272/#298) rule above explains them through +# its span class, which copies the halfwidth block since #594. + +[[change]] +issue = "fix(#594) a period-marked halfwidth katakana word is not a title" +# 'タナカ. John' and 'マイケル.' read title where their full-width twins read +# a name word; both read given now. Tolerated input (an edge full stop +# on a CJK word), so these are radar names and the rule classifies them +# for the release note rather than for the exit code. +name_regex = "^(?:タナカ\\. John|マイケル\\.)$" +fields = ["title", "given"] diff --git a/tools/differential/expected_since_2.0.0.toml b/tools/differential/expected_since_2.0.0.toml index 0632bf86..727ad133 100644 --- a/tools/differential/expected_since_2.0.0.toml +++ b/tools/differential/expected_since_2.0.0.toml @@ -377,7 +377,7 @@ issue = "fix(#271/#272/#298) native-script CJK: family-first order, hangul segme # entry removed from SUFFIX_WORDS or GLUED_HONORIFICS, now # forces this file to narrow with it instead of leaving a wide twin # behind to classify a real regression as intended. -name_regex = "[\\u3005-\\u3006\\u3040-\\u309F\\u30A0-\\u30FF\\u3400-\\u4DBF\\u4E00-\\u9FFF\\uF900-\\uFAFF\\uAC00-\\uD7A3\\uFF65-\\uFF65]" +name_regex = "[\\u3005-\\u3006\\u3040-\\u309F\\u30A0-\\u30FF\\u3400-\\u4DBF\\u4E00-\\u9FFF\\uF900-\\uFAFF\\uAC00-\\uD7A3\\uFF65-\\uFF9F]" fields = ["given", "middle", "family", "_ambiguities"] [[change]] @@ -603,7 +603,7 @@ issue = "fix(#298) 间隔号 division changes the comma reading, sending the cre # (U+30FB) and ・ (U+FF65) are three different separators with three # different rules in 2.1, and they are indistinguishable enough on # screen that a literal here would be unreviewable. -name_regex = "(?s)(?=.*,)(?=.*\\u00B7)(?=.*[\\u3005-\\u3006\\u3040-\\u309F\\u30A0-\\u30FF\\u3400-\\u4DBF\\u4E00-\\u9FFF\\uF900-\\uFAFF\\uAC00-\\uD7A3\\uFF65-\\uFF65])" +name_regex = "(?s)(?=.*,)(?=.*\\u00B7)(?=.*[\\u3005-\\u3006\\u3040-\\u309F\\u30A0-\\u30FF\\u3400-\\u4DBF\\u4E00-\\u9FFF\\uF900-\\uFAFF\\uAC00-\\uD7A3\\uFF65-\\uFF9F])" fields = ["given", "family", "title", "suffix"] [[change]] @@ -3989,3 +3989,38 @@ issue = "fix(#585) a decomposed initial keeps its whole first letter" # decomposition data. name_regex = "^(e\u0301mile zola|マイケル シ\u3099ャクソン)$" fields = ["_initials"] + +# --------------------------------------------------------------- +# #594: HALFWIDTH KATAKANA IS KATAKANA. The script table classifies +# the halfwidth kana block U+FF65-U+FF9F as Script.KATAKANA, so what +# reaches full-width katakana reaches halfwidth now: the kana license +# (rules.md#W4) reads 山田 タロウ family-first as it reads 山田 エミ, the +# 间隔号 divides between two halfwidth kana (rules.md#T3), and H2's +# opening-abbreviation shape declines a period-marked katakana word +# because katakana has no initials (rules.md#H2's Accepted clause), +# closing the limit decisions.md#cjk-full-stops recorded. A WHOLLY +# halfwidth name keeps the declared order, as wholly full-width +# katakana does, so 'ヤマダ タロウ' is parity and in neither rule. +# +# Literal-anchored to the names: no corpus line held a halfwidth kana +# before this change, so the movers ARE the six case rows #594 added, +# less the parity one. A shape spelling -- any halfwidth kana -- would +# be a hand copy of the span the canonical CJK rule already carries. +# LAST in the file: each diff reported unexplained, or unclassified on +# the radar, on the run before these rules existed (2026-10-03). +# --------------------------------------------------------------- + +# The license and 间隔号 movers need no rule at this baseline: the +# canonical fix(#271/#272/#298) rule above explains them through +# its span class, which copies the halfwidth block since #594. + +[[change]] +issue = "fix(#594) a period-marked halfwidth katakana word is not a title" +# 'タナカ. John' and 'マイケル.' read title where their full-width twins read +# a name word; both read given now. Tolerated input (an edge full stop +# on a CJK word), so these are radar names and the rule classifies them +# for the release note rather than for the exit code. +# `_ambiguities` is the v2 surface's: the lone 'マイケル.' reports +# given-or-family as 'マイケル.' does, where the title reported nothing. +name_regex = "^(?:タナカ\\. John|マイケル\\.)$" +fields = ["title", "given", "_ambiguities"] diff --git a/tools/differential/expected_since_2.1.0.toml b/tools/differential/expected_since_2.1.0.toml index db68dd94..046c0878 100644 --- a/tools/differential/expected_since_2.1.0.toml +++ b/tools/differential/expected_since_2.1.0.toml @@ -3900,3 +3900,41 @@ issue = "fix(#585) a decomposed initial keeps its whole first letter" # decomposition data. name_regex = "^(e\u0301mile zola|マイケル シ\u3099ャクソン)$" fields = ["_initials"] + +# --------------------------------------------------------------- +# #594: HALFWIDTH KATAKANA IS KATAKANA. The script table classifies +# the halfwidth kana block U+FF65-U+FF9F as Script.KATAKANA, so what +# reaches full-width katakana reaches halfwidth now: the kana license +# (rules.md#W4) reads 山田 タロウ family-first as it reads 山田 エミ, the +# 间隔号 divides between two halfwidth kana (rules.md#T3), and H2's +# opening-abbreviation shape declines a period-marked katakana word +# because katakana has no initials (rules.md#H2's Accepted clause), +# closing the limit decisions.md#cjk-full-stops recorded. A WHOLLY +# halfwidth name keeps the declared order, as wholly full-width +# katakana does, so 'ヤマダ タロウ' is parity and in neither rule. +# +# Literal-anchored to the names: no corpus line held a halfwidth kana +# before this change, so the movers ARE the six case rows #594 added, +# less the parity one. A shape spelling -- any halfwidth kana -- would +# be a hand copy of the span the canonical CJK rule already carries. +# LAST in the file: each diff reported unexplained, or unclassified on +# the radar, on the run before these rules existed (2026-10-03). +# --------------------------------------------------------------- + +[[change]] +issue = "fix(#594) halfwidth katakana takes the kana license and the 间隔号" +# '山田 タロウ' and '山田 ダイスケ' read family 山田 where every release read +# it given; 'タロウ·ヤマダ' divides where it was one given token. +name_regex = "^(?:山田 タロウ|山田 ダイスケ|タロウ\\u00B7ヤマダ)$" +fields = ["given", "family"] + +[[change]] +issue = "fix(#594) a period-marked halfwidth katakana word is not a title" +# 'タナカ. John' and 'マイケル.' read title where their full-width twins read +# a name word; both read given now. Tolerated input (an edge full stop +# on a CJK word), so these are radar names and the rule classifies them +# for the release note rather than for the exit code. +# `_ambiguities` is the v2 surface's: the lone 'マイケル.' reports +# given-or-family as 'マイケル.' does, where the title reported nothing. +name_regex = "^(?:タナカ\\. John|マイケル\\.)$" +fields = ["title", "given", "_ambiguities"] diff --git a/tools/differential/expected_since_2.2.0.toml b/tools/differential/expected_since_2.2.0.toml index e1b18890..ba8c7d99 100644 --- a/tools/differential/expected_since_2.2.0.toml +++ b/tools/differential/expected_since_2.2.0.toml @@ -2296,3 +2296,41 @@ issue = "fix(#585) a decomposed initial keeps its whole first letter" # decomposition data. name_regex = "^(e\u0301mile zola|マイケル シ\u3099ャクソン)$" fields = ["_initials"] + +# --------------------------------------------------------------- +# #594: HALFWIDTH KATAKANA IS KATAKANA. The script table classifies +# the halfwidth kana block U+FF65-U+FF9F as Script.KATAKANA, so what +# reaches full-width katakana reaches halfwidth now: the kana license +# (rules.md#W4) reads 山田 タロウ family-first as it reads 山田 エミ, the +# 间隔号 divides between two halfwidth kana (rules.md#T3), and H2's +# opening-abbreviation shape declines a period-marked katakana word +# because katakana has no initials (rules.md#H2's Accepted clause), +# closing the limit decisions.md#cjk-full-stops recorded. A WHOLLY +# halfwidth name keeps the declared order, as wholly full-width +# katakana does, so 'ヤマダ タロウ' is parity and in neither rule. +# +# Literal-anchored to the names: no corpus line held a halfwidth kana +# before this change, so the movers ARE the six case rows #594 added, +# less the parity one. A shape spelling -- any halfwidth kana -- would +# be a hand copy of the span the canonical CJK rule already carries. +# LAST in the file: each diff reported unexplained, or unclassified on +# the radar, on the run before these rules existed (2026-10-03). +# --------------------------------------------------------------- + +[[change]] +issue = "fix(#594) halfwidth katakana takes the kana license and the 间隔号" +# '山田 タロウ' and '山田 ダイスケ' read family 山田 where every release read +# it given; 'タロウ·ヤマダ' divides where it was one given token. +name_regex = "^(?:山田 タロウ|山田 ダイスケ|タロウ\\u00B7ヤマダ)$" +fields = ["given", "family"] + +[[change]] +issue = "fix(#594) a period-marked halfwidth katakana word is not a title" +# 'タナカ. John' and 'マイケル.' read title where their full-width twins read +# a name word; both read given now. Tolerated input (an edge full stop +# on a CJK word), so these are radar names and the rule classifies them +# for the release note rather than for the exit code. +# `_ambiguities` is the v2 surface's: the lone 'マイケル.' reports +# given-or-family as 'マイケル.' does, where the title reported nothing. +name_regex = "^(?:タナカ\\. John|マイケル\\.)$" +fields = ["title", "given", "_ambiguities"] diff --git a/tools/differential/expected_since_2.3.0.toml b/tools/differential/expected_since_2.3.0.toml index 9fb262c2..ebc0681a 100644 --- a/tools/differential/expected_since_2.3.0.toml +++ b/tools/differential/expected_since_2.3.0.toml @@ -1551,3 +1551,43 @@ issue = "fix(#585) a decomposed initial keeps its whole first letter" # decomposition data. name_regex = "^(e\u0301mile zola|マイケル シ\u3099ャクソン)$" fields = ["_initials"] + +# --------------------------------------------------------------- +# #594: HALFWIDTH KATAKANA IS KATAKANA. The script table classifies +# the halfwidth kana block U+FF65-U+FF9F as Script.KATAKANA, so what +# reaches full-width katakana reaches halfwidth now: the kana license +# (rules.md#W4) reads 山田 タロウ family-first as it reads 山田 エミ, the +# 间隔号 divides between two halfwidth kana (rules.md#T3), and H2's +# opening-abbreviation shape declines a period-marked katakana word +# because katakana has no initials (rules.md#H2's Accepted clause), +# closing the limit decisions.md#cjk-full-stops recorded. A WHOLLY +# halfwidth name keeps the declared order, as wholly full-width +# katakana does, so 'ヤマダ タロウ' is parity and in neither rule. +# +# Literal-anchored to the names: no corpus line held a halfwidth kana +# before this change, so the movers ARE the six case rows #594 added, +# less the parity one. A shape spelling -- any halfwidth kana -- would +# be a hand copy of the span the canonical CJK rule already carries. +# LAST in the file: each diff reported unexplained, or unclassified on +# the radar, on the run before these rules existed (2026-10-03). +# --------------------------------------------------------------- + +[[change]] +issue = "fix(#594) halfwidth katakana takes the kana license and the 间隔号" +# '山田 タロウ' and '山田 ダイスケ' read family 山田 where every release read +# it given; 'タロウ·ヤマダ' divides where it was one given token. +# `_ambiguities` moves on 'タロウ·ヤマダ' alone: one token reported +# given-or-family (#449, new in 2.3), and two read positionally do not. +name_regex = "^(?:山田 タロウ|山田 ダイスケ|タロウ\\u00B7ヤマダ)$" +fields = ["given", "family", "_ambiguities"] + +[[change]] +issue = "fix(#594) a period-marked halfwidth katakana word is not a title" +# 'タナカ. John' and 'マイケル.' read title where their full-width twins read +# a name word; both read given now. Tolerated input (an edge full stop +# on a CJK word), so these are radar names and the rule classifies them +# for the release note rather than for the exit code. +# `_ambiguities` is the v2 surface's: the lone 'マイケル.' reports +# given-or-family as 'マイケル.' does, where the title reported nothing. +name_regex = "^(?:タナカ\\. John|マイケル\\.)$" +fields = ["title", "given", "_ambiguities"] From 98094eb8ce75b7af38367c24e70e30b8cc0c1227 Mon Sep 17 00:00:00 2001 From: Derek Gulbranson Date: Sat, 3 Oct 2026 13:50:11 -0700 Subject: [PATCH 2/9] docs: halfwidth katakana in the Japanese sections and release log (#594) usage.rst explains halfwidth katakana where Japanese scripts are introduced and gives the script_orders recipe for callers whose katakana names are Japanese; locales.rst drops the claim that a halfwidth neighbour divides nothing. Co-Authored-By: Claude Opus 5.5 --- docs/locales.rst | 5 +++-- docs/release_log.rst | 2 ++ docs/usage.rst | 31 ++++++++++++++++++++++++++++--- 3 files changed, 33 insertions(+), 5 deletions(-) diff --git a/docs/locales.rst b/docs/locales.rst index c70a3888..603379d1 100644 --- a/docs/locales.rst +++ b/docs/locales.rst @@ -243,8 +243,9 @@ ones its own pack turned on. It is asked only where the surname list could not divide an unspaced token, and not where the name is already divided — by a second word in an East Asian script (the nakaguro ・ counts as a space), a family comma or a 间隔号. A word in any other -script beside the token — Latin (``"Dr. 高橋一郎"``), Cyrillic, even -halfwidth katakana — divides nothing, so the segmenter is still asked. Recognize the text +script beside the token — Latin (``"Dr. 高橋一郎"``), Cyrillic — divides +nothing, so the segmenter is still asked. Halfwidth katakana counts as +katakana, so ``"高橋一郎 タロウ"`` is already divided. Recognize the text you can actually read and return ``None`` for the rest, rather than answering for a script you never meant to handle. A segmenter is your code, so its failures do not stay inside the parse: its own exceptions diff --git a/docs/release_log.rst b/docs/release_log.rst index 8bc476b3..0dae50de 100644 --- a/docs/release_log.rst +++ b/docs/release_log.rst @@ -66,6 +66,8 @@ Release Log - **Fix a v1 ``Constants`` entry that is only a CJK full stop raising at the first parse.** After ``c.titles.add("。")``, ``HumanName("john smith", c)`` raised ``ValueError`` in 2.3; ``。``, ``.`` and ``。`` in any set or as a ``capitalization_exceptions`` key are now ignored with the same ``UserWarning`` as an entry with stray whitespace. 1.4.0 through 2.2.0 accepted such an entry and applied it to a name token that is nothing but that full stop, which 2.3 and later cannot match. An entry ending in such a full stop no longer raises either: ``c.suffix_not_acronyms.add("ma。")`` made ``HumanName("jack ma", c)`` raise ``ValueError`` in 2.3, and now gives last ``ma``, as without the entry. (closes #582) + - **Fix a Japanese name written in halfwidth katakana being read given-first.** ``HumanName("山田 タロウ")`` gives last ``山田``, first ``タロウ``, as ``山田 タロウ`` does, where every release gave first ``山田``, last ``タロウ``. Halfwidth katakana (U+FF65–U+FF9F), which legacy bank, payroll and CSV exports still carry, is now read as katakana everywhere the script matters: a second halfwidth word keeps the segmenter from re-dividing a kanji name (``高橋一郎 タロウ``), the Chinese ``·`` divides between halfwidth kana (``タロウ·ヤマダ`` gives first ``タロウ``, last ``ヤマダ`` where it was one first name), and a period-marked halfwidth word is no longer taken for an abbreviated title: ``タナカ. John`` gives first ``タナカ.`` where it gave title ``タナカ.``. A name written wholly in katakana, halfwidth or not, still keeps the order it was written in (``ヤマダ タロウ`` gives first ``ヤマダ``), because the script cannot say whether it is a Japanese name or a transcribed foreign one; to read your katakana names family-first, add ``(Script.KATAKANA, FAMILY_FIRST)`` to ``Policy.script_orders`` (see :ref:`east-asian-names`). See the ``W4`` entry of ``docs/design/decisions.md`` (closes #594) + **Additions** - **Add Lexicon.conjunctions_ambiguous, the one-letter connectives that read as initials.** A subset of ``conjunctions`` holding ``e`` and ``i`` by default; it is the knob for the change above rather than a switch. Portuguese data, where ``e`` links surnames the way ``y`` does in Spanish, takes it out: ``Lexicon.default().remove(conjunctions_ambiguous={"e"})`` restores the joining reading. Dutch data, where a bare single letter is an initial and never a connective, adds the other one: ``Lexicon.default().add(conjunctions_ambiguous={"y"})``. A v1 ``Constants`` has no manager of its own for it -- deleting the word from ``conjunctions`` is what turns the marking off, the same rule the glued-honorific tails follow. See ``docs/customize.rst`` (#383, #479) diff --git a/docs/usage.rst b/docs/usage.rst index c6ad560a..5353d636 100644 --- a/docs/usage.rst +++ b/docs/usage.rst @@ -303,7 +303,11 @@ ambiguous: native given names use it, but katakana is also how Japanese text writes a *foreign* name — マイケル・ジャクソン is Michael Jackson — and a transcription keeps the source language's order, given name first, its parts divided by the middle dot ・ (the nakaguro, -U+30FB) rather than by a space. +U+30FB) rather than by a space. Katakana also has a halfwidth form +(タロウ for タロウ), which older systems that could not store kanji used +for every name, Japanese or foreign; bank, payroll and CSV exports +still carry it. nameparser reads halfwidth katakana exactly as it reads +the full-width form. What happens automatically ^^^^^^^^^^^^^^^^^^^^^^^^^^ @@ -321,6 +325,8 @@ transcription, because a transcription is written in katakana alone. ('高橋', 'みなみ') >>> parse("山田 エミ").family '山田' + >>> parse("山田 タロウ").family + '山田' And the middle dot separates tokens the way a space does, so a transcribed foreign name divides into its parts — which, being wholly @@ -417,8 +423,27 @@ Romanized names ("Kim Min-jun", "Yamada Taro") are Latin script and follow the ordinary positional rules. Order genuinely varies in romanized data, so nothing script-based applies. A name written wholly in katakana stays positional for the reason given above, pack or no -pack: it is predominantly a transcription, and a transcription is -already in the order it should be read in. +pack: it may be a transcription, already in the order it should be +read in, or a Japanese name's reading written family-first — a +furigana field, or legacy halfwidth data — and the script cannot tell +the two apart. If you know your katakana names are Japanese, map +katakana to family-first yourself: + +.. doctest:: + + >>> from nameparser import (DEFAULT_SCRIPT_ORDERS, FAMILY_FIRST, Parser, + ... Policy, Script) + >>> parse("ヤマダ タロウ").family + 'タロウ' + >>> kana_family_first = Parser(policy=Policy(script_orders=( + ... *DEFAULT_SCRIPT_ORDERS, (Script.KATAKANA, FAMILY_FIRST)))) + >>> kana_family_first.parse("ヤマダ タロウ").family + 'ヤマダ' + +This reaches full-width katakana too, transcriptions included +(マイケル・ジャクソン would read family マイケル), which is why it is +yours to choose rather than a default. ``name_order=FAMILY_FIRST`` +would also work, but it reverses every name, Latin ones included. A Han transcription written with a space instead of the 间隔号 (威廉 莎士比亚) carries nothing to distinguish it from a native From 6451052b80ee599d88654ab6b74900d919ac4347 Mon Sep 17 00:00:00 2001 From: Derek Gulbranson Date: Sat, 3 Oct 2026 13:55:07 -0700 Subject: [PATCH 3/9] test: correct two comments that still called halfwidth kana unclassified (#594) Co-Authored-By: Claude Opus 5.5 --- tests/v2/cases.py | 5 +++-- tests/v2/test_parser.py | 8 ++++---- 2 files changed, 7 insertions(+), 6 deletions(-) diff --git a/tests/v2/cases.py b/tests/v2/cases.py index ebcbc88b..faa7f92d 100644 --- a/tests/v2/cases.py +++ b/tests/v2/cases.py @@ -81,8 +81,9 @@ def _stray_ascii(text: str) -> str: #: _script_matcher(Script.KATAKANA, whole=True) -- rather than a #: hand-copied codepoint range, so the KATAKANA span lives in exactly #: one place (nameparser._policy._SCRIPT_RANGES). That table's choice, -#: not this file's: halfwidth katakana (a different Unicode block, -#: U+FF65-U+FF9F, per _policy.py's own comment) is out of scope. +#: not this file's: since #594 it classifies halfwidth katakana +#: (U+FF65-U+FF9F) as katakana too, so a wholly-halfwidth text is a +#: shape-7 transcription here exactly as its full-width twin is. #: Applied to the text with whitespace stripped, so a spaced #: transcription still counts as wholly katakana; a whitespace-only #: string never reaches this predicate in practice, since the purity diff --git a/tests/v2/test_parser.py b/tests/v2/test_parser.py index 1294f228..06b31e25 100644 --- a/tests/v2/test_parser.py +++ b/tests/v2/test_parser.py @@ -1432,10 +1432,10 @@ def test_nakaguro_split_han_tokens_take_the_han_order() -> None: def test_halfwidth_nakaguro_splits_at_parse_level_too() -> None: - # decision, not accident: halfwidth kana classify as no script at - # all (_SCRIPT_RANGES only covers the fullwidth blocks), so this - # is order-agnostic positional fallback, not a script-order rule -- - # the dot still divides the tokens regardless + # decision, not accident: halfwidth kana is katakana (#594), and a + # name written wholly in katakana keeps the declared order + # (rules.md#W4) as マイケル・ジャクソン does -- KATAKANA has no + # script_orders entry -- while the dot divides the tokens regardless text = "マイケル・ジャクソン" n = parse(text) assert (n.given, n.family) == ( From 151ca9680320238020304d11055378d9c9f76f26 Mon Sep 17 00:00:00 2001 From: Derek Gulbranson Date: Sat, 3 Oct 2026 13:55:58 -0700 Subject: [PATCH 4/9] docs: fix what the review of #594's design-doc amendments found MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit - rules.md#H2's Accepted clause still called halfwidth katakana unclassified; #594 is what changed that, so the same PR amends it. - The NFKC fold was declined for the wrong reason. The interpunct flank guard reads raw text and only asks "classified?"; the real hazard is the kana license, since an uncomposable voicing mark (ア゙) folds into the hiragana block and would turn a wholly-katakana name family-first. Corrected in decisions.md#W4 and _policy.py's comment. - AGENTS.md's 间隔号 sentence still had pure katakana naming its convention, the rationale #594 retracted. - The cjk-full-stops header and the H2 pointer now say the halfwidth limit is closed. Co-Authored-By: Claude Opus 5.5 --- AGENTS.md | 2 +- docs/design/decisions.md | 6 +++--- docs/design/rules.md | 5 ++--- nameparser/_policy.py | 7 +++++-- 4 files changed, 11 insertions(+), 9 deletions(-) diff --git a/AGENTS.md b/AGENTS.md index 5cf3e3d6..8ab7d98a 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -263,7 +263,7 @@ logging.getLogger('HumanName').setLevel(logging.DEBUG) The library has two layers: `nameparser/config/` (data) and `nameparser/parser.py` (logic). -**Design philosophy — positional and language-agnostic.** The parser assigns parts by *position* plus small sets of words that join to neighbors; it never detects language. A name's language can't be reliably inferred from Latin-script transliteration ("Ali" is Arabic or Italian; "Van"/"Della"/"Bin" are first names in some cultures, particles in others), so language-specific rules belong in opt-in `Constants` config, never global defaults. Many "wrong for language X" reports (#133, #150, #130, #85, #103, #146, #83) are irreducible ambiguities — e.g. `de Mesnil` (want last name) vs `Van Johnson` (want first name) are the same `[prefix][word]` shape. Before adding a rule, confirm it doesn't break the opposite case (run the full suite — Portuguese and "Van Johnson" tests are the usual canaries). **The one scoped exception (2.1, #271/#272): script-conditional behavior is permitted exactly where the SCRIPT ITSELF — not statistics about it — determines the convention.** The never-detect-language rule above is about Latin *transliteration*, where the signal genuinely is destroyed; native script is a different question, and it is answered per behavior rather than per script. Five defaults fall out of it, plus a sixth that applies the same not-a-guess standard to specific WORDS rather than to a script (#308's honorific peel, below). Wholly-Han, wholly-Hangul and kana-licensed names read family-first (`Policy.script_orders`) — no language detection needed, because zh and ja both write family-first in native script, so order cannot be misread even though the language is unknowable. Unspaced hangul splits into surname + given name (`config/surnames.py` ships the Korean census list as DEFAULT vocabulary) — nothing but Korean is written in hangul and the surnames are a closed census set, and the vocabulary is self-selecting besides: a hangul entry can only ever match hangul text. Hiragana licenses Japanese (#272) — a name whose characters stay inside Han∪kana while carrying at least one kana cannot be Chinese (the kana rules it out) and is not a transcription (foreign names are transcribed in katakana ALONE, マイケル has no kanji), so 高橋みなみ and 山田 エミ read family-first too; mechanically they resolve to the HIRAGANA entry, the license's carrier key. PURE katakana — full-width or halfwidth, the halfwidth block U+FF65–FF9F being katakana in the script table since #594 — is excluded and keeps the positional default: マイケル・ジャクソン is a transcribed foreign name in its source order, and a Japanese reading written in katakana (legacy halfwidth data writes every name that way) cannot be told from one, so the script settles nothing and a caller who knows opts in through `script_orders`. And the 间隔号 U+00B7 (#298) is the transcription marker for scripts that HAVE no transcription script: a name it divides (威廉·莎士比亚 — flanked by classified characters on both sides, so Catalan's Gal·la is untouched) keeps its source order and never segments — the orthography names the convention, exactly as pure katakana does, with the divider carrying the signal instead of the script. And a listed CJK honorific glued to the END of a name token is split off it (#308) — 田中さん is 田中 plus さん — on the same orthography-settles-it test, narrowed for the glued position: an entry peels only where it can never end a name, so 씨/님/さん/様/先生 peel while 양/군/氏/博士/殿 stay spaced-only (김지양 and 田中博士 are names, and ~90 Japanese surnames end in 殿) and 君 is in NEITHER set (王君 is a complete Chinese name), though its kana spelling くん peels. Like the nakaguro's tokenize-level separation described next, it is reached by neither policy opt-out — but for its own reason: the vocabulary carries the license itself rather than borrowing the script's, so `segment_scripts` has nothing to say about it. Since #312 it also crosses the 间隔号, which still stops the surname split standing right beside it: it answers where a name DIVIDES into surname and given, and the peel never asks that question. Whether it also crosses the FAMILY comma is tolerated rather than settled: the 2026-09-01 demotion (rules.md#W3) narrowed that half from contract to best-effort, since no CJK writing system's own convention puts a comma between family and given at all — so `김, 민준씨` reads today exactly as the spaced `김 민준씨` does (family 김, given 민준, suffix 씨) while the split stands down as before, but that reading is watched on the differential's radar tier rather than pinned as contract. Its site is accordingly the name-bearing segment runs — `segments[:2]` under a family comma, and `segments[0]` as before otherwise, the family comma being the one structure that splits the name itself across two runs, with the honorific as often glued to the given side as to the family. That is the whole reach and nothing past it (`김, 민준 지훈씨` peels; `김, 민준, 지훈씨` and `김,, 민준씨` do not, both landing in a third run), and whether `segments[1]` is name text at all is now ASKED rather than inferred from the structure — `segment` does not guarantee it, since a one-word part before the comma reads as FAMILY_COMMA even when the part after it is entirely suffix-shaped, and the peel walking into such a run took `V.` for its site, found no listed tail and abandoned (#319). The question is `segment`'s own suffix-comma predicate, lifted into `_vocab.is_wholly_suffix` so the two stages cannot drift: a wholly suffix-shaped second run is declined and the scan stays in `segments[0]`, so `田中さん, V.` and `田中さん, Ph. D.` give さん up as `田中さん, PhD` always did. The test is necessary but NOT sufficient, and the second condition is not decoration: every honorific tail is also a suffix word, so a glued honorific is itself part of what makes its run read as suffix-shaped, and declining a run that holds the ONLY site loses the peel outright. `segments[0]` must therefore offer a peel site of its own before the second run is declined — `이, J.씨` and `선생님, J.씨` pass the suffix test and are scanned anyway, keeping the pre-#319 reading, while `김민준씨, J.씨` has a site on both sides and peels the person's own 씨 rather than the junk one behind the comma. Uniform in the PEEL, that is — where the credential itself lands is `assign`'s question and still differs by spelling (`V.` → `given`, `PhD` and `Ph. D.` → `suffix`). Not `_is_post_nominal` pluralized: the run predicate says yes both to what the token predicate vetoes (`V.`, `V`, `I` — the class the defect was reported as) and to what the token predicate never sees at all, since `period_joined_vocab` and the delimiter routes are the run predicate's alone (`Msc.Ed.` and `J.씨` reach it that way, and `田中さん, Msc.Ed.` moves with the rest). `Policy(lenient_comma_suffixes=False)` drops this call to the strict token test too — so those three read as name text again and keep the pre-#319 answer, while `Ph. D.` peels under the knob regardless, its merged `phd` passing the strict test. `田中さん, 太郎` is unchanged, and not because of its comma — the honorific there is not at the end of the name, 太郎 is. The nakaguro belongs to the same doctrine but is decided a layer down: U+30FB and its halfwidth twin U+FF65 separate tokens like whitespace, unconditionally and in tokenize, so neither policy opt-out (`script_orders={}`, `segment_scripts=()`) reaches it — the codepoints are CJK-only and appear in no other script's names, which is what licenses a tokenize-level rule where U+00B7 (also the Catalan punt volat, interior to Gal·la) needs the flanked-by-classified-script guard `_tokenize_region` gives it (#298). Han segmentation stays OPT-IN (`locales.ZH` for Chinese, `locales.JA` for Japanese) — a zh surname list corrupts Japanese kanji names, since 高 is a common Chinese surname and 高橋一郎 would split 高+橋一郎 where the correct reading is 高橋+一郎; no surname list divides a kanji name at all, so `locales.JA` activates the stage and a pluggable `Parser(segmenter=...)` does the dividing. Latin-script input is never touched by any of this: "Kim Min-jun" is genuinely order-ambiguous and stays governed by `name_order` and opt-in packs. Before adding a script-conditional rule, work out which of the three it is — certain, certain for this one behavior only, or a statistical guess wearing a script's clothes. +**Design philosophy — positional and language-agnostic.** The parser assigns parts by *position* plus small sets of words that join to neighbors; it never detects language. A name's language can't be reliably inferred from Latin-script transliteration ("Ali" is Arabic or Italian; "Van"/"Della"/"Bin" are first names in some cultures, particles in others), so language-specific rules belong in opt-in `Constants` config, never global defaults. Many "wrong for language X" reports (#133, #150, #130, #85, #103, #146, #83) are irreducible ambiguities — e.g. `de Mesnil` (want last name) vs `Van Johnson` (want first name) are the same `[prefix][word]` shape. Before adding a rule, confirm it doesn't break the opposite case (run the full suite — Portuguese and "Van Johnson" tests are the usual canaries). **The one scoped exception (2.1, #271/#272): script-conditional behavior is permitted exactly where the SCRIPT ITSELF — not statistics about it — determines the convention.** The never-detect-language rule above is about Latin *transliteration*, where the signal genuinely is destroyed; native script is a different question, and it is answered per behavior rather than per script. Five defaults fall out of it, plus a sixth that applies the same not-a-guess standard to specific WORDS rather than to a script (#308's honorific peel, below). Wholly-Han, wholly-Hangul and kana-licensed names read family-first (`Policy.script_orders`) — no language detection needed, because zh and ja both write family-first in native script, so order cannot be misread even though the language is unknowable. Unspaced hangul splits into surname + given name (`config/surnames.py` ships the Korean census list as DEFAULT vocabulary) — nothing but Korean is written in hangul and the surnames are a closed census set, and the vocabulary is self-selecting besides: a hangul entry can only ever match hangul text. Hiragana licenses Japanese (#272) — a name whose characters stay inside Han∪kana while carrying at least one kana cannot be Chinese (the kana rules it out) and is not a transcription (foreign names are transcribed in katakana ALONE, マイケル has no kanji), so 高橋みなみ and 山田 エミ read family-first too; mechanically they resolve to the HIRAGANA entry, the license's carrier key. PURE katakana — full-width or halfwidth, the halfwidth block U+FF65–FF9F being katakana in the script table since #594 — is excluded and keeps the positional default: マイケル・ジャクソン is a transcribed foreign name in its source order, and a Japanese reading written in katakana (legacy halfwidth data writes every name that way) cannot be told from one, so the script settles nothing and a caller who knows opts in through `script_orders`. And the 间隔号 U+00B7 (#298) is the transcription marker for scripts that HAVE no transcription script: a name it divides (威廉·莎士比亚 — flanked by classified characters on both sides, so Catalan's Gal·la is untouched) keeps its source order and never segments — the orthography names the convention, with the divider carrying the signal; pure katakana reaches the same reading by the opposite route, its script settling nothing so the declared order stands. And a listed CJK honorific glued to the END of a name token is split off it (#308) — 田中さん is 田中 plus さん — on the same orthography-settles-it test, narrowed for the glued position: an entry peels only where it can never end a name, so 씨/님/さん/様/先生 peel while 양/군/氏/博士/殿 stay spaced-only (김지양 and 田中博士 are names, and ~90 Japanese surnames end in 殿) and 君 is in NEITHER set (王君 is a complete Chinese name), though its kana spelling くん peels. Like the nakaguro's tokenize-level separation described next, it is reached by neither policy opt-out — but for its own reason: the vocabulary carries the license itself rather than borrowing the script's, so `segment_scripts` has nothing to say about it. Since #312 it also crosses the 间隔号, which still stops the surname split standing right beside it: it answers where a name DIVIDES into surname and given, and the peel never asks that question. Whether it also crosses the FAMILY comma is tolerated rather than settled: the 2026-09-01 demotion (rules.md#W3) narrowed that half from contract to best-effort, since no CJK writing system's own convention puts a comma between family and given at all — so `김, 민준씨` reads today exactly as the spaced `김 민준씨` does (family 김, given 민준, suffix 씨) while the split stands down as before, but that reading is watched on the differential's radar tier rather than pinned as contract. Its site is accordingly the name-bearing segment runs — `segments[:2]` under a family comma, and `segments[0]` as before otherwise, the family comma being the one structure that splits the name itself across two runs, with the honorific as often glued to the given side as to the family. That is the whole reach and nothing past it (`김, 민준 지훈씨` peels; `김, 민준, 지훈씨` and `김,, 민준씨` do not, both landing in a third run), and whether `segments[1]` is name text at all is now ASKED rather than inferred from the structure — `segment` does not guarantee it, since a one-word part before the comma reads as FAMILY_COMMA even when the part after it is entirely suffix-shaped, and the peel walking into such a run took `V.` for its site, found no listed tail and abandoned (#319). The question is `segment`'s own suffix-comma predicate, lifted into `_vocab.is_wholly_suffix` so the two stages cannot drift: a wholly suffix-shaped second run is declined and the scan stays in `segments[0]`, so `田中さん, V.` and `田中さん, Ph. D.` give さん up as `田中さん, PhD` always did. The test is necessary but NOT sufficient, and the second condition is not decoration: every honorific tail is also a suffix word, so a glued honorific is itself part of what makes its run read as suffix-shaped, and declining a run that holds the ONLY site loses the peel outright. `segments[0]` must therefore offer a peel site of its own before the second run is declined — `이, J.씨` and `선생님, J.씨` pass the suffix test and are scanned anyway, keeping the pre-#319 reading, while `김민준씨, J.씨` has a site on both sides and peels the person's own 씨 rather than the junk one behind the comma. Uniform in the PEEL, that is — where the credential itself lands is `assign`'s question and still differs by spelling (`V.` → `given`, `PhD` and `Ph. D.` → `suffix`). Not `_is_post_nominal` pluralized: the run predicate says yes both to what the token predicate vetoes (`V.`, `V`, `I` — the class the defect was reported as) and to what the token predicate never sees at all, since `period_joined_vocab` and the delimiter routes are the run predicate's alone (`Msc.Ed.` and `J.씨` reach it that way, and `田中さん, Msc.Ed.` moves with the rest). `Policy(lenient_comma_suffixes=False)` drops this call to the strict token test too — so those three read as name text again and keep the pre-#319 answer, while `Ph. D.` peels under the knob regardless, its merged `phd` passing the strict test. `田中さん, 太郎` is unchanged, and not because of its comma — the honorific there is not at the end of the name, 太郎 is. The nakaguro belongs to the same doctrine but is decided a layer down: U+30FB and its halfwidth twin U+FF65 separate tokens like whitespace, unconditionally and in tokenize, so neither policy opt-out (`script_orders={}`, `segment_scripts=()`) reaches it — the codepoints are CJK-only and appear in no other script's names, which is what licenses a tokenize-level rule where U+00B7 (also the Catalan punt volat, interior to Gal·la) needs the flanked-by-classified-script guard `_tokenize_region` gives it (#298). Han segmentation stays OPT-IN (`locales.ZH` for Chinese, `locales.JA` for Japanese) — a zh surname list corrupts Japanese kanji names, since 高 is a common Chinese surname and 高橋一郎 would split 高+橋一郎 where the correct reading is 高橋+一郎; no surname list divides a kanji name at all, so `locales.JA` activates the stage and a pluggable `Parser(segmenter=...)` does the dividing. Latin-script input is never touched by any of this: "Kim Min-jun" is genuinely order-ambiguous and stays governed by `name_order` and opt-in packs. Before adding a script-conditional rule, work out which of the three it is — certain, certain for this one behavior only, or a statistical guess wearing a script's clothes. **A constant's membership is a question you may reopen.** Proposing that a word be ADDED, REMOVED or MOVED between vocabulary sets is ordinary design work — a shipped entry is not evidence that anyone judged it. `SUFFIX_ACRONYMS` arrived in `af5bdab` as a bulk Wikipedia import never reviewed against surname collisions: 572 of its 577 alphabetic entries leave `family` empty in `"John "` against five ambiguous-gated exceptions (recomputed 2026-09-07; the fifth is `ba`; 568 of 575 against seven, measured 2026-09-25 after #540), and `sa`, `se` and `om` are borne as surnames (measured 2026-08-23), as were `rai`, `cha`, `ba` and `mc` before `decisions.md#suffix-acronym-collisions` decided all four — `rai` and `cha` removed, `ba` marked ambiguous, `mc` left alone as no borne name at all. The same entry marked `meng` and `lac` ambiguous on 2026-09-25 (#540). When a fix starts to look like new machinery, check the vocabulary first. Criterion: `decisions.md#vocabulary-collisions`, with #360's positional qualifier. diff --git a/docs/design/decisions.md b/docs/design/decisions.md index 7843185f..d3d8664f 100644 --- a/docs/design/decisions.md +++ b/docs/design/decisions.md @@ -468,7 +468,7 @@ Declined: - 2026-07-27 (script-scoped order amendment) — the family-first override is keyed to the SCRIPT of the written name, never to a guessed language: wholly-Han, wholly-Hangul, and kana-licensed Japanese read family-first because zh, ko and ja all write family-first in native script; wholly-katakana names are predominantly transcriptions and keep the declared order. Latin transliterations are never touched. - 2026-07-29 #272 — the kana license: Han∪kana with at least one kana cannot be Chinese and is not a transcription, so 高橋みなみ reads family-first though it is written in two scripts. -- 2026-10-03 #594 — HALFWIDTH KATAKANA IS KATAKANA. The script table classifies the halfwidth kana block U+FF65–FF9F as `Script.KATAKANA`, so `山田 タロウ` reads family 山田 as `山田 エミ` does; every release through 2.3.0 left the block unclassified and read it given 山田, family タロウ, the positional reading. Legacy data written for JIS X 0201 systems, which had no kanji, spells names this way — bank, payroll and CSV exports. DECIDED by widening the table rather than by folding: the T rules forbid rewriting the text (decisions.md#T1), so normalization at tokenize was never available, and an NFKC fold at classification — the other place the issue offered — reaches far past the kana (`JOHN` to `JOHN`, `㈱` to `(株)`) and maps a lone voicing mark `゙` (U+FF9E) to the combining U+3099, which sits in the HIRAGANA block, so the per-character interpunct flank guard would read it as hiragana. A block range is how the table already decides (U+30FB and the voicing marks are in by block), and it takes the voicing marks U+FF9E/U+FF9F in because they are spacing characters following their base. The span starts at U+FF65, the halfwidth nakaguro, which tokenize already treats as a separator (rules.md#T2), and stops short of U+FF61–FF64, whose U+FF61 is a full stop. Everything keyed on the table follows: the kana license, the second-East-Asian-word count that keeps a segmenter from being asked (`高橋一郎 タロウ`), the 间隔号's flank guard (`タロウ·ヤマダ` divides), and `_NO_INITIALS`, which closes the limit #cjk-full-stops recorded (`タナカ. John` read title `タナカ.`). THE RATIONALE FOR WHOLLY-KATAKANA NAMES IS AMENDED, NOT THE BEHAVIOR: a wholly-halfwidth name keeps the declared order exactly as a wholly-katakana one does (`ヤマダ タロウ` reads given ヤマダ), but the 2026-07-27 reason — that such names are predominantly transcriptions — is weaker here, since a system with no kanji wrote Japanese names in katakana too, family-first. What stands is the narrower claim: the script cannot tell a Japanese reading from a transcription, so it settles no order, and a caller who knows their katakana is Japanese adds `(Script.KATAKANA, FAMILY_FIRST)` to `script_orders` (docs/usage.rst's Boundaries section carries the recipe). `name_order=FAMILY_FIRST` works too, at the cost of reversing Latin names as well. MEASURED 2026-10-03: the differential gate saw nothing before this change because no corpus line held halfwidth kana; the six case rows #594 added are its whole population, three of them contract movers (`山田 タロウ`, `山田 ダイスケ`, `タロウ·ヤマダ`), two tolerated radar movers (`タナカ. John`, `マイケル.`) and one parity row (`ヤマダ タロウ`). RECOMPUTE: parse the six on this tree and on 2.3.0 from PyPI and read `title` beside the name fields. +- 2026-10-03 #594 — HALFWIDTH KATAKANA IS KATAKANA. The script table classifies the halfwidth kana block U+FF65–FF9F as `Script.KATAKANA`, so `山田 タロウ` reads family 山田 as `山田 エミ` does; every release through 2.3.0 left the block unclassified and read it given 山田, family タロウ, the positional reading. Legacy data written for JIS X 0201 systems, which had no kanji, spells names this way — bank, payroll and CSV exports. DECIDED by widening the table rather than by folding: the T rules forbid rewriting the text (decisions.md#T1), so normalization at tokenize was never available, and an NFKC fold at classification — the other place the issue offered — reaches far past the kana (`JOHN` to `JOHN`, `㈱` to `(株)`), and it would move a wholly-katakana name across W4's line: a voicing mark with no precomposed form to join (`ア゙`, `ン゙`) folds to the COMBINING U+3099, which sits in the HIRAGANA block, so the folded token reads as kana-licensed Japanese and `ア゙イ タロウ` would read family-first where the unfolded one keeps the declared order (measured 2026-10-03 by parsing the NFKC spelling `ア゙イ タロウ`, which reads family `ア゙イ`; the review's two-input sweep of halfwidth against folded spellings, 600 pairs, differed only on such marks). The interpunct flank guard is NOT the hazard, though an earlier wording of this bullet said it was: it reads raw text and asks only whether a character is classified at all, which hiragana and katakana both are. The same measurement shows a pre-existing gap the fold would have widened rather than created: full-width katakana typed with an uncomposable combining U+3099 already takes the kana license today. A block range is how the table already decides (U+30FB and the voicing marks are in by block), and it takes the voicing marks U+FF9E/U+FF9F in because they are spacing characters following their base. The span starts at U+FF65, the halfwidth nakaguro, which tokenize already treats as a separator (rules.md#T2), and stops short of U+FF61–FF64, whose U+FF61 is a full stop. Everything keyed on the table follows: the kana license, the second-East-Asian-word count that keeps a segmenter from being asked (`高橋一郎 タロウ`), the 间隔号's flank guard (`タロウ·ヤマダ` divides), and `_NO_INITIALS`, which closes the limit #cjk-full-stops recorded (`タナカ. John` read title `タナカ.`). THE RATIONALE FOR WHOLLY-KATAKANA NAMES IS AMENDED, NOT THE BEHAVIOR: a wholly-halfwidth name keeps the declared order exactly as a wholly-katakana one does (`ヤマダ タロウ` reads given ヤマダ), but the 2026-07-27 reason — that such names are predominantly transcriptions — is weaker here, since a system with no kanji wrote Japanese names in katakana too, family-first. What stands is the narrower claim: the script cannot tell a Japanese reading from a transcription, so it settles no order, and a caller who knows their katakana is Japanese adds `(Script.KATAKANA, FAMILY_FIRST)` to `script_orders` (docs/usage.rst's Boundaries section carries the recipe). `name_order=FAMILY_FIRST` works too, at the cost of reversing Latin names as well. MEASURED 2026-10-03: the differential gate saw nothing before this change because no corpus line held halfwidth kana; the six case rows #594 added are its whole population, three of them contract movers (`山田 タロウ`, `山田 ダイスケ`, `タロウ·ヤマダ`), two tolerated radar movers (`タナカ. John`, `マイケル.`) and one parity row (`ヤマダ タロウ`). RECOMPUTE: parse the six on this tree and on 2.3.0 from PyPI and read `title` beside the name fields. ### D1 — the segmenterless-activation warning @@ -513,7 +513,7 @@ Declined (ambiguity kinds for script-resolved names, 2026-07-27): Decided 2026-09-08 (was Open: [#316](https://github.com/derek73/python-nameparser/issues/316), what a trailing title-vocabulary word should do while the comma paths disagreed): the trailing slot reads VOCABULARY, not shape (#H5), and the leading inference here stays unconditional — #316's open question 2, a symmetric leading rule that would read "Esq. Smith" as a suffix, is DECLINED. #109 shipped the leading inference on purpose, the family comma already carries the credential case ("Smith, Esq." → suffix, this rule's own Accepted clause), and reversing S2's "a suffix never opens the string" for period-marked words for the sake of one word is the wrong trade. The two comma paths agree now: "John Smith Prof." and "Smith, Prof." both read title Prof. See #H5. -- 2026-09-10 — the shape declines a word carrying a script with no initials (#323); the argument, the halfwidth-kana limit and the measurements are in #cjk-full-stops. +- 2026-09-10 — the shape declines a word carrying a script with no initials (#323); the argument, the halfwidth-kana limit and the measurements are in #cjk-full-stops. That limit was closed on 2026-10-03 by #594, which classifies halfwidth katakana (decisions.md#W4). ### H3 — the title run's floor @@ -1515,7 +1515,7 @@ ONE GAP SEEN FROM TWO SIDES, which is why the two issues were answered together. - **#323 turned out broader than it was filed.** The issue named the honorific peel. Making the classification fold read through an edge stop moved three readers of the `None` it used to return, and only one of them was the peel's neighbour test: the SURNAME SITE stepped past the family name onto the given name (`양. 지훈` cut 지훈 in half — given `양.`, middle `지`, family `훈` — and now reads given `지훈`, family `양.`), the ORDER RULE fell back to positional (`양 지훈.` lost family-first and now keeps it; `田中 太郎.` likewise), and the segmenter's neighbour precondition missed a writer-drawn boundary (`山田太郎 田中.` consulted a pluggable segmenter on `山田太郎` as if it stood alone, and is blocked now). The issue weighed two candidate fixes and chose neither: teaching the surname site to consult `is_suffix_strict` beside `effective_script`, or making the classification tolerant of a trailing period, which it called the broader one and expected to interact with #322. The broader one shipped, and the narrow one would have reached the filed name alone — `김.` is no suffix vocabulary, so `김. 민준` sits outside it entirely. - **The surname-site amendment was found by measurement during execution, not designed.** With the classification fold alone, `김. 민준` read given `김`, middle `.`, family `민준`: the token classified as hangul, the site matched `김` against its own head, and what was left over — the stop — BECAME the remainder. So the site matches on `text.rstrip(FULL_STOPS)` and cuts the core, a head being a prefix, which makes the offset that cuts the core cut the text; the stop rides with the remainder (`김민준.` divides as `김` + `민준.`, `김. 민준` as family `김.` plus given `민준`). `rstrip`, not `strip`: a LEADING stop would break the prefix argument. The reading is recorded in `ebf64db`'s comment at the site. - **Two readings JOINED the bundle, approved in session on 2026-09-09.** Neither was filed. FIRST, the glued-period honorific (rules.md#W3): a stop on the honorific's own word stood between the listed tail and the token's end, so nothing peeled and `田中さん.` and `김민준씨.` read as titles. The peel now matches the tail on the core and cuts BEFORE it, so both divide where their stop-less spellings do and the stop rides with the honorific. The move was free because W3 had recorded that reading as measured on 2026-09-05 and pinned by nothing, its own words being that neither string is a case row or a corpus line, so no row pinned those two readings; both are case rows now, tolerated. SECOND, H2's opening-abbreviation shape on a CJK word: `田中.` read title and now reads family `田中.`, `田中. 太郎` reads family `田中.`, given `太郎`. Han is what WITNESSES this one, and hangul cannot: hangul segmentation runs first, so `김민준.` never reaches the shape test and divides as `김` + `민준.` by the surname site instead. Latin and Cyrillic are untouched (`Smith. John` still reads title `Smith.`, `Проф. Иванов` title `Проф.`), and the veto is contains-any, so `Kim김. Smith` is refused a title too — deliberately: a word carrying a script with no abbreviations is not wearing an abbreviation's period. -- **Recorded limits, all five measured 2026-09-10 on this tree — and the segmenter one is a limit the review round CLOSED, kept in this bullet because this is where the claim stood.** HALFWIDTH KANA sits outside the script table, so it reads by the LATIN shapes at both period sites and not just at the initial one: `_vocab.is_initial("ラ.")` is `True` where `is_initial("ラ.")` is `False`, and H2's opening-abbreviation shape MISROUTES a longer halfwidth opener — `parse("タナカ. John")` gives title `タナカ.` where `parse("タナカ. John")` gives given `タナカ.`. Both because `_policy._SCRIPT_RANGES` excludes the halfwidth kana block U+FF65–FF9F on purpose — legacy bank and CSV data writes it, and the table's comment calls that a separate normalization problem, #272's separator handling touching only the halfwidth DOT U+FF65 and none of the kana behind it — and `_NO_INITIALS`, being a tuple of Script members, silently inherits the exclusion. That is #322's second mechanism, restated here because it is the one the bundle did NOT close. An earlier wording of this bullet said "nothing misroutes today, every kana honorific being multi-character", which had the direction backwards: the misroute is at the SHAPE site rather than the peel, and being multi-character is what carries a halfwidth opener INTO H2's shape rather than what keeps it out. NOTHING PINS EITHER READING — no case row, no corpus line, for the halfwidth block or for `タナカ. John` — so both can move without the suite or the differential saying so. It stays a recorded limit rather than a fix: what to do with halfwidth kana is a normalization question (widen the script table, or fold widths at classification) and not a full-stop one, and the bundle answered the full-stop one. CLOSED 2026-10-03 by #594, which widened the table (decisions.md#W4 carries why not the fold): `タナカ. John` reads given `タナカ.` and `is_initial("ラ.")` is `False`, and both readings are pinned now — the case rows `ja_halfwidth_katakana_opener_with_a_period_is_a_name` and `ja_halfwidth_katakana_lone_name_with_a_period_is_not_a_title`, and `test_is_initial_script_repertoire`. RECOMPUTE: parse `タナカ. John` and `タナカ. John` and read `title` beside the name fields; `is_initial` alone reports only the shorter half. A PLUGGABLE SEGMENTER receives the CORE, the token with its trailing stops removed — the same string the vocabulary match reads, the surname site handing it `core` rather than `text` (corrected 2026-09-10, in the review round after this entry was first written: the site handed over the raw token, which this bullet recorded as a limit, and the fix costs nothing because a head is a prefix — every offset reported against the core cuts `text` in the same place, so the stops ride with the last piece and no answer of the segmenter's can make a stop a piece of its own). The shipped `ja` pack's `_wholly_japanese` therefore ACCEPTS `山田太郎.` now, where the raw token failed its repertoire test: measured 2026-09-10 with `namedivider-python` installed (`uv run --extra ja`), `parser_for(locales.JA, segmenter=locales.ja_segmenter())` reads `山田太郎.` as family 山田, given `太郎.` with a SEGMENTATION report scoring 0.44, where `d37b8ec` read title `山田太郎.` and the un-corrected branch read family `山田太郎.`; the already-divided `山田 太郎.` reads family 山田, given `太郎.` on this tree and given 山, middle 田, family `太郎.` at `d37b8ec`. Pinned by `test_a_consulted_segmenter_receives_the_core_without_the_stop`. A LEADING stop keeps the token out of the SURNAME site only, because the classification fold that gates it rstrips as it does: `.김민준` classifies as no script, never becomes a surname site and stays one whole word, reading given `.김민준` exactly as `d37b8ec` reads it (corrected 2026-09-10 in the same round: the fold stripped both edges when this entry was written, which moved the ROLE to family and — with a segmenter configured — offered the token to it raw, an answer of offset 1 dividing it into the stop and the name; the fold now matches the sites, and hiding such a token is the no-split rules.md#W1 already accepts). It does NOT keep the token out of the HONORIFIC PEEL, which consults no script at all — the tail alone is its license (rules.md#W2's Background) — so a leading stop does not defeat it: `.김민준씨` peels to `.김민준` and `씨` exactly as `김민준씨` does, pinned by `test_the_peel_reads_the_trailing_stop_only`. A BRACKETED CJK CREDENTIAL degrades: rules.md#S1's escape in `_extract` calls a clause suffix-shaped when it ends in an ASCII period, so `(김민준.) John Smith` is unwrapped and the H2 veto then leaves `김민준.` as name text the surname site divides — given 김, middle `민준. John`, family Smith, where 2.2.0 and `d37b8ec` read title `김민준.` — pinned as the tolerated `ko_name_with_a_period_in_a_bracketed_credential`. And a caller-authored NFD surname entry is composed at ingest, so raw NFD input then finds no entry and the name goes unsplit; pinned by `test_an_nfd_authored_entry_is_stored_composed`. +- **Recorded limits, all five measured 2026-09-10 on this tree — and the segmenter one is a limit the review round CLOSED, kept in this bullet because this is where the claim stood. The halfwidth-kana one was CLOSED on 2026-10-03 by #594; see the CLOSED note after its paragraph below, and decisions.md#W4.** HALFWIDTH KANA sits outside the script table, so it reads by the LATIN shapes at both period sites and not just at the initial one: `_vocab.is_initial("ラ.")` is `True` where `is_initial("ラ.")` is `False`, and H2's opening-abbreviation shape MISROUTES a longer halfwidth opener — `parse("タナカ. John")` gives title `タナカ.` where `parse("タナカ. John")` gives given `タナカ.`. Both because `_policy._SCRIPT_RANGES` excludes the halfwidth kana block U+FF65–FF9F on purpose — legacy bank and CSV data writes it, and the table's comment calls that a separate normalization problem, #272's separator handling touching only the halfwidth DOT U+FF65 and none of the kana behind it — and `_NO_INITIALS`, being a tuple of Script members, silently inherits the exclusion. That is #322's second mechanism, restated here because it is the one the bundle did NOT close. An earlier wording of this bullet said "nothing misroutes today, every kana honorific being multi-character", which had the direction backwards: the misroute is at the SHAPE site rather than the peel, and being multi-character is what carries a halfwidth opener INTO H2's shape rather than what keeps it out. NOTHING PINS EITHER READING — no case row, no corpus line, for the halfwidth block or for `タナカ. John` — so both can move without the suite or the differential saying so. It stays a recorded limit rather than a fix: what to do with halfwidth kana is a normalization question (widen the script table, or fold widths at classification) and not a full-stop one, and the bundle answered the full-stop one. CLOSED 2026-10-03 by #594, which widened the table (decisions.md#W4 carries why not the fold): `タナカ. John` reads given `タナカ.` and `is_initial("ラ.")` is `False`, and both readings are pinned now — the case rows `ja_halfwidth_katakana_opener_with_a_period_is_a_name` and `ja_halfwidth_katakana_lone_name_with_a_period_is_not_a_title`, and `test_is_initial_script_repertoire`. RECOMPUTE: parse `タナカ. John` and `タナカ. John` and read `title` beside the name fields; `is_initial` alone reports only the shorter half. A PLUGGABLE SEGMENTER receives the CORE, the token with its trailing stops removed — the same string the vocabulary match reads, the surname site handing it `core` rather than `text` (corrected 2026-09-10, in the review round after this entry was first written: the site handed over the raw token, which this bullet recorded as a limit, and the fix costs nothing because a head is a prefix — every offset reported against the core cuts `text` in the same place, so the stops ride with the last piece and no answer of the segmenter's can make a stop a piece of its own). The shipped `ja` pack's `_wholly_japanese` therefore ACCEPTS `山田太郎.` now, where the raw token failed its repertoire test: measured 2026-09-10 with `namedivider-python` installed (`uv run --extra ja`), `parser_for(locales.JA, segmenter=locales.ja_segmenter())` reads `山田太郎.` as family 山田, given `太郎.` with a SEGMENTATION report scoring 0.44, where `d37b8ec` read title `山田太郎.` and the un-corrected branch read family `山田太郎.`; the already-divided `山田 太郎.` reads family 山田, given `太郎.` on this tree and given 山, middle 田, family `太郎.` at `d37b8ec`. Pinned by `test_a_consulted_segmenter_receives_the_core_without_the_stop`. A LEADING stop keeps the token out of the SURNAME site only, because the classification fold that gates it rstrips as it does: `.김민준` classifies as no script, never becomes a surname site and stays one whole word, reading given `.김민준` exactly as `d37b8ec` reads it (corrected 2026-09-10 in the same round: the fold stripped both edges when this entry was written, which moved the ROLE to family and — with a segmenter configured — offered the token to it raw, an answer of offset 1 dividing it into the stop and the name; the fold now matches the sites, and hiding such a token is the no-split rules.md#W1 already accepts). It does NOT keep the token out of the HONORIFIC PEEL, which consults no script at all — the tail alone is its license (rules.md#W2's Background) — so a leading stop does not defeat it: `.김민준씨` peels to `.김민준` and `씨` exactly as `김민준씨` does, pinned by `test_the_peel_reads_the_trailing_stop_only`. A BRACKETED CJK CREDENTIAL degrades: rules.md#S1's escape in `_extract` calls a clause suffix-shaped when it ends in an ASCII period, so `(김민준.) John Smith` is unwrapped and the H2 veto then leaves `김민준.` as name text the surname site divides — given 김, middle `민준. John`, family Smith, where 2.2.0 and `d37b8ec` read title `김민준.` — pinned as the tolerated `ko_name_with_a_period_in_a_bracketed_credential`. And a caller-authored NFD surname entry is composed at ingest, so raw NFD input then finds no entry and the name goes unsplit; pinned by `test_an_nfd_authored_entry_is_stored_composed`. - **Pickle compatibility, decided 2026-09-10 (Derek): a release-note line, not a forgiving load.** `Lexicon.__setstate__` re-runs `_normalize` over the loaded state and rejects any entry the fold changes, so a `Lexicon` pickled by 2.1.x or 2.2.x whose caller-added entries the widened fold now touches (a non-ASCII entry written with a wide stop such as `씨。`, or authored in NFD) raises `ValueError: incompatible Lexicon pickle: entries are not normalized (titles: 씨。); this state was not written by this version of nameparser` — a message that is wrong about provenance for exactly these pickles, which a released version did write. Measured 2026-09-10 (recompute: take `Lexicon.default().add(titles={"씨"}).__getstate__()`, add `씨。` to its `titles`, and load it into `Lexicon.__new__(Lexicon)`). The shipped vocabulary is unaffected (zero of 1735 strings change under NFC, and no shipped entry carries a wide stop). Two remedies were weighed: a release-note line with the fix a caller applies (rebuild the `Lexicon` from its source rather than unpickle it), or a load path that accepts state whose only drift is one fold pass away, re-folds it and warns naming the entries. The note was chosen: the load guard was designed to refuse rewritten data (the guarded-raise design #3-0-reevaluations records as right regardless of the in-a-minor friction) rather than become a fourth place caller data is corrected without a word, the population is caller-added CJK entries pickled across a minor, and a rebuild is the documented way to carry a `Lexicon` across versions. The 2.3.0 release note carries the line; `tests/v2/test_lexicon.py::test_unpickling_rejects_unnormalized_entries` is the pin of the refusal itself. - **CORRECTIONS this entry makes to entries above it.** #indic-honorifics' "Abbreviation marks and periods" bullet said `_normalize` "lowercases and strips edge whitespace and ASCII periods only — no NFC, no NFKC, no casefold anywhere in that module"; it now composes NFC for a non-ASCII word and strips all four stops. The VISARGA ARGUMENT still holds and holds for the same reason: ঃ U+0983 has no canonical decomposition, so NFC leaves it exactly where it was — recompute with that bullet's own one-liner, `Lexicon.default().add(titles={"মোঃ"}).titles & {"মোঃ"}`, still non-empty. #W1's 2026-07-29 §1a bullet said "segmentation MATCHING stays raw" and called the classification fold "the one deliberate exception"; matching is now folded too, so what stays raw is SEGMENTATION — the surname site's membership test and the peel's tail slice, which index the token's own text — and the folds are two, not one. #cjk-comma-demotion's 2026-09-05 five-parse bullet is amended in place: `田中さん 太郎.` no longer flips the order (the period is invisible to the classification, so it reads family `田中さん`, given `太郎.` as its no-period twin does), and `田中さん.` and `김민준씨.` no longer read as titles, so that bullet's summary — that the ONLY one of the five where a period changes nothing is the one where it sits on a word the split-off steps past — is superseded: the period changes nothing in four of the five now, and `田中さん 様.` is joined rather than alone. #initials-repertoire's principle bullet restated the veto phonologically ("morphemes or syllables"), which `_policy._NO_INITIALS`' own comment forbids in as many words — Devanagari is an abugida and Arabic an abjad, neither has letters in that sense and both abbreviate — so the bullet now states the CLDR criterion the constant actually uses. - **What moved, measured 2026-09-10 and re-measured the same day after the review round.** SEVENTEEN names in `tools/differential/corpus*.jsonl` read differently against `d37b8ec`, and all seventeen are the bundle's own new rows, every one of them on the RADAR tier in `corpus_cjk_tolerated.jsonl`. Twelve are the CJK movers this bundle's landing commits wrote; the thirteenth, the wholly-katakana `マイケル.`, joined in the whole-branch review that produced this entry, reading given `マイケル.` where `d37b8ec` reads title `マイケル.` — the katakana arm of the same H2 veto, pinned as `ja_katakana_lone_name_with_a_period_is_not_a_title` in `tests/v2/cases.py`. The last FOUR joined in the review round after that, as pins on readings nothing held: three stop-bearing FAMILY_COMMA spellings (`김민준씨., J.씨`, `田中さん., V.`, `이, J.씨.`) and the bracketed `(김민준.) John Smith`, the recorded degradation in the limits bullet above. A leading-stop shape moves the same way but sits in no corpus at all: `.김민준씨.` — a stop no script ever writes, glued before a honorific-bearing word — read one whole given `.김민준씨.` at `d37b8ec` and now reads given `.김민준`, suffix `씨.`, pinned only by the stage tests for its two edges taken separately (`test_the_peel_reads_the_trailing_stop_only` for the leading stop, `test_peels_a_listed_tail_through_a_trailing_full_stop` for the trailing one) rather than by any corpus row or a combined case of its own. The three comma rows agree with their stop-less twins in every field once the riding stop is removed from the suffix (`씨., J.씨` vs `씨, J.씨`; `さん.` vs `さん`; `씨.` vs `씨`) — each differs from its twin in exactly that one field, by exactly the stop that rides — which is the peel's own "one name, two spellings" argument holding as far as a raw field comparison can show it; the bracketed row does not agree with anything and is not meant to. `田中. 太郎` stays there with the rest, though it is what WITNESSES rules.md#H2's Accepted clause added here: `tools/differential/compare.py`'s tier comment records the 2026-09-05 precedent for exactly this collision — where a demoted file holds a text a rules.md example line names, the EXAMPLE LINE moves into the tolerated rule and the row is not promoted, because marking the row alone leaves the name enforced and documented as demoted. So the clause's CJK reading is carried in W3's tolerated example block, H2's own block keeps the Latin control `Smith. John`, and all four #323 shape rows stay `tolerated`. THE POPULATION THAT COULD HAVE MOVED AND DID NOT is two different sets depending on which criterion is read, and both are worth having in front of you. Under the SHAPE this bundle is about — an edge full stop glued to a word carrying a classified character — the corpora hold 22 names: the seventeen movers and FIVE non-movers (`田中さん 様.`, `田中さん, 様.`, `김민준 씨.`, `김민준 양.`, `김민준, 씨.`), all five already in `corpus_cjk_tolerated.jsonl` before this bundle. Two of the twenty-two need the detector to be written carefully, which is why the recipe is spelled out below: the stop on `김민준씨., J.씨` and `田中さん., V.` sits before a COMMA, so a detector splitting on whitespace alone sees `김민준씨.,` and finds no edge stop (an 18 recorded earlier in the day came from such a detector, over a corpus four names smaller), and the stop on `(김민준.)` is inside a bracketed clause until rules.md#S1's escape unwraps it. Under the looser reading of an ASCII PERIOD ANYWHERE in a CJK-bearing name, the corpora held 20 distinct names on 21 rows before this bundle (19 rows in the tolerated file and 2 in `corpus_issues.jsonl`, `Dr 김민준씨, Jr.` being the one name on two of them) — unmoved by the katakana row, which is a name this bundle's own work wrote, not a pre-bundle count. The looser set is wider because it sweeps in periods sitting on LATIN tokens inside a CJK name — `毛 泽东 Dr.` and `田中さん, V.` — which is precisely what the veto is scoped not to touch, it reading the word and not the name. NEITHER the five nor the twenty moves — the seventeen movers are the whole of what did, and every one of them is a row this bundle wrote. An earlier wording of this bullet said "the twenty-one CJK names carrying an ASCII period": that counted rows as names and quoted the looser criterion beside the shape's argument. RECOMPUTE: parse the union of the corpus files on both trees and diff the seven fields and the ambiguity kinds; for the SHAPE population, take each name's tokens (`_pipeline._tokenize`, not a whitespace split) and ask whether any token's `strip(FULL_STOPS)` is shorter and still carries a classified character, then add the bracketed clause by hand; for the looser one, count distinct names and rows separately. The gate exits 0 at all four baselines with `radar unclassified: 0`; one ledger rule carries them, `fix(#322/#323)`, ELEVEN members at 1.4.0 and 2.0.0 and SEVENTEEN at 2.1.0 and 2.2.0 — the three hangul names the older ledgers hand to their native-script CJK rule are one part of the difference, 2.1.0 being the release that shipped that behavior, and the three FAMILY_COMMA rows are the other: at 2.0.0 none of the three diffs at all, and at 1.4.0 the two that do (`田中さん., V.`, `이, J.씨.`) are already claimed by the broad `fix(cjk-comma-compound)` and `fix(cjk-comma-honorific-peel)` rules that ledger carries, so adding them to this rule would only take a name off a rule that describes it — the same narrow-first reasoning that keeps the three hangul names out. diff --git a/docs/design/rules.md b/docs/design/rules.md index 69054647..33cbf98a 100644 --- a/docs/design/rules.md +++ b/docs/design/rules.md @@ -116,9 +116,8 @@ H2. Rationale: before a name, an abbreviation is almost always a Accepted: the shape reads a Latin convention, and a script with no initials has no period abbreviations either, so a period- marked opening word carrying a Han, kana or hangul character — - as the script table classifies them; halfwidth katakana sits - outside it and still reads by the Latin shape, the limit - decisions.md#cjk-full-stops records — is a name word and never a + as the script table classifies them, halfwidth katakana + included since #594 — is a name word and never a title by shape (#323; decisions.md#cjk-full-stops) — the same veto that keeps 씨. from reading as an initial. The Latin reading is unchanged, and W3's example block carries the CJK diff --git a/nameparser/_policy.py b/nameparser/_policy.py index bbed93b3..b9052166 100644 --- a/nameparser/_policy.py +++ b/nameparser/_policy.py @@ -134,8 +134,11 @@ class Script(StrEnum): # 0201 systems spells names in it, and 山田 タロウ must read as 山田 タロウ # does. Classifying it is a range, not a fold -- T1 forbids rewriting # the text, and an NFKC fold at classification would reach far past -# the kana (fullwidth Latin, ㈱) while mapping a lone voicing mark -# ゙ to the HIRAGANA block's combining U+3099. The span takes in the +# the kana (fullwidth Latin, ㈱), and a voicing mark with no +# precomposed form to join (ア゙) folds to the combining U+3099, which +# is in the HIRAGANA block -- so the folded token would take the kana +# license and a wholly-katakana name would turn family-first +# (decisions.md#W4). The span takes in the # halfwidth nakaguro U+FF65, which tokenize turns into a separator, # on the same direct-call grounds as U+30FB below, and the voicing # marks U+FF9E/U+FF9F, which are spacing characters following their From 6803db66667a9920b39033f6e7ccb4cfdcd0c790 Mon Sep 17 00:00:00 2001 From: Derek Gulbranson Date: Sat, 3 Oct 2026 13:57:01 -0700 Subject: [PATCH 5/9] docs: give the NFKC sweep a recompute instead of a bare count (#594) Co-Authored-By: Claude Opus 5.5 --- docs/design/decisions.md | 4 ++-- 1 file changed, 2 insertions(+), 2 deletions(-) diff --git a/docs/design/decisions.md b/docs/design/decisions.md index d3d8664f..008f0f6f 100644 --- a/docs/design/decisions.md +++ b/docs/design/decisions.md @@ -468,7 +468,7 @@ Declined: - 2026-07-27 (script-scoped order amendment) — the family-first override is keyed to the SCRIPT of the written name, never to a guessed language: wholly-Han, wholly-Hangul, and kana-licensed Japanese read family-first because zh, ko and ja all write family-first in native script; wholly-katakana names are predominantly transcriptions and keep the declared order. Latin transliterations are never touched. - 2026-07-29 #272 — the kana license: Han∪kana with at least one kana cannot be Chinese and is not a transcription, so 高橋みなみ reads family-first though it is written in two scripts. -- 2026-10-03 #594 — HALFWIDTH KATAKANA IS KATAKANA. The script table classifies the halfwidth kana block U+FF65–FF9F as `Script.KATAKANA`, so `山田 タロウ` reads family 山田 as `山田 エミ` does; every release through 2.3.0 left the block unclassified and read it given 山田, family タロウ, the positional reading. Legacy data written for JIS X 0201 systems, which had no kanji, spells names this way — bank, payroll and CSV exports. DECIDED by widening the table rather than by folding: the T rules forbid rewriting the text (decisions.md#T1), so normalization at tokenize was never available, and an NFKC fold at classification — the other place the issue offered — reaches far past the kana (`JOHN` to `JOHN`, `㈱` to `(株)`), and it would move a wholly-katakana name across W4's line: a voicing mark with no precomposed form to join (`ア゙`, `ン゙`) folds to the COMBINING U+3099, which sits in the HIRAGANA block, so the folded token reads as kana-licensed Japanese and `ア゙イ タロウ` would read family-first where the unfolded one keeps the declared order (measured 2026-10-03 by parsing the NFKC spelling `ア゙イ タロウ`, which reads family `ア゙イ`; the review's two-input sweep of halfwidth against folded spellings, 600 pairs, differed only on such marks). The interpunct flank guard is NOT the hazard, though an earlier wording of this bullet said it was: it reads raw text and asks only whether a character is classified at all, which hiragana and katakana both are. The same measurement shows a pre-existing gap the fold would have widened rather than created: full-width katakana typed with an uncomposable combining U+3099 already takes the kana license today. A block range is how the table already decides (U+30FB and the voicing marks are in by block), and it takes the voicing marks U+FF9E/U+FF9F in because they are spacing characters following their base. The span starts at U+FF65, the halfwidth nakaguro, which tokenize already treats as a separator (rules.md#T2), and stops short of U+FF61–FF64, whose U+FF61 is a full stop. Everything keyed on the table follows: the kana license, the second-East-Asian-word count that keeps a segmenter from being asked (`高橋一郎 タロウ`), the 间隔号's flank guard (`タロウ·ヤマダ` divides), and `_NO_INITIALS`, which closes the limit #cjk-full-stops recorded (`タナカ. John` read title `タナカ.`). THE RATIONALE FOR WHOLLY-KATAKANA NAMES IS AMENDED, NOT THE BEHAVIOR: a wholly-halfwidth name keeps the declared order exactly as a wholly-katakana one does (`ヤマダ タロウ` reads given ヤマダ), but the 2026-07-27 reason — that such names are predominantly transcriptions — is weaker here, since a system with no kanji wrote Japanese names in katakana too, family-first. What stands is the narrower claim: the script cannot tell a Japanese reading from a transcription, so it settles no order, and a caller who knows their katakana is Japanese adds `(Script.KATAKANA, FAMILY_FIRST)` to `script_orders` (docs/usage.rst's Boundaries section carries the recipe). `name_order=FAMILY_FIRST` works too, at the cost of reversing Latin names as well. MEASURED 2026-10-03: the differential gate saw nothing before this change because no corpus line held halfwidth kana; the six case rows #594 added are its whole population, three of them contract movers (`山田 タロウ`, `山田 ダイスケ`, `タロウ·ヤマダ`), two tolerated radar movers (`タナカ. John`, `マイケル.`) and one parity row (`ヤマダ タロウ`). RECOMPUTE: parse the six on this tree and on 2.3.0 from PyPI and read `title` beside the name fields. +- 2026-10-03 #594 — HALFWIDTH KATAKANA IS KATAKANA. The script table classifies the halfwidth kana block U+FF65–FF9F as `Script.KATAKANA`, so `山田 タロウ` reads family 山田 as `山田 エミ` does; every release through 2.3.0 left the block unclassified and read it given 山田, family タロウ, the positional reading. Legacy data written for JIS X 0201 systems, which had no kanji, spells names this way — bank, payroll and CSV exports. DECIDED by widening the table rather than by folding: the T rules forbid rewriting the text (decisions.md#T1), so normalization at tokenize was never available, and an NFKC fold at classification — the other place the issue offered — reaches far past the kana (`JOHN` to `JOHN`, `㈱` to `(株)`), and it would move a wholly-katakana name across W4's line: a voicing mark with no precomposed form to join (`ア゙`, `ン゙`) folds to the COMBINING U+3099, which sits in the HIRAGANA block, so the folded token reads as kana-licensed Japanese and `ア゙イ タロウ` would read family-first where the unfolded one keeps the declared order (measured 2026-10-03 by parsing the NFKC spelling `ア゙イ タロウ`, which reads family `ア゙イ`; a two-input sweep of halfwidth names against their NFC(NFKC) spellings differed only where such a mark stood; RECOMPUTE: parse each halfwidth name and `unicodedata.normalize("NFC", unicodedata.normalize("NFKC", name))` and compare the fields). The interpunct flank guard is NOT the hazard, though an earlier wording of this bullet said it was: it reads raw text and asks only whether a character is classified at all, which hiragana and katakana both are. The same measurement shows a pre-existing gap the fold would have widened rather than created: full-width katakana typed with an uncomposable combining U+3099 already takes the kana license today. A block range is how the table already decides (U+30FB and the voicing marks are in by block), and it takes the voicing marks U+FF9E/U+FF9F in because they are spacing characters following their base. The span starts at U+FF65, the halfwidth nakaguro, which tokenize already treats as a separator (rules.md#T2), and stops short of U+FF61–FF64, whose U+FF61 is a full stop. Everything keyed on the table follows: the kana license, the second-East-Asian-word count that keeps a segmenter from being asked (`高橋一郎 タロウ`), the 间隔号's flank guard (`タロウ·ヤマダ` divides), and `_NO_INITIALS`, which closes the limit #cjk-full-stops recorded (`タナカ. John` read title `タナカ.`). THE RATIONALE FOR WHOLLY-KATAKANA NAMES IS AMENDED, NOT THE BEHAVIOR: a wholly-halfwidth name keeps the declared order exactly as a wholly-katakana one does (`ヤマダ タロウ` reads given ヤマダ), but the 2026-07-27 reason — that such names are predominantly transcriptions — is weaker here, since a system with no kanji wrote Japanese names in katakana too, family-first. What stands is the narrower claim: the script cannot tell a Japanese reading from a transcription, so it settles no order, and a caller who knows their katakana is Japanese adds `(Script.KATAKANA, FAMILY_FIRST)` to `script_orders` (docs/usage.rst's Boundaries section carries the recipe). `name_order=FAMILY_FIRST` works too, at the cost of reversing Latin names as well. MEASURED 2026-10-03: the differential gate saw nothing before this change because no corpus line held halfwidth kana; the six case rows #594 added are its whole population, three of them contract movers (`山田 タロウ`, `山田 ダイスケ`, `タロウ·ヤマダ`), two tolerated radar movers (`タナカ. John`, `マイケル.`) and one parity row (`ヤマダ タロウ`). RECOMPUTE: parse the six on this tree and on 2.3.0 from PyPI and read `title` beside the name fields. ### D1 — the segmenterless-activation warning @@ -1515,7 +1515,7 @@ ONE GAP SEEN FROM TWO SIDES, which is why the two issues were answered together. - **#323 turned out broader than it was filed.** The issue named the honorific peel. Making the classification fold read through an edge stop moved three readers of the `None` it used to return, and only one of them was the peel's neighbour test: the SURNAME SITE stepped past the family name onto the given name (`양. 지훈` cut 지훈 in half — given `양.`, middle `지`, family `훈` — and now reads given `지훈`, family `양.`), the ORDER RULE fell back to positional (`양 지훈.` lost family-first and now keeps it; `田中 太郎.` likewise), and the segmenter's neighbour precondition missed a writer-drawn boundary (`山田太郎 田中.` consulted a pluggable segmenter on `山田太郎` as if it stood alone, and is blocked now). The issue weighed two candidate fixes and chose neither: teaching the surname site to consult `is_suffix_strict` beside `effective_script`, or making the classification tolerant of a trailing period, which it called the broader one and expected to interact with #322. The broader one shipped, and the narrow one would have reached the filed name alone — `김.` is no suffix vocabulary, so `김. 민준` sits outside it entirely. - **The surname-site amendment was found by measurement during execution, not designed.** With the classification fold alone, `김. 민준` read given `김`, middle `.`, family `민준`: the token classified as hangul, the site matched `김` against its own head, and what was left over — the stop — BECAME the remainder. So the site matches on `text.rstrip(FULL_STOPS)` and cuts the core, a head being a prefix, which makes the offset that cuts the core cut the text; the stop rides with the remainder (`김민준.` divides as `김` + `민준.`, `김. 민준` as family `김.` plus given `민준`). `rstrip`, not `strip`: a LEADING stop would break the prefix argument. The reading is recorded in `ebf64db`'s comment at the site. - **Two readings JOINED the bundle, approved in session on 2026-09-09.** Neither was filed. FIRST, the glued-period honorific (rules.md#W3): a stop on the honorific's own word stood between the listed tail and the token's end, so nothing peeled and `田中さん.` and `김민준씨.` read as titles. The peel now matches the tail on the core and cuts BEFORE it, so both divide where their stop-less spellings do and the stop rides with the honorific. The move was free because W3 had recorded that reading as measured on 2026-09-05 and pinned by nothing, its own words being that neither string is a case row or a corpus line, so no row pinned those two readings; both are case rows now, tolerated. SECOND, H2's opening-abbreviation shape on a CJK word: `田中.` read title and now reads family `田中.`, `田中. 太郎` reads family `田中.`, given `太郎`. Han is what WITNESSES this one, and hangul cannot: hangul segmentation runs first, so `김민준.` never reaches the shape test and divides as `김` + `민준.` by the surname site instead. Latin and Cyrillic are untouched (`Smith. John` still reads title `Smith.`, `Проф. Иванов` title `Проф.`), and the veto is contains-any, so `Kim김. Smith` is refused a title too — deliberately: a word carrying a script with no abbreviations is not wearing an abbreviation's period. -- **Recorded limits, all five measured 2026-09-10 on this tree — and the segmenter one is a limit the review round CLOSED, kept in this bullet because this is where the claim stood. The halfwidth-kana one was CLOSED on 2026-10-03 by #594; see the CLOSED note after its paragraph below, and decisions.md#W4.** HALFWIDTH KANA sits outside the script table, so it reads by the LATIN shapes at both period sites and not just at the initial one: `_vocab.is_initial("ラ.")` is `True` where `is_initial("ラ.")` is `False`, and H2's opening-abbreviation shape MISROUTES a longer halfwidth opener — `parse("タナカ. John")` gives title `タナカ.` where `parse("タナカ. John")` gives given `タナカ.`. Both because `_policy._SCRIPT_RANGES` excludes the halfwidth kana block U+FF65–FF9F on purpose — legacy bank and CSV data writes it, and the table's comment calls that a separate normalization problem, #272's separator handling touching only the halfwidth DOT U+FF65 and none of the kana behind it — and `_NO_INITIALS`, being a tuple of Script members, silently inherits the exclusion. That is #322's second mechanism, restated here because it is the one the bundle did NOT close. An earlier wording of this bullet said "nothing misroutes today, every kana honorific being multi-character", which had the direction backwards: the misroute is at the SHAPE site rather than the peel, and being multi-character is what carries a halfwidth opener INTO H2's shape rather than what keeps it out. NOTHING PINS EITHER READING — no case row, no corpus line, for the halfwidth block or for `タナカ. John` — so both can move without the suite or the differential saying so. It stays a recorded limit rather than a fix: what to do with halfwidth kana is a normalization question (widen the script table, or fold widths at classification) and not a full-stop one, and the bundle answered the full-stop one. CLOSED 2026-10-03 by #594, which widened the table (decisions.md#W4 carries why not the fold): `タナカ. John` reads given `タナカ.` and `is_initial("ラ.")` is `False`, and both readings are pinned now — the case rows `ja_halfwidth_katakana_opener_with_a_period_is_a_name` and `ja_halfwidth_katakana_lone_name_with_a_period_is_not_a_title`, and `test_is_initial_script_repertoire`. RECOMPUTE: parse `タナカ. John` and `タナカ. John` and read `title` beside the name fields; `is_initial` alone reports only the shorter half. A PLUGGABLE SEGMENTER receives the CORE, the token with its trailing stops removed — the same string the vocabulary match reads, the surname site handing it `core` rather than `text` (corrected 2026-09-10, in the review round after this entry was first written: the site handed over the raw token, which this bullet recorded as a limit, and the fix costs nothing because a head is a prefix — every offset reported against the core cuts `text` in the same place, so the stops ride with the last piece and no answer of the segmenter's can make a stop a piece of its own). The shipped `ja` pack's `_wholly_japanese` therefore ACCEPTS `山田太郎.` now, where the raw token failed its repertoire test: measured 2026-09-10 with `namedivider-python` installed (`uv run --extra ja`), `parser_for(locales.JA, segmenter=locales.ja_segmenter())` reads `山田太郎.` as family 山田, given `太郎.` with a SEGMENTATION report scoring 0.44, where `d37b8ec` read title `山田太郎.` and the un-corrected branch read family `山田太郎.`; the already-divided `山田 太郎.` reads family 山田, given `太郎.` on this tree and given 山, middle 田, family `太郎.` at `d37b8ec`. Pinned by `test_a_consulted_segmenter_receives_the_core_without_the_stop`. A LEADING stop keeps the token out of the SURNAME site only, because the classification fold that gates it rstrips as it does: `.김민준` classifies as no script, never becomes a surname site and stays one whole word, reading given `.김민준` exactly as `d37b8ec` reads it (corrected 2026-09-10 in the same round: the fold stripped both edges when this entry was written, which moved the ROLE to family and — with a segmenter configured — offered the token to it raw, an answer of offset 1 dividing it into the stop and the name; the fold now matches the sites, and hiding such a token is the no-split rules.md#W1 already accepts). It does NOT keep the token out of the HONORIFIC PEEL, which consults no script at all — the tail alone is its license (rules.md#W2's Background) — so a leading stop does not defeat it: `.김민준씨` peels to `.김민준` and `씨` exactly as `김민준씨` does, pinned by `test_the_peel_reads_the_trailing_stop_only`. A BRACKETED CJK CREDENTIAL degrades: rules.md#S1's escape in `_extract` calls a clause suffix-shaped when it ends in an ASCII period, so `(김민준.) John Smith` is unwrapped and the H2 veto then leaves `김민준.` as name text the surname site divides — given 김, middle `민준. John`, family Smith, where 2.2.0 and `d37b8ec` read title `김민준.` — pinned as the tolerated `ko_name_with_a_period_in_a_bracketed_credential`. And a caller-authored NFD surname entry is composed at ingest, so raw NFD input then finds no entry and the name goes unsplit; pinned by `test_an_nfd_authored_entry_is_stored_composed`. +- **Recorded limits, all five measured 2026-09-10 on this tree — and the segmenter one is a limit the review round CLOSED, kept in this bullet because this is where the claim stood. The halfwidth-kana one was CLOSED on 2026-10-03 by #594; see the CLOSED note that ends the halfwidth passage below, and decisions.md#W4.** HALFWIDTH KANA sits outside the script table, so it reads by the LATIN shapes at both period sites and not just at the initial one: `_vocab.is_initial("ラ.")` is `True` where `is_initial("ラ.")` is `False`, and H2's opening-abbreviation shape MISROUTES a longer halfwidth opener — `parse("タナカ. John")` gives title `タナカ.` where `parse("タナカ. John")` gives given `タナカ.`. Both because `_policy._SCRIPT_RANGES` excludes the halfwidth kana block U+FF65–FF9F on purpose — legacy bank and CSV data writes it, and the table's comment calls that a separate normalization problem, #272's separator handling touching only the halfwidth DOT U+FF65 and none of the kana behind it — and `_NO_INITIALS`, being a tuple of Script members, silently inherits the exclusion. That is #322's second mechanism, restated here because it is the one the bundle did NOT close. An earlier wording of this bullet said "nothing misroutes today, every kana honorific being multi-character", which had the direction backwards: the misroute is at the SHAPE site rather than the peel, and being multi-character is what carries a halfwidth opener INTO H2's shape rather than what keeps it out. NOTHING PINS EITHER READING — no case row, no corpus line, for the halfwidth block or for `タナカ. John` — so both can move without the suite or the differential saying so. It stays a recorded limit rather than a fix: what to do with halfwidth kana is a normalization question (widen the script table, or fold widths at classification) and not a full-stop one, and the bundle answered the full-stop one. CLOSED 2026-10-03 by #594, which widened the table (decisions.md#W4 carries why not the fold): `タナカ. John` reads given `タナカ.` and `is_initial("ラ.")` is `False`, and both readings are pinned now — the case rows `ja_halfwidth_katakana_opener_with_a_period_is_a_name` and `ja_halfwidth_katakana_lone_name_with_a_period_is_not_a_title`, and `test_is_initial_script_repertoire`. RECOMPUTE: parse `タナカ. John` and `タナカ. John` and read `title` beside the name fields; `is_initial` alone reports only the shorter half. A PLUGGABLE SEGMENTER receives the CORE, the token with its trailing stops removed — the same string the vocabulary match reads, the surname site handing it `core` rather than `text` (corrected 2026-09-10, in the review round after this entry was first written: the site handed over the raw token, which this bullet recorded as a limit, and the fix costs nothing because a head is a prefix — every offset reported against the core cuts `text` in the same place, so the stops ride with the last piece and no answer of the segmenter's can make a stop a piece of its own). The shipped `ja` pack's `_wholly_japanese` therefore ACCEPTS `山田太郎.` now, where the raw token failed its repertoire test: measured 2026-09-10 with `namedivider-python` installed (`uv run --extra ja`), `parser_for(locales.JA, segmenter=locales.ja_segmenter())` reads `山田太郎.` as family 山田, given `太郎.` with a SEGMENTATION report scoring 0.44, where `d37b8ec` read title `山田太郎.` and the un-corrected branch read family `山田太郎.`; the already-divided `山田 太郎.` reads family 山田, given `太郎.` on this tree and given 山, middle 田, family `太郎.` at `d37b8ec`. Pinned by `test_a_consulted_segmenter_receives_the_core_without_the_stop`. A LEADING stop keeps the token out of the SURNAME site only, because the classification fold that gates it rstrips as it does: `.김민준` classifies as no script, never becomes a surname site and stays one whole word, reading given `.김민준` exactly as `d37b8ec` reads it (corrected 2026-09-10 in the same round: the fold stripped both edges when this entry was written, which moved the ROLE to family and — with a segmenter configured — offered the token to it raw, an answer of offset 1 dividing it into the stop and the name; the fold now matches the sites, and hiding such a token is the no-split rules.md#W1 already accepts). It does NOT keep the token out of the HONORIFIC PEEL, which consults no script at all — the tail alone is its license (rules.md#W2's Background) — so a leading stop does not defeat it: `.김민준씨` peels to `.김민준` and `씨` exactly as `김민준씨` does, pinned by `test_the_peel_reads_the_trailing_stop_only`. A BRACKETED CJK CREDENTIAL degrades: rules.md#S1's escape in `_extract` calls a clause suffix-shaped when it ends in an ASCII period, so `(김민준.) John Smith` is unwrapped and the H2 veto then leaves `김민준.` as name text the surname site divides — given 김, middle `민준. John`, family Smith, where 2.2.0 and `d37b8ec` read title `김민준.` — pinned as the tolerated `ko_name_with_a_period_in_a_bracketed_credential`. And a caller-authored NFD surname entry is composed at ingest, so raw NFD input then finds no entry and the name goes unsplit; pinned by `test_an_nfd_authored_entry_is_stored_composed`. - **Pickle compatibility, decided 2026-09-10 (Derek): a release-note line, not a forgiving load.** `Lexicon.__setstate__` re-runs `_normalize` over the loaded state and rejects any entry the fold changes, so a `Lexicon` pickled by 2.1.x or 2.2.x whose caller-added entries the widened fold now touches (a non-ASCII entry written with a wide stop such as `씨。`, or authored in NFD) raises `ValueError: incompatible Lexicon pickle: entries are not normalized (titles: 씨。); this state was not written by this version of nameparser` — a message that is wrong about provenance for exactly these pickles, which a released version did write. Measured 2026-09-10 (recompute: take `Lexicon.default().add(titles={"씨"}).__getstate__()`, add `씨。` to its `titles`, and load it into `Lexicon.__new__(Lexicon)`). The shipped vocabulary is unaffected (zero of 1735 strings change under NFC, and no shipped entry carries a wide stop). Two remedies were weighed: a release-note line with the fix a caller applies (rebuild the `Lexicon` from its source rather than unpickle it), or a load path that accepts state whose only drift is one fold pass away, re-folds it and warns naming the entries. The note was chosen: the load guard was designed to refuse rewritten data (the guarded-raise design #3-0-reevaluations records as right regardless of the in-a-minor friction) rather than become a fourth place caller data is corrected without a word, the population is caller-added CJK entries pickled across a minor, and a rebuild is the documented way to carry a `Lexicon` across versions. The 2.3.0 release note carries the line; `tests/v2/test_lexicon.py::test_unpickling_rejects_unnormalized_entries` is the pin of the refusal itself. - **CORRECTIONS this entry makes to entries above it.** #indic-honorifics' "Abbreviation marks and periods" bullet said `_normalize` "lowercases and strips edge whitespace and ASCII periods only — no NFC, no NFKC, no casefold anywhere in that module"; it now composes NFC for a non-ASCII word and strips all four stops. The VISARGA ARGUMENT still holds and holds for the same reason: ঃ U+0983 has no canonical decomposition, so NFC leaves it exactly where it was — recompute with that bullet's own one-liner, `Lexicon.default().add(titles={"মোঃ"}).titles & {"মোঃ"}`, still non-empty. #W1's 2026-07-29 §1a bullet said "segmentation MATCHING stays raw" and called the classification fold "the one deliberate exception"; matching is now folded too, so what stays raw is SEGMENTATION — the surname site's membership test and the peel's tail slice, which index the token's own text — and the folds are two, not one. #cjk-comma-demotion's 2026-09-05 five-parse bullet is amended in place: `田中さん 太郎.` no longer flips the order (the period is invisible to the classification, so it reads family `田中さん`, given `太郎.` as its no-period twin does), and `田中さん.` and `김민준씨.` no longer read as titles, so that bullet's summary — that the ONLY one of the five where a period changes nothing is the one where it sits on a word the split-off steps past — is superseded: the period changes nothing in four of the five now, and `田中さん 様.` is joined rather than alone. #initials-repertoire's principle bullet restated the veto phonologically ("morphemes or syllables"), which `_policy._NO_INITIALS`' own comment forbids in as many words — Devanagari is an abugida and Arabic an abjad, neither has letters in that sense and both abbreviate — so the bullet now states the CLDR criterion the constant actually uses. - **What moved, measured 2026-09-10 and re-measured the same day after the review round.** SEVENTEEN names in `tools/differential/corpus*.jsonl` read differently against `d37b8ec`, and all seventeen are the bundle's own new rows, every one of them on the RADAR tier in `corpus_cjk_tolerated.jsonl`. Twelve are the CJK movers this bundle's landing commits wrote; the thirteenth, the wholly-katakana `マイケル.`, joined in the whole-branch review that produced this entry, reading given `マイケル.` where `d37b8ec` reads title `マイケル.` — the katakana arm of the same H2 veto, pinned as `ja_katakana_lone_name_with_a_period_is_not_a_title` in `tests/v2/cases.py`. The last FOUR joined in the review round after that, as pins on readings nothing held: three stop-bearing FAMILY_COMMA spellings (`김민준씨., J.씨`, `田中さん., V.`, `이, J.씨.`) and the bracketed `(김민준.) John Smith`, the recorded degradation in the limits bullet above. A leading-stop shape moves the same way but sits in no corpus at all: `.김민준씨.` — a stop no script ever writes, glued before a honorific-bearing word — read one whole given `.김민준씨.` at `d37b8ec` and now reads given `.김민준`, suffix `씨.`, pinned only by the stage tests for its two edges taken separately (`test_the_peel_reads_the_trailing_stop_only` for the leading stop, `test_peels_a_listed_tail_through_a_trailing_full_stop` for the trailing one) rather than by any corpus row or a combined case of its own. The three comma rows agree with their stop-less twins in every field once the riding stop is removed from the suffix (`씨., J.씨` vs `씨, J.씨`; `さん.` vs `さん`; `씨.` vs `씨`) — each differs from its twin in exactly that one field, by exactly the stop that rides — which is the peel's own "one name, two spellings" argument holding as far as a raw field comparison can show it; the bracketed row does not agree with anything and is not meant to. `田中. 太郎` stays there with the rest, though it is what WITNESSES rules.md#H2's Accepted clause added here: `tools/differential/compare.py`'s tier comment records the 2026-09-05 precedent for exactly this collision — where a demoted file holds a text a rules.md example line names, the EXAMPLE LINE moves into the tolerated rule and the row is not promoted, because marking the row alone leaves the name enforced and documented as demoted. So the clause's CJK reading is carried in W3's tolerated example block, H2's own block keeps the Latin control `Smith. John`, and all four #323 shape rows stay `tolerated`. THE POPULATION THAT COULD HAVE MOVED AND DID NOT is two different sets depending on which criterion is read, and both are worth having in front of you. Under the SHAPE this bundle is about — an edge full stop glued to a word carrying a classified character — the corpora hold 22 names: the seventeen movers and FIVE non-movers (`田中さん 様.`, `田中さん, 様.`, `김민준 씨.`, `김민준 양.`, `김민준, 씨.`), all five already in `corpus_cjk_tolerated.jsonl` before this bundle. Two of the twenty-two need the detector to be written carefully, which is why the recipe is spelled out below: the stop on `김민준씨., J.씨` and `田中さん., V.` sits before a COMMA, so a detector splitting on whitespace alone sees `김민준씨.,` and finds no edge stop (an 18 recorded earlier in the day came from such a detector, over a corpus four names smaller), and the stop on `(김민준.)` is inside a bracketed clause until rules.md#S1's escape unwraps it. Under the looser reading of an ASCII PERIOD ANYWHERE in a CJK-bearing name, the corpora held 20 distinct names on 21 rows before this bundle (19 rows in the tolerated file and 2 in `corpus_issues.jsonl`, `Dr 김민준씨, Jr.` being the one name on two of them) — unmoved by the katakana row, which is a name this bundle's own work wrote, not a pre-bundle count. The looser set is wider because it sweeps in periods sitting on LATIN tokens inside a CJK name — `毛 泽东 Dr.` and `田中さん, V.` — which is precisely what the veto is scoped not to touch, it reading the word and not the name. NEITHER the five nor the twenty moves — the seventeen movers are the whole of what did, and every one of them is a row this bundle wrote. An earlier wording of this bullet said "the twenty-one CJK names carrying an ASCII period": that counted rows as names and quoted the looser criterion beside the shape's argument. RECOMPUTE: parse the union of the corpus files on both trees and diff the seven fields and the ambiguity kinds; for the SHAPE population, take each name's tokens (`_pipeline._tokenize`, not a whitespace split) and ask whether any token's `strip(FULL_STOPS)` is shorter and still carries a classified character, then add the bracketed clause by hand; for the looser one, count distinct names and rows separately. The gate exits 0 at all four baselines with `radar unclassified: 0`; one ledger rule carries them, `fix(#322/#323)`, ELEVEN members at 1.4.0 and 2.0.0 and SEVENTEEN at 2.1.0 and 2.2.0 — the three hangul names the older ledgers hand to their native-script CJK rule are one part of the difference, 2.1.0 being the release that shipped that behavior, and the three FAMILY_COMMA rows are the other: at 2.0.0 none of the three diffs at all, and at 1.4.0 the two that do (`田中さん., V.`, `이, J.씨.`) are already claimed by the broad `fix(cjk-comma-compound)` and `fix(cjk-comma-honorific-peel)` rules that ledger carries, so adding them to this rule would only take a name off a rule that describes it — the same narrow-first reasoning that keeps the three hangul names out. From 1ea6e68fa2ce902bc848c84b16e2738cc9837f71 Mon Sep 17 00:00:00 2001 From: Derek Gulbranson Date: Sat, 3 Oct 2026 14:16:14 -0700 Subject: [PATCH 6/9] =?UTF-8?q?test:=20pin=20the=20span's=20=EF=BE=9F/?= =?UTF-8?q?=EF=BD=B0=20edges,=20the=20unspaced=20twin=20and=20the=20ja=20a?= =?UTF-8?q?dapter=20(#594)?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit From the PR review's coverage pass: - 山田 ペーター pins U+FF9F and U+FF70, which only the ledger sync test held: narrowing the range by either one now fails a behavior row. - 山田タロウ, the halfwidth twin of ja_unspaced_unsegmented_default, pins a single Han+halfwidth token taking the license (and 2.3's given-or-family report going away). - The ja adapter stub accepts 山田タロウ, the one changed path no test reached. The license rule in the 2.1/2.2/2.3 ledgers takes the two new movers, and the #594 ledger blocks are regenerated per baseline: the title rule's comment claimed the full-width twins read a name word, true only at 2.3.0, and the 1.4.0/2.0.0 headers described rules those files do not carry. _CORPUS_CLAIMS records the wider reach. Co-Authored-By: Claude Opus 5.5 --- tests/v2/cases.py | 17 +++++++++ tests/v2/test_ledger_guards.py | 38 +++++++++++--------- tests/v2/test_locales.py | 3 ++ tools/differential/corpus_cjk.jsonl | 2 ++ tools/differential/expected_since_1.4.0.toml | 37 +++++++++---------- tools/differential/expected_since_2.0.0.toml | 32 ++++++++--------- tools/differential/expected_since_2.1.0.toml | 30 +++++++++------- tools/differential/expected_since_2.2.0.toml | 30 +++++++++------- tools/differential/expected_since_2.3.0.toml | 35 ++++++++++-------- 9 files changed, 132 insertions(+), 92 deletions(-) diff --git a/tests/v2/cases.py b/tests/v2/cases.py index faa7f92d..28b26238 100644 --- a/tests/v2/cases.py +++ b/tests/v2/cases.py @@ -9141,6 +9141,23 @@ def _check_cjk_shape_purity(self) -> None: "only the block span classifies it; a range ending at " "U+FF9D would leave ダイスケ mixed-script and the name " "positional"), + Case("ja_halfwidth_semi_voiced_and_long_vowel_marks_are_katakana", + "山田 ペーター", + {"family": "山田", "given": "ペーター"}, + classification="fix(#594)", + notes="the span's other two edges a name actually reaches: " + "the semi-voiced mark ゚ (U+FF9F, the block's last " + "codepoint) and the prolonged sound mark ー (U+FF70). " + "Narrowing the range by either one leaves ペーター " + "mixed-script and the name positional"), + Case("ja_halfwidth_unspaced_unsegmented_default", "山田タロウ", + {"family": "山田タロウ"}, + classification="fix(#594)", + notes="the halfwidth twin of ja_unspaced_unsegmented_default: " + "one Han+halfwidth token takes the kana license on its " + "own and reads family. 2.3.0 read it given, reporting " + "given-or-family; the license decides it, so nothing is " + "reported now"), Case("ja_halfwidth_pure_katakana_positional", "ヤマダ タロウ", {"given": "ヤマダ", "family": "タロウ"}, notes="parity row guarding the license's boundary in " diff --git a/tests/v2/test_ledger_guards.py b/tests/v2/test_ledger_guards.py index b91ed016..74df9515 100644 --- a/tests/v2/test_ledger_guards.py +++ b/tests/v2/test_ledger_guards.py @@ -63,8 +63,8 @@ class declares, which members an alternation offers. Those are exact # table does not classify it. Empty since #594. Its one member was the # halfwidth middle dot U+FF65, which separates tokens without having # been classified while halfwidth kana stayed out of the table; #594 -# classified the whole halfwidth kana block, U+FF65 with it, and the -# membership guard below moved it into the table as it exists to. +# classified the whole halfwidth kana block, U+FF65 with it, and it +# moved into the table, as the membership guard below requires. # U+00B7 is deliberately NOT here -- its flank guard means every name # it can change matches through a classified flanking character # already. Single-sourced: read by the span sweep below, and by the @@ -3641,15 +3641,17 @@ def _claim(rule: dict) -> _Claim: # 2026-10-02, #585: 138 -> 139, one new corpus name and not a # wider rule: the decomposed katakana R3 row 'マイケル ジャクソン' # (NFD) lies in its script span, and fix(#585) explains it. - # 2026-10-03, #594: 139 -> 145, and this time the regex DID + # 2026-10-03, #594: 139 -> 147, and this time the regex DID # widen: its class copies the halfwidth kana block U+FF65-U+FF9F # whole, where it held U+FF65 alone. What it reaches beyond that - # is exactly the six halfwidth case rows #594 added -- no corpus - # line held halfwidth kana before them -- and it explains the - # three whose name parts move; 'ヤマダ タロウ' is parity, as - # 'マイケル ジャクソン' already was inside the same class. + # is exactly the eight halfwidth case rows #594 added -- no + # corpus line held halfwidth kana before them. It explains the + # five that move only name fields; the two title movers also + # move `title`, outside its fields, and go to the #594 rule; + # 'ヤマダ タロウ' is parity, as 'マイケル ジャクソン' already was + # inside the same class. "fix(#271/#272/#298) native-script CJK: family-first order, hangul segmentation, the kana license and the dots": - _Claim(145, ('family', 'given', 'middle'), "cfd0d19b46d5", None), + _Claim(147, ('family', 'given', 'middle'), "b2dd5ac30ae4", None), # 2026-09-19, #533: 33 -> 68. The count grew with the CORPUS # rather than with the rule -- this change added 35 # maiden-clause names as rules.md example lines and @@ -4797,15 +4799,17 @@ def _claim(rule: dict) -> _Claim: # 2026-10-02, #585: 138 -> 139, one new corpus name and not a # wider rule: the decomposed katakana R3 row 'マイケル ジャクソン' # (NFD) lies in its script span, and fix(#585) explains it. - # 2026-10-03, #594: 139 -> 145, and this time the regex DID + # 2026-10-03, #594: 139 -> 147, and this time the regex DID # widen: its class copies the halfwidth kana block U+FF65-U+FF9F # whole, where it held U+FF65 alone. What it reaches beyond that - # is exactly the six halfwidth case rows #594 added -- no corpus - # line held halfwidth kana before them -- and it explains the - # three whose name parts move; 'ヤマダ タロウ' is parity, as - # 'マイケル ジャクソン' already was inside the same class. + # is exactly the eight halfwidth case rows #594 added -- no + # corpus line held halfwidth kana before them. It explains the + # five that move only name fields; the two title movers also + # move `title`, outside its fields, and go to the #594 rule; + # 'ヤマダ タロウ' is parity, as 'マイケル ジャクソン' already was + # inside the same class. "fix(#271/#272/#298) native-script CJK: family-first order, hangul segmentation, the kana license and the dots": - _Claim(145, ('_ambiguities', 'family', 'given', 'middle'), "cfd0d19b46d5", None), + _Claim(147, ('_ambiguities', 'family', 'given', 'middle'), "b2dd5ac30ae4", None), # 37 -> 35 with the same 2026-09-05 narrowing as the 1.4 twin, # whose entry carries the reason. Here the one name that # changed hands, '김민준 박사님', goes to the spaced rule @@ -5861,7 +5865,7 @@ def _claim(rule: dict) -> _Claim: "fix(#585) a decomposed initial keeps its whole first letter": _Claim(2, ('_initials',), "c4f045134f0f", None), "fix(#594) halfwidth katakana takes the kana license and the 间隔号": - _Claim(3, ('family', 'given'), "fdaa80515e52", None), + _Claim(5, ('family', 'given'), "86805db7d55e", None), "fix(#594) a period-marked halfwidth katakana word is not a title": _Claim(2, ('_ambiguities', 'given', 'title'), "8cdafcf56c45", None), }, @@ -6586,7 +6590,7 @@ def _claim(rule: dict) -> _Claim: "fix(#585) a decomposed initial keeps its whole first letter": _Claim(2, ('_initials',), "c4f045134f0f", None), "fix(#594) halfwidth katakana takes the kana license and the 间隔号": - _Claim(3, ('family', 'given'), "fdaa80515e52", None), + _Claim(5, ('family', 'given'), "86805db7d55e", None), "fix(#594) a period-marked halfwidth katakana word is not a title": _Claim(2, ('_ambiguities', 'given', 'title'), "8cdafcf56c45", None), }, @@ -6911,7 +6915,7 @@ def _claim(rule: dict) -> _Claim: "fix(#585) a decomposed initial keeps its whole first letter": _Claim(2, ('_initials',), "c4f045134f0f", None), "fix(#594) halfwidth katakana takes the kana license and the 间隔号": - _Claim(3, ('_ambiguities', 'family', 'given'), "fdaa80515e52", None), + _Claim(5, ('_ambiguities', 'family', 'given'), "86805db7d55e", None), "fix(#594) a period-marked halfwidth katakana word is not a title": _Claim(2, ('_ambiguities', 'given', 'title'), "8cdafcf56c45", None), }, diff --git a/tests/v2/test_locales.py b/tests/v2/test_locales.py index 30a14c69..15d2665c 100644 --- a/tests/v2/test_locales.py +++ b/tests/v2/test_locales.py @@ -292,6 +292,9 @@ def test_ja_adapter_guard_stack_against_a_stub( # contains JA chars but is not WHOLLY JA assert seg("Yamada太郎") is None assert seg("林") is None # namedivider raises below 2 + # halfwidth kana is katakana, so a Han+halfwidth token is wholly + # Japanese and reaches the divider (#594) + assert seg("山田タロウ") is not None def test_ja_adapter_declines_shime_tokens( diff --git a/tools/differential/corpus_cjk.jsonl b/tools/differential/corpus_cjk.jsonl index e4cf05f5..e531a466 100644 --- a/tools/differential/corpus_cjk.jsonl +++ b/tools/differential/corpus_cjk.jsonl @@ -22,9 +22,11 @@ "山田 花子(旧姓 佐藤)" "山田 タロウ" "山田 ダイスケ" +"山田 ペーター" "山田「タロ」太郎" "山田太郎様" "山田花子 旧姓 佐藤" +"山田タロウ" "张伟" "毛 泽东" "毛 김" diff --git a/tools/differential/expected_since_1.4.0.toml b/tools/differential/expected_since_1.4.0.toml index ad2caa1d..93baefc0 100644 --- a/tools/differential/expected_since_1.4.0.toml +++ b/tools/differential/expected_since_1.4.0.toml @@ -185,8 +185,9 @@ issue = "fix(#271/#272/#298) native-script CJK: family-first order, hangul segme # #594. Until then U+FF65 alone stood here, a span _SCRIPT_RANGES did # not have: the halfwidth middle dot separated tokens like its # fullwidth twin, so 'マイケル・ジャクソン' split where 1.4 left one token, -# while the rest of the halfwidth block was left out of both because -# a dotless halfwidth name was parity -- which it still is where it +# while the rest of the halfwidth block was left out of the table as a +# separate normalization problem, and out of this class because a +# dotless halfwidth name was parity -- which it still is where it # is wholly katakana ('ヤマダ タロウ' keeps the declared order, as # 'マイケル ジャクソン' does). What #594 moved is the kana license and # the 间隔号's flank guard reaching halfwidth kana: '山田 タロウ' reads @@ -4939,25 +4940,25 @@ fields = ["_initials"] # because katakana has no initials (rules.md#H2's Accepted clause), # closing the limit decisions.md#cjk-full-stops recorded. A WHOLLY # halfwidth name keeps the declared order, as wholly full-width -# katakana does, so 'ヤマダ タロウ' is parity and in neither rule. -# -# Literal-anchored to the names: no corpus line held a halfwidth kana -# before this change, so the movers ARE the six case rows #594 added, -# less the parity one. A shape spelling -- any halfwidth kana -- would -# be a hand copy of the span the canonical CJK rule already carries. -# LAST in the file: each diff reported unexplained, or unclassified on -# the radar, on the run before these rules existed (2026-10-03). +# katakana does, so 'ヤマダ タロウ' is parity. +# +# At this baseline the five license and 间隔号 movers need no rule of +# their own: the canonical fix(#271/#272/#298) rule above explains +# them through its span class, which copies the halfwidth block since +# #594 (its class reaches the parity name too, as it reaches +# 'マイケル ジャクソン'). The two title movers move `title`, outside +# that rule's fields, so the one rule below takes them. Literal- +# anchored, and LAST in the file: both reported unclassified on the +# radar on the run before it existed (2026-10-03). # --------------------------------------------------------------- -# The license and 间隔号 movers need no rule at this baseline: the -# canonical fix(#271/#272/#298) rule above explains them through -# its span class, which copies the halfwidth block since #594. - [[change]] issue = "fix(#594) a period-marked halfwidth katakana word is not a title" -# 'タナカ. John' and 'マイケル.' read title where their full-width twins read -# a name word; both read given now. Tolerated input (an edge full stop -# on a CJK word), so these are radar names and the rule classifies them -# for the release note rather than for the exit code. +# 'タナカ. John' and 'マイケル.' read title at 1.4.0, as their full-width +# twins did until #cjk-full-stops (2.3); both read given now, as the +# twins do. +# Tolerated input (an edge full stop on a CJK word), so these are +# radar names and the rule classifies them for the release note rather +# than for the exit code. name_regex = "^(?:タナカ\\. John|マイケル\\.)$" fields = ["title", "given"] diff --git a/tools/differential/expected_since_2.0.0.toml b/tools/differential/expected_since_2.0.0.toml index 727ad133..85ce09bc 100644 --- a/tools/differential/expected_since_2.0.0.toml +++ b/tools/differential/expected_since_2.0.0.toml @@ -4000,26 +4000,26 @@ fields = ["_initials"] # because katakana has no initials (rules.md#H2's Accepted clause), # closing the limit decisions.md#cjk-full-stops recorded. A WHOLLY # halfwidth name keeps the declared order, as wholly full-width -# katakana does, so 'ヤマダ タロウ' is parity and in neither rule. -# -# Literal-anchored to the names: no corpus line held a halfwidth kana -# before this change, so the movers ARE the six case rows #594 added, -# less the parity one. A shape spelling -- any halfwidth kana -- would -# be a hand copy of the span the canonical CJK rule already carries. -# LAST in the file: each diff reported unexplained, or unclassified on -# the radar, on the run before these rules existed (2026-10-03). +# katakana does, so 'ヤマダ タロウ' is parity. +# +# At this baseline the five license and 间隔号 movers need no rule of +# their own: the canonical fix(#271/#272/#298) rule above explains +# them through its span class, which copies the halfwidth block since +# #594 (its class reaches the parity name too, as it reaches +# 'マイケル ジャクソン'). The two title movers move `title`, outside +# that rule's fields, so the one rule below takes them. Literal- +# anchored, and LAST in the file: both reported unclassified on the +# radar on the run before it existed (2026-10-03). # --------------------------------------------------------------- -# The license and 间隔号 movers need no rule at this baseline: the -# canonical fix(#271/#272/#298) rule above explains them through -# its span class, which copies the halfwidth block since #594. - [[change]] issue = "fix(#594) a period-marked halfwidth katakana word is not a title" -# 'タナカ. John' and 'マイケル.' read title where their full-width twins read -# a name word; both read given now. Tolerated input (an edge full stop -# on a CJK word), so these are radar names and the rule classifies them -# for the release note rather than for the exit code. +# 'タナカ. John' and 'マイケル.' read title at 2.0.0, as their full-width +# twins did until #cjk-full-stops (2.3); both read given now, as the +# twins do. +# Tolerated input (an edge full stop on a CJK word), so these are +# radar names and the rule classifies them for the release note rather +# than for the exit code. # `_ambiguities` is the v2 surface's: the lone 'マイケル.' reports # given-or-family as 'マイケル.' does, where the title reported nothing. name_regex = "^(?:タナカ\\. John|マイケル\\.)$" diff --git a/tools/differential/expected_since_2.1.0.toml b/tools/differential/expected_since_2.1.0.toml index 046c0878..5933da94 100644 --- a/tools/differential/expected_since_2.1.0.toml +++ b/tools/differential/expected_since_2.1.0.toml @@ -3911,29 +3911,33 @@ fields = ["_initials"] # because katakana has no initials (rules.md#H2's Accepted clause), # closing the limit decisions.md#cjk-full-stops recorded. A WHOLLY # halfwidth name keeps the declared order, as wholly full-width -# katakana does, so 'ヤマダ タロウ' is parity and in neither rule. +# katakana does, so 'ヤマダ タロウ' is parity. # # Literal-anchored to the names: no corpus line held a halfwidth kana -# before this change, so the movers ARE the six case rows #594 added, -# less the parity one. A shape spelling -- any halfwidth kana -- would -# be a hand copy of the span the canonical CJK rule already carries. -# LAST in the file: each diff reported unexplained, or unclassified on -# the radar, on the run before these rules existed (2026-10-03). +# before this change, so the movers ARE the eight case rows #594 added, +# less the parity one -- five here, two in the title rule. A shape +# spelling -- any halfwidth kana -- would be a hand copy of the span +# the canonical CJK rule already carries. LAST in the file: each diff +# reported unexplained, or unclassified on the radar, on the run +# before these rules existed (2026-10-03). # --------------------------------------------------------------- [[change]] issue = "fix(#594) halfwidth katakana takes the kana license and the 间隔号" -# '山田 タロウ' and '山田 ダイスケ' read family 山田 where every release read -# it given; 'タロウ·ヤマダ' divides where it was one given token. -name_regex = "^(?:山田 タロウ|山田 ダイスケ|タロウ\\u00B7ヤマダ)$" +# '山田 タロウ', '山田 ダイスケ' and '山田 ペーター' read family 山田 where every +# release read it given, the unspaced '山田タロウ' reads family where it +# read given, and 'タロウ·ヤマダ' divides where it was one given token. +name_regex = "^(?:山田 タロウ|山田 ダイスケ|山田 ペーター|山田タロウ|タロウ\\u00B7ヤマダ)$" fields = ["given", "family"] [[change]] issue = "fix(#594) a period-marked halfwidth katakana word is not a title" -# 'タナカ. John' and 'マイケル.' read title where their full-width twins read -# a name word; both read given now. Tolerated input (an edge full stop -# on a CJK word), so these are radar names and the rule classifies them -# for the release note rather than for the exit code. +# 'タナカ. John' and 'マイケル.' read title at 2.1.0, as their full-width +# twins did until #cjk-full-stops (2.3); both read given now, as the +# twins do. +# Tolerated input (an edge full stop on a CJK word), so these are +# radar names and the rule classifies them for the release note rather +# than for the exit code. # `_ambiguities` is the v2 surface's: the lone 'マイケル.' reports # given-or-family as 'マイケル.' does, where the title reported nothing. name_regex = "^(?:タナカ\\. John|マイケル\\.)$" diff --git a/tools/differential/expected_since_2.2.0.toml b/tools/differential/expected_since_2.2.0.toml index ba8c7d99..f578f864 100644 --- a/tools/differential/expected_since_2.2.0.toml +++ b/tools/differential/expected_since_2.2.0.toml @@ -2307,29 +2307,33 @@ fields = ["_initials"] # because katakana has no initials (rules.md#H2's Accepted clause), # closing the limit decisions.md#cjk-full-stops recorded. A WHOLLY # halfwidth name keeps the declared order, as wholly full-width -# katakana does, so 'ヤマダ タロウ' is parity and in neither rule. +# katakana does, so 'ヤマダ タロウ' is parity. # # Literal-anchored to the names: no corpus line held a halfwidth kana -# before this change, so the movers ARE the six case rows #594 added, -# less the parity one. A shape spelling -- any halfwidth kana -- would -# be a hand copy of the span the canonical CJK rule already carries. -# LAST in the file: each diff reported unexplained, or unclassified on -# the radar, on the run before these rules existed (2026-10-03). +# before this change, so the movers ARE the eight case rows #594 added, +# less the parity one -- five here, two in the title rule. A shape +# spelling -- any halfwidth kana -- would be a hand copy of the span +# the canonical CJK rule already carries. LAST in the file: each diff +# reported unexplained, or unclassified on the radar, on the run +# before these rules existed (2026-10-03). # --------------------------------------------------------------- [[change]] issue = "fix(#594) halfwidth katakana takes the kana license and the 间隔号" -# '山田 タロウ' and '山田 ダイスケ' read family 山田 where every release read -# it given; 'タロウ·ヤマダ' divides where it was one given token. -name_regex = "^(?:山田 タロウ|山田 ダイスケ|タロウ\\u00B7ヤマダ)$" +# '山田 タロウ', '山田 ダイスケ' and '山田 ペーター' read family 山田 where every +# release read it given, the unspaced '山田タロウ' reads family where it +# read given, and 'タロウ·ヤマダ' divides where it was one given token. +name_regex = "^(?:山田 タロウ|山田 ダイスケ|山田 ペーター|山田タロウ|タロウ\\u00B7ヤマダ)$" fields = ["given", "family"] [[change]] issue = "fix(#594) a period-marked halfwidth katakana word is not a title" -# 'タナカ. John' and 'マイケル.' read title where their full-width twins read -# a name word; both read given now. Tolerated input (an edge full stop -# on a CJK word), so these are radar names and the rule classifies them -# for the release note rather than for the exit code. +# 'タナカ. John' and 'マイケル.' read title at 2.2.0, as their full-width +# twins did until #cjk-full-stops (2.3); both read given now, as the +# twins do. +# Tolerated input (an edge full stop on a CJK word), so these are +# radar names and the rule classifies them for the release note rather +# than for the exit code. # `_ambiguities` is the v2 surface's: the lone 'マイケル.' reports # given-or-family as 'マイケル.' does, where the title reported nothing. name_regex = "^(?:タナカ\\. John|マイケル\\.)$" diff --git a/tools/differential/expected_since_2.3.0.toml b/tools/differential/expected_since_2.3.0.toml index ebc0681a..5a1bdfe9 100644 --- a/tools/differential/expected_since_2.3.0.toml +++ b/tools/differential/expected_since_2.3.0.toml @@ -1562,31 +1562,36 @@ fields = ["_initials"] # because katakana has no initials (rules.md#H2's Accepted clause), # closing the limit decisions.md#cjk-full-stops recorded. A WHOLLY # halfwidth name keeps the declared order, as wholly full-width -# katakana does, so 'ヤマダ タロウ' is parity and in neither rule. +# katakana does, so 'ヤマダ タロウ' is parity. # # Literal-anchored to the names: no corpus line held a halfwidth kana -# before this change, so the movers ARE the six case rows #594 added, -# less the parity one. A shape spelling -- any halfwidth kana -- would -# be a hand copy of the span the canonical CJK rule already carries. -# LAST in the file: each diff reported unexplained, or unclassified on -# the radar, on the run before these rules existed (2026-10-03). +# before this change, so the movers ARE the eight case rows #594 added, +# less the parity one -- five here, two in the title rule. A shape +# spelling -- any halfwidth kana -- would be a hand copy of the span +# the canonical CJK rule already carries. LAST in the file: each diff +# reported unexplained, or unclassified on the radar, on the run +# before these rules existed (2026-10-03). # --------------------------------------------------------------- [[change]] issue = "fix(#594) halfwidth katakana takes the kana license and the 间隔号" -# '山田 タロウ' and '山田 ダイスケ' read family 山田 where every release read -# it given; 'タロウ·ヤマダ' divides where it was one given token. -# `_ambiguities` moves on 'タロウ·ヤマダ' alone: one token reported -# given-or-family (#449, new in 2.3), and two read positionally do not. -name_regex = "^(?:山田 タロウ|山田 ダイスケ|タロウ\\u00B7ヤマダ)$" +# '山田 タロウ', '山田 ダイスケ' and '山田 ペーター' read family 山田 where every +# release read it given, the unspaced '山田タロウ' reads family where it +# read given, and 'タロウ·ヤマダ' divides where it was one given token. +# `_ambiguities` moves on the two single-token readings 2.3 reported +# given-or-family (#449, new in 2.3): 'タロウ·ヤマダ' now reads as two +# tokens, and the license decides '山田タロウ'. +name_regex = "^(?:山田 タロウ|山田 ダイスケ|山田 ペーター|山田タロウ|タロウ\\u00B7ヤマダ)$" fields = ["given", "family", "_ambiguities"] [[change]] issue = "fix(#594) a period-marked halfwidth katakana word is not a title" -# 'タナカ. John' and 'マイケル.' read title where their full-width twins read -# a name word; both read given now. Tolerated input (an edge full stop -# on a CJK word), so these are radar names and the rule classifies them -# for the release note rather than for the exit code. +# 'タナカ. John' and 'マイケル.' read title where their full-width twins +# already read a name word (#cjk-full-stops, new in 2.3); both read +# given now, as the twins do. +# Tolerated input (an edge full stop on a CJK word), so these are +# radar names and the rule classifies them for the release note rather +# than for the exit code. # `_ambiguities` is the v2 surface's: the lone 'マイケル.' reports # given-or-family as 'マイケル.' does, where the title reported nothing. name_regex = "^(?:タナカ\\. John|マイケル\\.)$" From 15e03588ee36b23b25c75c22b8cec36abfb462cd Mon Sep 17 00:00:00 2001 From: Derek Gulbranson Date: Sat, 3 Oct 2026 14:16:14 -0700 Subject: [PATCH 7/9] docs: fix what the PR review found in #594's comments and prose - _policy.py cited T1 for "nothing rewrites the text"; that is rules.md's T Background, now quoted. Same correction in decisions.md#W4. - The voicing-mark sentence states the mechanism: NFC never composes a halfwidth mark into its base. - decisions.md#W4's population is eight rows, five contract movers. - migrate.rst still gave the retracted "usually a transcription" reason. - The release log says "declared order", which is what W4 keeps. Co-Authored-By: Claude Opus 5.5 --- docs/design/decisions.md | 2 +- docs/migrate.rst | 6 ++++-- docs/release_log.rst | 2 +- nameparser/_policy.py | 17 ++++++++++------- 4 files changed, 16 insertions(+), 11 deletions(-) diff --git a/docs/design/decisions.md b/docs/design/decisions.md index 008f0f6f..0976234c 100644 --- a/docs/design/decisions.md +++ b/docs/design/decisions.md @@ -468,7 +468,7 @@ Declined: - 2026-07-27 (script-scoped order amendment) — the family-first override is keyed to the SCRIPT of the written name, never to a guessed language: wholly-Han, wholly-Hangul, and kana-licensed Japanese read family-first because zh, ko and ja all write family-first in native script; wholly-katakana names are predominantly transcriptions and keep the declared order. Latin transliterations are never touched. - 2026-07-29 #272 — the kana license: Han∪kana with at least one kana cannot be Chinese and is not a transcription, so 高橋みなみ reads family-first though it is written in two scripts. -- 2026-10-03 #594 — HALFWIDTH KATAKANA IS KATAKANA. The script table classifies the halfwidth kana block U+FF65–FF9F as `Script.KATAKANA`, so `山田 タロウ` reads family 山田 as `山田 エミ` does; every release through 2.3.0 left the block unclassified and read it given 山田, family タロウ, the positional reading. Legacy data written for JIS X 0201 systems, which had no kanji, spells names this way — bank, payroll and CSV exports. DECIDED by widening the table rather than by folding: the T rules forbid rewriting the text (decisions.md#T1), so normalization at tokenize was never available, and an NFKC fold at classification — the other place the issue offered — reaches far past the kana (`JOHN` to `JOHN`, `㈱` to `(株)`), and it would move a wholly-katakana name across W4's line: a voicing mark with no precomposed form to join (`ア゙`, `ン゙`) folds to the COMBINING U+3099, which sits in the HIRAGANA block, so the folded token reads as kana-licensed Japanese and `ア゙イ タロウ` would read family-first where the unfolded one keeps the declared order (measured 2026-10-03 by parsing the NFKC spelling `ア゙イ タロウ`, which reads family `ア゙イ`; a two-input sweep of halfwidth names against their NFC(NFKC) spellings differed only where such a mark stood; RECOMPUTE: parse each halfwidth name and `unicodedata.normalize("NFC", unicodedata.normalize("NFKC", name))` and compare the fields). The interpunct flank guard is NOT the hazard, though an earlier wording of this bullet said it was: it reads raw text and asks only whether a character is classified at all, which hiragana and katakana both are. The same measurement shows a pre-existing gap the fold would have widened rather than created: full-width katakana typed with an uncomposable combining U+3099 already takes the kana license today. A block range is how the table already decides (U+30FB and the voicing marks are in by block), and it takes the voicing marks U+FF9E/U+FF9F in because they are spacing characters following their base. The span starts at U+FF65, the halfwidth nakaguro, which tokenize already treats as a separator (rules.md#T2), and stops short of U+FF61–FF64, whose U+FF61 is a full stop. Everything keyed on the table follows: the kana license, the second-East-Asian-word count that keeps a segmenter from being asked (`高橋一郎 タロウ`), the 间隔号's flank guard (`タロウ·ヤマダ` divides), and `_NO_INITIALS`, which closes the limit #cjk-full-stops recorded (`タナカ. John` read title `タナカ.`). THE RATIONALE FOR WHOLLY-KATAKANA NAMES IS AMENDED, NOT THE BEHAVIOR: a wholly-halfwidth name keeps the declared order exactly as a wholly-katakana one does (`ヤマダ タロウ` reads given ヤマダ), but the 2026-07-27 reason — that such names are predominantly transcriptions — is weaker here, since a system with no kanji wrote Japanese names in katakana too, family-first. What stands is the narrower claim: the script cannot tell a Japanese reading from a transcription, so it settles no order, and a caller who knows their katakana is Japanese adds `(Script.KATAKANA, FAMILY_FIRST)` to `script_orders` (docs/usage.rst's Boundaries section carries the recipe). `name_order=FAMILY_FIRST` works too, at the cost of reversing Latin names as well. MEASURED 2026-10-03: the differential gate saw nothing before this change because no corpus line held halfwidth kana; the six case rows #594 added are its whole population, three of them contract movers (`山田 タロウ`, `山田 ダイスケ`, `タロウ·ヤマダ`), two tolerated radar movers (`タナカ. John`, `マイケル.`) and one parity row (`ヤマダ タロウ`). RECOMPUTE: parse the six on this tree and on 2.3.0 from PyPI and read `title` beside the name fields. +- 2026-10-03 #594 — HALFWIDTH KATAKANA IS KATAKANA. The script table classifies the halfwidth kana block U+FF65–FF9F as `Script.KATAKANA`, so `山田 タロウ` reads family 山田 as `山田 エミ` does; every release through 2.3.0 left the block unclassified and read it given 山田, family タロウ, the positional reading. Legacy data written for JIS X 0201 systems, which had no kanji, spells names this way — bank, payroll and CSV exports. DECIDED by widening the table rather than by folding: nothing rewrites the text before parsing (rules.md's T Background), so normalization at tokenize was never available, and an NFKC fold at classification — the other place the issue offered — reaches far past the kana (`JOHN` to `JOHN`, `㈱` to `(株)`), and it would move a wholly-katakana name across W4's line: a voicing mark with no precomposed form to join (`ア゙`, `ン゙`) folds to the COMBINING U+3099, which sits in the HIRAGANA block, so the folded token reads as kana-licensed Japanese and `ア゙イ タロウ` would read family-first where the unfolded one keeps the declared order (measured 2026-10-03 by parsing the NFKC spelling `ア゙イ タロウ`, which reads family `ア゙イ`; a two-input sweep of halfwidth names against their NFC(NFKC) spellings differed only where such a mark stood; RECOMPUTE: parse each halfwidth name and `unicodedata.normalize("NFC", unicodedata.normalize("NFKC", name))` and compare the fields). The interpunct flank guard is NOT the hazard, though an earlier wording of this bullet said it was: it reads raw text and asks only whether a character is classified at all, which hiragana and katakana both are. The same measurement shows a pre-existing gap the fold would have widened rather than created: full-width katakana typed with an uncomposable combining U+3099 already takes the kana license today. A block range is how the table already decides (U+30FB and the voicing marks are in by block), and it takes the voicing marks U+FF9E/U+FF9F in because they are spacing characters following their base. The span starts at U+FF65, the halfwidth nakaguro, which tokenize already treats as a separator (rules.md#T2), and stops short of U+FF61–FF64, whose U+FF61 is a full stop. Everything keyed on the table follows: the kana license, the second-East-Asian-word count that keeps a segmenter from being asked (`高橋一郎 タロウ`), the 间隔号's flank guard (`タロウ·ヤマダ` divides), and `_NO_INITIALS`, which closes the limit #cjk-full-stops recorded (`タナカ. John` read title `タナカ.`). THE RATIONALE FOR WHOLLY-KATAKANA NAMES IS AMENDED, NOT THE BEHAVIOR: a wholly-halfwidth name keeps the declared order exactly as a wholly-katakana one does (`ヤマダ タロウ` reads given ヤマダ), but the 2026-07-27 reason — that such names are predominantly transcriptions — is weaker here, since a system with no kanji wrote Japanese names in katakana too, family-first. What stands is the narrower claim: the script cannot tell a Japanese reading from a transcription, so it settles no order, and a caller who knows their katakana is Japanese adds `(Script.KATAKANA, FAMILY_FIRST)` to `script_orders` (docs/usage.rst's Boundaries section carries the recipe). `name_order=FAMILY_FIRST` works too, at the cost of reversing Latin names as well. MEASURED 2026-10-03: the differential gate saw nothing before this change because no corpus line held halfwidth kana; the eight case rows #594 added are its whole population, five of them contract movers (`山田 タロウ`, `山田 ダイスケ`, `山田 ペーター`, the unspaced `山田タロウ`, `タロウ·ヤマダ`), two tolerated radar movers (`タナカ. John`, `マイケル.`) and one parity row (`ヤマダ タロウ`). RECOMPUTE: parse the eight on this tree and on 2.3.0 from PyPI and read `title` beside the name fields. ### D1 — the segmenterless-activation warning diff --git a/docs/migrate.rst b/docs/migrate.rst index 6ff8adc4..f0bb02b4 100644 --- a/docs/migrate.rst +++ b/docs/migrate.rst @@ -600,8 +600,10 @@ token moves from ``first`` to ``last`` exactly as ``毛泽东`` does: ``first``. The spaced shapes change what the name renders as too: ``str(HumanName("高橋 みなみ"))`` was ``"高橋 みなみ"`` and is now ``"みなみ 高橋"``. A name written *wholly* in katakana is deliberately -left alone — it is usually a transcribed foreign name already in -given-first order — so ``HumanName("マイケル ジャクソン")`` reads +left alone — it may be a transcribed foreign name already in +given-first order or a Japanese name's reading written family-first, +and the script cannot tell which (see :ref:`east-asian-names` to opt +in) — so ``HumanName("マイケル ジャクソン")`` reads ``first="マイケル"``/``last="ジャクソン"`` on both versions. One more shape changes for a different reason: the katakana middle dot diff --git a/docs/release_log.rst b/docs/release_log.rst index 0dae50de..c488f3aa 100644 --- a/docs/release_log.rst +++ b/docs/release_log.rst @@ -66,7 +66,7 @@ Release Log - **Fix a v1 ``Constants`` entry that is only a CJK full stop raising at the first parse.** After ``c.titles.add("。")``, ``HumanName("john smith", c)`` raised ``ValueError`` in 2.3; ``。``, ``.`` and ``。`` in any set or as a ``capitalization_exceptions`` key are now ignored with the same ``UserWarning`` as an entry with stray whitespace. 1.4.0 through 2.2.0 accepted such an entry and applied it to a name token that is nothing but that full stop, which 2.3 and later cannot match. An entry ending in such a full stop no longer raises either: ``c.suffix_not_acronyms.add("ma。")`` made ``HumanName("jack ma", c)`` raise ``ValueError`` in 2.3, and now gives last ``ma``, as without the entry. (closes #582) - - **Fix a Japanese name written in halfwidth katakana being read given-first.** ``HumanName("山田 タロウ")`` gives last ``山田``, first ``タロウ``, as ``山田 タロウ`` does, where every release gave first ``山田``, last ``タロウ``. Halfwidth katakana (U+FF65–U+FF9F), which legacy bank, payroll and CSV exports still carry, is now read as katakana everywhere the script matters: a second halfwidth word keeps the segmenter from re-dividing a kanji name (``高橋一郎 タロウ``), the Chinese ``·`` divides between halfwidth kana (``タロウ·ヤマダ`` gives first ``タロウ``, last ``ヤマダ`` where it was one first name), and a period-marked halfwidth word is no longer taken for an abbreviated title: ``タナカ. John`` gives first ``タナカ.`` where it gave title ``タナカ.``. A name written wholly in katakana, halfwidth or not, still keeps the order it was written in (``ヤマダ タロウ`` gives first ``ヤマダ``), because the script cannot say whether it is a Japanese name or a transcribed foreign one; to read your katakana names family-first, add ``(Script.KATAKANA, FAMILY_FIRST)`` to ``Policy.script_orders`` (see :ref:`east-asian-names`). See the ``W4`` entry of ``docs/design/decisions.md`` (closes #594) + - **Fix a Japanese name written in halfwidth katakana being read given-first.** ``HumanName("山田 タロウ")`` gives last ``山田``, first ``タロウ``, as ``山田 タロウ`` does, where every release gave first ``山田``, last ``タロウ``. Halfwidth katakana (U+FF65–U+FF9F), which legacy bank, payroll and CSV exports still carry, is now read as katakana everywhere the script matters: a second halfwidth word keeps the segmenter from re-dividing a kanji name (``高橋一郎 タロウ``), the Chinese ``·`` divides between halfwidth kana (``タロウ·ヤマダ`` gives first ``タロウ``, last ``ヤマダ`` where it was one first name), and a period-marked halfwidth word is no longer taken for an abbreviated title: ``タナカ. John`` gives first ``タナカ.`` where it gave title ``タナカ.``. A name written wholly in katakana, halfwidth or not, still keeps the declared order, given-first by default (``ヤマダ タロウ`` gives first ``ヤマダ``), because the script cannot say whether it is a Japanese name or a transcribed foreign one; to read your katakana names family-first, add ``(Script.KATAKANA, FAMILY_FIRST)`` to ``Policy.script_orders`` (see :ref:`east-asian-names`). See the ``W4`` entry of ``docs/design/decisions.md`` (closes #594) **Additions** diff --git a/nameparser/_policy.py b/nameparser/_policy.py index b9052166..2b685b2d 100644 --- a/nameparser/_policy.py +++ b/nameparser/_policy.py @@ -132,17 +132,20 @@ class Script(StrEnum): # personal name uses them. Halfwidth kana (U+FF65-U+FF9F, #594) IS # katakana here: legacy bank, payroll and CSV data written for JIS X # 0201 systems spells names in it, and 山田 タロウ must read as 山田 タロウ -# does. Classifying it is a range, not a fold -- T1 forbids rewriting -# the text, and an NFKC fold at classification would reach far past -# the kana (fullwidth Latin, ㈱), and a voicing mark with no +# does. Classifying it is a range, not a fold. rules.md's T Background +# says "nothing rewrites the text before parsing", which rules out a +# fold at tokenize, and an NFKC fold at classification would reach far +# past the kana (fullwidth Latin, ㈱), and a voicing mark with no # precomposed form to join (ア゙) folds to the combining U+3099, which # is in the HIRAGANA block -- so the folded token would take the kana # license and a wholly-katakana name would turn family-first # (decisions.md#W4). The span takes in the # halfwidth nakaguro U+FF65, which tokenize turns into a separator, # on the same direct-call grounds as U+30FB below, and the voicing -# marks U+FF9E/U+FF9F, which are spacing characters following their -# base and so need the block to be classified at all. It stops short +# marks U+FF9E/U+FF9F: they are spacing characters (category Lm) that +# NFC never composes into their base, there being no precomposed +# halfwidth voiced kana, so a span ending at U+FF9D would leave ダ +# unclassified. It stops short # of U+FF61-U+FF64, the halfwidth CJK punctuation, whose U+FF61 is a # full stop (_lexicon.FULL_STOPS). # This table classifies by Unicode BLOCK, not the UAX #24 Script @@ -319,8 +322,8 @@ def _order_repr(value: tuple[Role, ...]) -> str: #: classified fix, #294's mechanism. KATAKANA is deliberately absent: #: a PURE-katakana token may be a transcribed foreign name kept in its #: source (usually given-first) order, or a Japanese reading written -#: family-first -- legacy halfwidth data (#594) writes every name in -#: katakana -- and the script cannot tell them apart, so nothing should +#: family-first -- legacy halfwidth data (#594) spells Japanese names in +#: katakana too -- and the script cannot tell them apart, so nothing should #: default on it. Canonical form: sorted (Script, order) pairs, #: matching the field's storage. DEFAULT_SCRIPT_ORDERS: tuple[ From f4dda946e67f0bbf4c5644c08fd6a5ec3619abf9 Mon Sep 17 00:00:00 2001 From: Derek Gulbranson Date: Sat, 3 Oct 2026 14:18:33 -0700 Subject: [PATCH 8/9] =?UTF-8?q?test:=20say=20which=20of=20the=20span's=20c?= =?UTF-8?q?odepoints=20the=20=EF=BE=8D=EF=BE=9F=EF=BD=B0=EF=BE=80=EF=BD=B0?= =?UTF-8?q?=20row=20pins=20(#594)?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit U+FF70 is interior to the span, not an edge. Co-Authored-By: Claude Opus 5.5 --- tests/v2/cases.py | 10 +++++----- 1 file changed, 5 insertions(+), 5 deletions(-) diff --git a/tests/v2/cases.py b/tests/v2/cases.py index 28b26238..48593c4b 100644 --- a/tests/v2/cases.py +++ b/tests/v2/cases.py @@ -9145,11 +9145,11 @@ def _check_cjk_shape_purity(self) -> None: "山田 ペーター", {"family": "山田", "given": "ペーター"}, classification="fix(#594)", - notes="the span's other two edges a name actually reaches: " - "the semi-voiced mark ゚ (U+FF9F, the block's last " - "codepoint) and the prolonged sound mark ー (U+FF70). " - "Narrowing the range by either one leaves ペーター " - "mixed-script and the name positional"), + notes="the block's last codepoint, the semi-voiced mark ゚ " + "(U+FF9F), and the prolonged sound mark ー (U+FF70, which " + "a range starting at the ordinary letters, U+FF71, " + "would drop). Narrowing the range past either leaves " + "ペーター mixed-script and the name positional"), Case("ja_halfwidth_unspaced_unsegmented_default", "山田タロウ", {"family": "山田タロウ"}, classification="fix(#594)", From 357f685bbf3668553218b0c5d155ba02eebc9b23 Mon Sep 17 00:00:00 2001 From: Derek Gulbranson Date: Sat, 3 Oct 2026 14:29:24 -0700 Subject: [PATCH 9/9] docs: AGENTS.md no longer says halfwidth data writes every name in katakana (#594) The same over-broad claim the review corrected in _policy.py. Co-Authored-By: Claude Opus 5.5 --- AGENTS.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/AGENTS.md b/AGENTS.md index 8ab7d98a..8bd30e4e 100644 --- a/AGENTS.md +++ b/AGENTS.md @@ -263,7 +263,7 @@ logging.getLogger('HumanName').setLevel(logging.DEBUG) The library has two layers: `nameparser/config/` (data) and `nameparser/parser.py` (logic). -**Design philosophy — positional and language-agnostic.** The parser assigns parts by *position* plus small sets of words that join to neighbors; it never detects language. A name's language can't be reliably inferred from Latin-script transliteration ("Ali" is Arabic or Italian; "Van"/"Della"/"Bin" are first names in some cultures, particles in others), so language-specific rules belong in opt-in `Constants` config, never global defaults. Many "wrong for language X" reports (#133, #150, #130, #85, #103, #146, #83) are irreducible ambiguities — e.g. `de Mesnil` (want last name) vs `Van Johnson` (want first name) are the same `[prefix][word]` shape. Before adding a rule, confirm it doesn't break the opposite case (run the full suite — Portuguese and "Van Johnson" tests are the usual canaries). **The one scoped exception (2.1, #271/#272): script-conditional behavior is permitted exactly where the SCRIPT ITSELF — not statistics about it — determines the convention.** The never-detect-language rule above is about Latin *transliteration*, where the signal genuinely is destroyed; native script is a different question, and it is answered per behavior rather than per script. Five defaults fall out of it, plus a sixth that applies the same not-a-guess standard to specific WORDS rather than to a script (#308's honorific peel, below). Wholly-Han, wholly-Hangul and kana-licensed names read family-first (`Policy.script_orders`) — no language detection needed, because zh and ja both write family-first in native script, so order cannot be misread even though the language is unknowable. Unspaced hangul splits into surname + given name (`config/surnames.py` ships the Korean census list as DEFAULT vocabulary) — nothing but Korean is written in hangul and the surnames are a closed census set, and the vocabulary is self-selecting besides: a hangul entry can only ever match hangul text. Hiragana licenses Japanese (#272) — a name whose characters stay inside Han∪kana while carrying at least one kana cannot be Chinese (the kana rules it out) and is not a transcription (foreign names are transcribed in katakana ALONE, マイケル has no kanji), so 高橋みなみ and 山田 エミ read family-first too; mechanically they resolve to the HIRAGANA entry, the license's carrier key. PURE katakana — full-width or halfwidth, the halfwidth block U+FF65–FF9F being katakana in the script table since #594 — is excluded and keeps the positional default: マイケル・ジャクソン is a transcribed foreign name in its source order, and a Japanese reading written in katakana (legacy halfwidth data writes every name that way) cannot be told from one, so the script settles nothing and a caller who knows opts in through `script_orders`. And the 间隔号 U+00B7 (#298) is the transcription marker for scripts that HAVE no transcription script: a name it divides (威廉·莎士比亚 — flanked by classified characters on both sides, so Catalan's Gal·la is untouched) keeps its source order and never segments — the orthography names the convention, with the divider carrying the signal; pure katakana reaches the same reading by the opposite route, its script settling nothing so the declared order stands. And a listed CJK honorific glued to the END of a name token is split off it (#308) — 田中さん is 田中 plus さん — on the same orthography-settles-it test, narrowed for the glued position: an entry peels only where it can never end a name, so 씨/님/さん/様/先生 peel while 양/군/氏/博士/殿 stay spaced-only (김지양 and 田中博士 are names, and ~90 Japanese surnames end in 殿) and 君 is in NEITHER set (王君 is a complete Chinese name), though its kana spelling くん peels. Like the nakaguro's tokenize-level separation described next, it is reached by neither policy opt-out — but for its own reason: the vocabulary carries the license itself rather than borrowing the script's, so `segment_scripts` has nothing to say about it. Since #312 it also crosses the 间隔号, which still stops the surname split standing right beside it: it answers where a name DIVIDES into surname and given, and the peel never asks that question. Whether it also crosses the FAMILY comma is tolerated rather than settled: the 2026-09-01 demotion (rules.md#W3) narrowed that half from contract to best-effort, since no CJK writing system's own convention puts a comma between family and given at all — so `김, 민준씨` reads today exactly as the spaced `김 민준씨` does (family 김, given 민준, suffix 씨) while the split stands down as before, but that reading is watched on the differential's radar tier rather than pinned as contract. Its site is accordingly the name-bearing segment runs — `segments[:2]` under a family comma, and `segments[0]` as before otherwise, the family comma being the one structure that splits the name itself across two runs, with the honorific as often glued to the given side as to the family. That is the whole reach and nothing past it (`김, 민준 지훈씨` peels; `김, 민준, 지훈씨` and `김,, 민준씨` do not, both landing in a third run), and whether `segments[1]` is name text at all is now ASKED rather than inferred from the structure — `segment` does not guarantee it, since a one-word part before the comma reads as FAMILY_COMMA even when the part after it is entirely suffix-shaped, and the peel walking into such a run took `V.` for its site, found no listed tail and abandoned (#319). The question is `segment`'s own suffix-comma predicate, lifted into `_vocab.is_wholly_suffix` so the two stages cannot drift: a wholly suffix-shaped second run is declined and the scan stays in `segments[0]`, so `田中さん, V.` and `田中さん, Ph. D.` give さん up as `田中さん, PhD` always did. The test is necessary but NOT sufficient, and the second condition is not decoration: every honorific tail is also a suffix word, so a glued honorific is itself part of what makes its run read as suffix-shaped, and declining a run that holds the ONLY site loses the peel outright. `segments[0]` must therefore offer a peel site of its own before the second run is declined — `이, J.씨` and `선생님, J.씨` pass the suffix test and are scanned anyway, keeping the pre-#319 reading, while `김민준씨, J.씨` has a site on both sides and peels the person's own 씨 rather than the junk one behind the comma. Uniform in the PEEL, that is — where the credential itself lands is `assign`'s question and still differs by spelling (`V.` → `given`, `PhD` and `Ph. D.` → `suffix`). Not `_is_post_nominal` pluralized: the run predicate says yes both to what the token predicate vetoes (`V.`, `V`, `I` — the class the defect was reported as) and to what the token predicate never sees at all, since `period_joined_vocab` and the delimiter routes are the run predicate's alone (`Msc.Ed.` and `J.씨` reach it that way, and `田中さん, Msc.Ed.` moves with the rest). `Policy(lenient_comma_suffixes=False)` drops this call to the strict token test too — so those three read as name text again and keep the pre-#319 answer, while `Ph. D.` peels under the knob regardless, its merged `phd` passing the strict test. `田中さん, 太郎` is unchanged, and not because of its comma — the honorific there is not at the end of the name, 太郎 is. The nakaguro belongs to the same doctrine but is decided a layer down: U+30FB and its halfwidth twin U+FF65 separate tokens like whitespace, unconditionally and in tokenize, so neither policy opt-out (`script_orders={}`, `segment_scripts=()`) reaches it — the codepoints are CJK-only and appear in no other script's names, which is what licenses a tokenize-level rule where U+00B7 (also the Catalan punt volat, interior to Gal·la) needs the flanked-by-classified-script guard `_tokenize_region` gives it (#298). Han segmentation stays OPT-IN (`locales.ZH` for Chinese, `locales.JA` for Japanese) — a zh surname list corrupts Japanese kanji names, since 高 is a common Chinese surname and 高橋一郎 would split 高+橋一郎 where the correct reading is 高橋+一郎; no surname list divides a kanji name at all, so `locales.JA` activates the stage and a pluggable `Parser(segmenter=...)` does the dividing. Latin-script input is never touched by any of this: "Kim Min-jun" is genuinely order-ambiguous and stays governed by `name_order` and opt-in packs. Before adding a script-conditional rule, work out which of the three it is — certain, certain for this one behavior only, or a statistical guess wearing a script's clothes. +**Design philosophy — positional and language-agnostic.** The parser assigns parts by *position* plus small sets of words that join to neighbors; it never detects language. A name's language can't be reliably inferred from Latin-script transliteration ("Ali" is Arabic or Italian; "Van"/"Della"/"Bin" are first names in some cultures, particles in others), so language-specific rules belong in opt-in `Constants` config, never global defaults. Many "wrong for language X" reports (#133, #150, #130, #85, #103, #146, #83) are irreducible ambiguities — e.g. `de Mesnil` (want last name) vs `Van Johnson` (want first name) are the same `[prefix][word]` shape. Before adding a rule, confirm it doesn't break the opposite case (run the full suite — Portuguese and "Van Johnson" tests are the usual canaries). **The one scoped exception (2.1, #271/#272): script-conditional behavior is permitted exactly where the SCRIPT ITSELF — not statistics about it — determines the convention.** The never-detect-language rule above is about Latin *transliteration*, where the signal genuinely is destroyed; native script is a different question, and it is answered per behavior rather than per script. Five defaults fall out of it, plus a sixth that applies the same not-a-guess standard to specific WORDS rather than to a script (#308's honorific peel, below). Wholly-Han, wholly-Hangul and kana-licensed names read family-first (`Policy.script_orders`) — no language detection needed, because zh and ja both write family-first in native script, so order cannot be misread even though the language is unknowable. Unspaced hangul splits into surname + given name (`config/surnames.py` ships the Korean census list as DEFAULT vocabulary) — nothing but Korean is written in hangul and the surnames are a closed census set, and the vocabulary is self-selecting besides: a hangul entry can only ever match hangul text. Hiragana licenses Japanese (#272) — a name whose characters stay inside Han∪kana while carrying at least one kana cannot be Chinese (the kana rules it out) and is not a transcription (foreign names are transcribed in katakana ALONE, マイケル has no kanji), so 高橋みなみ and 山田 エミ read family-first too; mechanically they resolve to the HIRAGANA entry, the license's carrier key. PURE katakana — full-width or halfwidth, the halfwidth block U+FF65–FF9F being katakana in the script table since #594 — is excluded and keeps the positional default: マイケル・ジャクソン is a transcribed foreign name in its source order, and a Japanese reading written in katakana (legacy halfwidth data spells Japanese names that way too) cannot be told from one, so the script settles nothing and a caller who knows opts in through `script_orders`. And the 间隔号 U+00B7 (#298) is the transcription marker for scripts that HAVE no transcription script: a name it divides (威廉·莎士比亚 — flanked by classified characters on both sides, so Catalan's Gal·la is untouched) keeps its source order and never segments — the orthography names the convention, with the divider carrying the signal; pure katakana reaches the same reading by the opposite route, its script settling nothing so the declared order stands. And a listed CJK honorific glued to the END of a name token is split off it (#308) — 田中さん is 田中 plus さん — on the same orthography-settles-it test, narrowed for the glued position: an entry peels only where it can never end a name, so 씨/님/さん/様/先生 peel while 양/군/氏/博士/殿 stay spaced-only (김지양 and 田中博士 are names, and ~90 Japanese surnames end in 殿) and 君 is in NEITHER set (王君 is a complete Chinese name), though its kana spelling くん peels. Like the nakaguro's tokenize-level separation described next, it is reached by neither policy opt-out — but for its own reason: the vocabulary carries the license itself rather than borrowing the script's, so `segment_scripts` has nothing to say about it. Since #312 it also crosses the 间隔号, which still stops the surname split standing right beside it: it answers where a name DIVIDES into surname and given, and the peel never asks that question. Whether it also crosses the FAMILY comma is tolerated rather than settled: the 2026-09-01 demotion (rules.md#W3) narrowed that half from contract to best-effort, since no CJK writing system's own convention puts a comma between family and given at all — so `김, 민준씨` reads today exactly as the spaced `김 민준씨` does (family 김, given 민준, suffix 씨) while the split stands down as before, but that reading is watched on the differential's radar tier rather than pinned as contract. Its site is accordingly the name-bearing segment runs — `segments[:2]` under a family comma, and `segments[0]` as before otherwise, the family comma being the one structure that splits the name itself across two runs, with the honorific as often glued to the given side as to the family. That is the whole reach and nothing past it (`김, 민준 지훈씨` peels; `김, 민준, 지훈씨` and `김,, 민준씨` do not, both landing in a third run), and whether `segments[1]` is name text at all is now ASKED rather than inferred from the structure — `segment` does not guarantee it, since a one-word part before the comma reads as FAMILY_COMMA even when the part after it is entirely suffix-shaped, and the peel walking into such a run took `V.` for its site, found no listed tail and abandoned (#319). The question is `segment`'s own suffix-comma predicate, lifted into `_vocab.is_wholly_suffix` so the two stages cannot drift: a wholly suffix-shaped second run is declined and the scan stays in `segments[0]`, so `田中さん, V.` and `田中さん, Ph. D.` give さん up as `田中さん, PhD` always did. The test is necessary but NOT sufficient, and the second condition is not decoration: every honorific tail is also a suffix word, so a glued honorific is itself part of what makes its run read as suffix-shaped, and declining a run that holds the ONLY site loses the peel outright. `segments[0]` must therefore offer a peel site of its own before the second run is declined — `이, J.씨` and `선생님, J.씨` pass the suffix test and are scanned anyway, keeping the pre-#319 reading, while `김민준씨, J.씨` has a site on both sides and peels the person's own 씨 rather than the junk one behind the comma. Uniform in the PEEL, that is — where the credential itself lands is `assign`'s question and still differs by spelling (`V.` → `given`, `PhD` and `Ph. D.` → `suffix`). Not `_is_post_nominal` pluralized: the run predicate says yes both to what the token predicate vetoes (`V.`, `V`, `I` — the class the defect was reported as) and to what the token predicate never sees at all, since `period_joined_vocab` and the delimiter routes are the run predicate's alone (`Msc.Ed.` and `J.씨` reach it that way, and `田中さん, Msc.Ed.` moves with the rest). `Policy(lenient_comma_suffixes=False)` drops this call to the strict token test too — so those three read as name text again and keep the pre-#319 answer, while `Ph. D.` peels under the knob regardless, its merged `phd` passing the strict test. `田中さん, 太郎` is unchanged, and not because of its comma — the honorific there is not at the end of the name, 太郎 is. The nakaguro belongs to the same doctrine but is decided a layer down: U+30FB and its halfwidth twin U+FF65 separate tokens like whitespace, unconditionally and in tokenize, so neither policy opt-out (`script_orders={}`, `segment_scripts=()`) reaches it — the codepoints are CJK-only and appear in no other script's names, which is what licenses a tokenize-level rule where U+00B7 (also the Catalan punt volat, interior to Gal·la) needs the flanked-by-classified-script guard `_tokenize_region` gives it (#298). Han segmentation stays OPT-IN (`locales.ZH` for Chinese, `locales.JA` for Japanese) — a zh surname list corrupts Japanese kanji names, since 高 is a common Chinese surname and 高橋一郎 would split 高+橋一郎 where the correct reading is 高橋+一郎; no surname list divides a kanji name at all, so `locales.JA` activates the stage and a pluggable `Parser(segmenter=...)` does the dividing. Latin-script input is never touched by any of this: "Kim Min-jun" is genuinely order-ambiguous and stays governed by `name_order` and opt-in packs. Before adding a script-conditional rule, work out which of the three it is — certain, certain for this one behavior only, or a statistical guess wearing a script's clothes. **A constant's membership is a question you may reopen.** Proposing that a word be ADDED, REMOVED or MOVED between vocabulary sets is ordinary design work — a shipped entry is not evidence that anyone judged it. `SUFFIX_ACRONYMS` arrived in `af5bdab` as a bulk Wikipedia import never reviewed against surname collisions: 572 of its 577 alphabetic entries leave `family` empty in `"John "` against five ambiguous-gated exceptions (recomputed 2026-09-07; the fifth is `ba`; 568 of 575 against seven, measured 2026-09-25 after #540), and `sa`, `se` and `om` are borne as surnames (measured 2026-08-23), as were `rai`, `cha`, `ba` and `mc` before `decisions.md#suffix-acronym-collisions` decided all four — `rai` and `cha` removed, `ba` marked ambiguous, `mc` left alone as no borne name at all. The same entry marked `meng` and `lac` ambiguous on 2026-09-25 (#540). When a fix starts to look like new machinery, check the vocabulary first. Criterion: `decisions.md#vocabulary-collisions`, with #360's positional qualifier.