Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 1 addition & 1 deletion AGENTS.md

Large diffs are not rendered by default.

5 changes: 3 additions & 2 deletions docs/design/decisions.md

Large diffs are not rendered by default.

16 changes: 9 additions & 7 deletions docs/design/rules.md
Original file line number Diff line number Diff line change
Expand Up @@ -116,9 +116,8 @@ H2. Rationale: before a name, an abbreviation is almost always a
Accepted: the shape reads a Latin convention, and a script with
no initials has no period abbreviations either, so a period-
marked opening word carrying a Han, kana or hangul character —
as the script table classifies them; halfwidth katakana sits
outside it and still reads by the Latin shape, the limit
decisions.md#cjk-full-stops records — is a name word and never a
as the script table classifies them, halfwidth katakana
included since #594 — is a name word and never a
title by shape (#323; decisions.md#cjk-full-stops) — the same
veto that keeps 씨. from reading as an initial. The Latin
reading is unchanged, and W3's example block carries the CJK
Expand Down Expand Up @@ -2188,7 +2187,7 @@ O5. Rationale: O4 reads a name by comparing where its words stand,

## Scripts & writing systems (W)

Background: script-conditional behavior is permitted exactly where the writing system itself — not statistics about it — settles the convention; a language can never be inferred from Latin-script text, because transliteration destroys the signal. The facts this section builds on: Chinese and Japanese both write the family name first in native script, so the script settles the order without knowing the language. Hangul is written by exactly one language and Korean family names are a small closed census set. Han text does not identify its language — a Chinese surname list would divide Japanese 高橋一郎 as 高 + 橋一郎 — which is why Han division is opt-in and there is no Korean pack to opt into. Hiragana never transcribes a foreign name (transcriptions are katakana alone), so kanji-plus-kana is a Japanese name in Japanese order, while wholly-katakana is predominantly a transcribed foreign name already in given-first order. Real Chinese text is unspaced (毛泽东); the spaced 毛 泽东 is an artifact. A fuller narrative lives in docs/usage.rst's East Asian section. One fact carries its own consequence: none of the three writing systems marks the family name with a comma — position in the written form is what identifies it, so a comma standing between the family name and the given name is a listing convention carried in from elsewhere rather than a form the script produces. That is why the rule reading one (W3) is tolerated rather than normative. CLDR's own locale data says the same where a contrary convention would have had to appear: across its ko, zh and ja personName patterns not one of the 126 pattern strings carries a comma of any width, the surname-first referring patterns separating surname from given by a single space, and the only comma in reach belongs to the locale-neutral root's sorting format — a list-ordering format, which ko and zh override comma-free and ja all but one inherited slot (decisions.md#cjk-comma-demotion carries the pull verbatim, with its URLs, its commit and its date). A full stop of any width — the ASCII period, the fullwidth ., the ideographic 。 and its halfwidth 。 — glued after a script-written word is punctuation and not part of the word: it is invisible to the script reading and to the vocabulary, and it stays in the text on the word it arrived with, because no East Asian script writes an initial or an abbreviation with a period (#322, #323; decisions.md#cjk-full-stops). A stop glued BEFORE the word is punctuation to the vocabulary lookup, which folds both edges away, so .씨 is still the honorific. The classification fold that feeds the two division sites reads the trailing edge only, so a word wearing a leading stop is given no script at all and never becomes a surname site: .김민준 stays one whole word, given, no script rule reaching it. The honorific peel (W2) is not gated by that fold — the tail alone is its license — so it reads such a token regardless of a leading stop, and its own trailing-edge fold is what decides there: .김민준씨 peels to .김민준 and 씨, and .김민준씨. peels to .김민준 and 씨. (tests/v2/pipeline/test_script_segment.py's test_the_peel_reads_the_trailing_stop_only).
Background: script-conditional behavior is permitted exactly where the writing system itself — not statistics about it — settles the convention; a language can never be inferred from Latin-script text, because transliteration destroys the signal. The facts this section builds on: Chinese and Japanese both write the family name first in native script, so the script settles the order without knowing the language. Hangul is written by exactly one language and Korean family names are a small closed census set. Han text does not identify its language — a Chinese surname list would divide Japanese 高橋一郎 as 高 + 橋一郎 — which is why Han division is opt-in and there is no Korean pack to opt into. Hiragana never transcribes a foreign name (transcriptions are katakana alone), so kanji-plus-kana is a Japanese name in Japanese order, while wholly-katakana may be a transcribed foreign name already in given-first order or a Japanese name's reading written family-first, and nothing in the script says which. Katakana has a halfwidth form (タロウ for タロウ), the only kana systems without kanji could store, and legacy data still carries it: it is katakana, read the same way. Real Chinese text is unspaced (毛泽东); the spaced 毛 泽东 is an artifact. A fuller narrative lives in docs/usage.rst's East Asian section. One fact carries its own consequence: none of the three writing systems marks the family name with a comma — position in the written form is what identifies it, so a comma standing between the family name and the given name is a listing convention carried in from elsewhere rather than a form the script produces. That is why the rule reading one (W3) is tolerated rather than normative. CLDR's own locale data says the same where a contrary convention would have had to appear: across its ko, zh and ja personName patterns not one of the 126 pattern strings carries a comma of any width, the surname-first referring patterns separating surname from given by a single space, and the only comma in reach belongs to the locale-neutral root's sorting format — a list-ordering format, which ko and zh override comma-free and ja all but one inherited slot (decisions.md#cjk-comma-demotion carries the pull verbatim, with its URLs, its commit and its date). A full stop of any width — the ASCII period, the fullwidth ., the ideographic 。 and its halfwidth 。 — glued after a script-written word is punctuation and not part of the word: it is invisible to the script reading and to the vocabulary, and it stays in the text on the word it arrived with, because no East Asian script writes an initial or an abbreviation with a period (#322, #323; decisions.md#cjk-full-stops). A stop glued BEFORE the word is punctuation to the vocabulary lookup, which folds both edges away, so .씨 is still the honorific. The classification fold that feeds the two division sites reads the trailing edge only, so a word wearing a leading stop is given no script at all and never becomes a surname site: .김민준 stays one whole word, given, no script rule reaching it. The honorific peel (W2) is not gated by that fold — the tail alone is its license — so it reads such a token regardless of a leading stop, and its own trailing-edge fold is what decides there: .김민준씨 peels to .김민준 and 씨, and .김민준씨. peels to .김민준 and 씨. (tests/v2/pipeline/test_script_segment.py's test_the_peel_reads_the_trailing_stop_only).

W1. Rationale: hangul is monoglot Korean and its surnames are a
closed census set, so an unspaced hangul name divides at a
Expand Down Expand Up @@ -2298,17 +2297,20 @@ W3. Rationale: a family name declared by a comma is the writer's

W4. Rationale: Chinese, Japanese and Korean all write the family
name first in native script — the script settles the order
without knowing the language — while a wholly-katakana name is
predominantly a transcribed foreign name already in its source
order.
without knowing the language — while a wholly-katakana name may
be a transcribed foreign name already in its source order or a
Japanese reading written family-first, so its script settles
nothing.
A name written wholly in one East Asian script, or in the
kana-licensed Japanese repertoire, reads family-first whatever
order the caller declared; a wholly-katakana name keeps the
declared order.
"김 민준" → family="김"
"山田 太郎" → family="山田"
"高橋 みなみ" → family="高橋"
"山田 タロウ" → family="山田"
"マイケル ジャクソン" → given="マイケル" · boundary
"ヤマダ タロウ" → given="ヤマダ" · boundary
Accepted: a name the interpunct divides keeps its source order —
the divider itself marks a transcription (T3) — so the override
stands down there; the katakana middle dot (T2) carries no such
Expand Down
5 changes: 3 additions & 2 deletions docs/locales.rst
Original file line number Diff line number Diff line change
Expand Up @@ -243,8 +243,9 @@ ones its own pack turned on. It is asked only where the surname list
could not divide an unspaced token, and not where the name is already
divided — by a second word in an East Asian script (the nakaguro ・
counts as a space), a family comma or a 间隔号. A word in any other
script beside the token — Latin (``"Dr. 高橋一郎"``), Cyrillic, even
halfwidth katakana — divides nothing, so the segmenter is still asked. Recognize the text
script beside the token — Latin (``"Dr. 高橋一郎"``), Cyrillic — divides
nothing, so the segmenter is still asked. Halfwidth katakana counts as
katakana, so ``"高橋一郎 タロウ"`` is already divided. Recognize the text
you can actually read and return ``None`` for the rest, rather than
answering for a script you never meant to handle. A segmenter is your
code, so its failures do not stay inside the parse: its own exceptions
Expand Down
6 changes: 4 additions & 2 deletions docs/migrate.rst
Original file line number Diff line number Diff line change
Expand Up @@ -600,8 +600,10 @@ token moves from ``first`` to ``last`` exactly as ``毛泽东`` does:
``first``. The spaced shapes change what the name renders as too:
``str(HumanName("高橋 みなみ"))`` was ``"高橋 みなみ"`` and is now
``"みなみ 高橋"``. A name written *wholly* in katakana is deliberately
left alone — it is usually a transcribed foreign name already in
given-first order — so ``HumanName("マイケル ジャクソン")`` reads
left alone — it may be a transcribed foreign name already in
given-first order or a Japanese name's reading written family-first,
and the script cannot tell which (see :ref:`east-asian-names` to opt
in) — so ``HumanName("マイケル ジャクソン")`` reads
``first="マイケル"``/``last="ジャクソン"`` on both versions.

One more shape changes for a different reason: the katakana middle dot
Expand Down
2 changes: 2 additions & 0 deletions docs/release_log.rst
Original file line number Diff line number Diff line change
Expand Up @@ -66,6 +66,8 @@ Release Log

- **Fix a v1 ``Constants`` entry that is only a CJK full stop raising at the first parse.** After ``c.titles.add("。")``, ``HumanName("john smith", c)`` raised ``ValueError`` in 2.3; ``。``, ``.`` and ``。`` in any set or as a ``capitalization_exceptions`` key are now ignored with the same ``UserWarning`` as an entry with stray whitespace. 1.4.0 through 2.2.0 accepted such an entry and applied it to a name token that is nothing but that full stop, which 2.3 and later cannot match. An entry ending in such a full stop no longer raises either: ``c.suffix_not_acronyms.add("ma。")`` made ``HumanName("jack ma", c)`` raise ``ValueError`` in 2.3, and now gives last ``ma``, as without the entry. (closes #582)

- **Fix a Japanese name written in halfwidth katakana being read given-first.** ``HumanName("山田 タロウ")`` gives last ``山田``, first ``タロウ``, as ``山田 タロウ`` does, where every release gave first ``山田``, last ``タロウ``. Halfwidth katakana (U+FF65–U+FF9F), which legacy bank, payroll and CSV exports still carry, is now read as katakana everywhere the script matters: a second halfwidth word keeps the segmenter from re-dividing a kanji name (``高橋一郎 タロウ``), the Chinese ``·`` divides between halfwidth kana (``タロウ·ヤマダ`` gives first ``タロウ``, last ``ヤマダ`` where it was one first name), and a period-marked halfwidth word is no longer taken for an abbreviated title: ``タナカ. John`` gives first ``タナカ.`` where it gave title ``タナカ.``. A name written wholly in katakana, halfwidth or not, still keeps the declared order, given-first by default (``ヤマダ タロウ`` gives first ``ヤマダ``), because the script cannot say whether it is a Japanese name or a transcribed foreign one; to read your katakana names family-first, add ``(Script.KATAKANA, FAMILY_FIRST)`` to ``Policy.script_orders`` (see :ref:`east-asian-names`). See the ``W4`` entry of ``docs/design/decisions.md`` (closes #594)

**Additions**

- **Add Lexicon.conjunctions_ambiguous, the one-letter connectives that read as initials.** A subset of ``conjunctions`` holding ``e`` and ``i`` by default; it is the knob for the change above rather than a switch. Portuguese data, where ``e`` links surnames the way ``y`` does in Spanish, takes it out: ``Lexicon.default().remove(conjunctions_ambiguous={"e"})`` restores the joining reading. Dutch data, where a bare single letter is an initial and never a connective, adds the other one: ``Lexicon.default().add(conjunctions_ambiguous={"y"})``. A v1 ``Constants`` has no manager of its own for it -- deleting the word from ``conjunctions`` is what turns the marking off, the same rule the glued-honorific tails follow. See ``docs/customize.rst`` (#383, #479)
Expand Down
31 changes: 28 additions & 3 deletions docs/usage.rst
Original file line number Diff line number Diff line change
Expand Up @@ -303,7 +303,11 @@ ambiguous: native given names use it, but katakana is also how
Japanese text writes a *foreign* name — マイケル・ジャクソン is Michael
Jackson — and a transcription keeps the source language's order, given
name first, its parts divided by the middle dot ・ (the nakaguro,
U+30FB) rather than by a space.
U+30FB) rather than by a space. Katakana also has a halfwidth form
(タロウ for タロウ), which older systems that could not store kanji used
for every name, Japanese or foreign; bank, payroll and CSV exports
still carry it. nameparser reads halfwidth katakana exactly as it reads
the full-width form.

What happens automatically
^^^^^^^^^^^^^^^^^^^^^^^^^^
Expand All @@ -321,6 +325,8 @@ transcription, because a transcription is written in katakana alone.
('高橋', 'みなみ')
>>> parse("山田 エミ").family
'山田'
>>> parse("山田 タロウ").family
'山田'

And the middle dot separates tokens the way a space does, so a
transcribed foreign name divides into its parts — which, being wholly
Expand Down Expand Up @@ -417,8 +423,27 @@ Romanized names ("Kim Min-jun", "Yamada Taro") are Latin script and
follow the ordinary positional rules. Order genuinely varies in
romanized data, so nothing script-based applies. A name written wholly
in katakana stays positional for the reason given above, pack or no
pack: it is predominantly a transcription, and a transcription is
already in the order it should be read in.
pack: it may be a transcription, already in the order it should be
read in, or a Japanese name's reading written family-first — a
furigana field, or legacy halfwidth data — and the script cannot tell
the two apart. If you know your katakana names are Japanese, map
katakana to family-first yourself:

.. doctest::

>>> from nameparser import (DEFAULT_SCRIPT_ORDERS, FAMILY_FIRST, Parser,
... Policy, Script)
>>> parse("ヤマダ タロウ").family
'タロウ'
>>> kana_family_first = Parser(policy=Policy(script_orders=(
... *DEFAULT_SCRIPT_ORDERS, (Script.KATAKANA, FAMILY_FIRST))))
>>> kana_family_first.parse("ヤマダ タロウ").family
'ヤマダ'

This reaches full-width katakana too, transcriptions included
(マイケル・ジャクソン would read family マイケル), which is why it is
yours to choose rather than a default. ``name_order=FAMILY_FIRST``
would also work, but it reverses every name, Latin ones included.

A Han transcription written with a space instead of the 间隔号
(威廉 莎士比亚) carries nothing to distinguish it from a native
Expand Down
6 changes: 4 additions & 2 deletions nameparser/_pipeline/_vocab.py
Original file line number Diff line number Diff line change
Expand Up @@ -1431,8 +1431,10 @@ def effective_script(text: str) -> Script | None:
katakana-only: マイケル has no kanji, but さくらエミ -- hiragana
plus katakana -- is kana-only AND licensed) -- and resolves to the
HIRAGANA carrier entry. Pure-katakana stays KATAKANA
(single_script's answer): a lone katakana token is predominantly a
transcribed foreign name, so nothing defaults on it."""
(single_script's answer): a lone katakana token may be a
transcribed foreign name or a Japanese reading, and the script
cannot say which, so nothing defaults on it. Halfwidth kana is
katakana by the table (#594), so 山田タロウ is licensed too."""
# None for both shapes _wholly_ja could never match anyway (empty
# text, or all-ASCII text): real work, not a leftover "if text"
# guard, since the ASCII case is one a bare emptiness check would
Expand Down
Loading
Loading