parse("山田 エミ") reads family 山田, given エミ: kanji with kana is licensed as Japanese and assigned family-first. The same name written with halfwidth katakana, parse("山田 タロウ"), reads given 山田, family タロウ, which is the positional reading. Script.KATAKANA covers U+30A0–30FF only, so halfwidth kana (U+FF65–FF9F) is unclassified and the name never reaches the kana license.
>>> from nameparser import parse
>>> parse("山田 エミ").family
'山田'
>>> parse("山田 タロウ").family
'タロウ'
The same gap reaches the segmenter. A halfwidth-katakana word beside a Han token doesn't count as a second East Asian word, so "高橋一郎 タロウ" still asks the segmenter to divide 高橋一郎, where "高橋一郎 タロウ" doesn't.
Halfwidth katakana still turns up in legacy Japanese data (old JIS systems, bank and payroll exports). The question is whether to treat it as katakana for script classification (NFKC maps タロウ → タロウ), and if so, whether that belongs in classification (widen the range) or as normalization at tokenize. Note that U+FF65, the halfwidth nakaguro, is already a token separator.
Found during the docs review in #593. The docs describe the license as "kanji and kana other than katakana alone" and leave halfwidth unmentioned until this is decided.
parse("山田 エミ")reads family山田, givenエミ: kanji with kana is licensed as Japanese and assigned family-first. The same name written with halfwidth katakana,parse("山田 タロウ"), reads given山田, familyタロウ, which is the positional reading.Script.KATAKANAcovers U+30A0–30FF only, so halfwidth kana (U+FF65–FF9F) is unclassified and the name never reaches the kana license.The same gap reaches the segmenter. A halfwidth-katakana word beside a Han token doesn't count as a second East Asian word, so
"高橋一郎 タロウ"still asks the segmenter to divide高橋一郎, where"高橋一郎 タロウ"doesn't.Halfwidth katakana still turns up in legacy Japanese data (old JIS systems, bank and payroll exports). The question is whether to treat it as katakana for script classification (NFKC maps
タロウ→タロウ), and if so, whether that belongs in classification (widen the range) or as normalization at tokenize. Note that U+FF65, the halfwidth nakaguro, is already a token separator.Found during the docs review in #593. The docs describe the license as "kanji and kana other than katakana alone" and leave halfwidth unmentioned until this is decided.