Skip to content

docs: review locales.rst and migrate.rst against the tree; state the kana license correctly - #593

Merged
derek73 merged 7 commits into
masterfrom
docs/locales-migrate-review
Oct 3, 2026
Merged

derek73 merged 7 commits into
masterfrom
docs/locales-migrate-review

Conversation

@derek73

@derek73 derek73 commented Oct 3, 2026 •

Copy link
Copy Markdown
Owner

Two parallel reviewers checked locales.rst and migrate.rst against the current tree. Each claim was measured, then fixed. There are five commits: one cross-page fix, then drift and structure for each page.

Cross-page: the kana license (638d2315)
Five pages paraphrased the script_orders kana rule as "kanji mixed with kana", including concepts.rst's wording from #592. The predicate admits any name within kanji and kana that holds some kana and isn't katakana alone. Wholly-hiragana やまだ はなこ and unspaced さくらエミ read family-first; ヤマダ タロウ stays positional. Fixed in concepts, customize (three places), locales, modules and usage.

locales.rst drift (4fb7e253), each measured:

  • given_name_titles alone doesn't raise; the entry is inert.
  • The "a fragment is validated alone" example used given_name_titles, which works fine in a fragment. The rule really applies to particles_ambiguous, suffix_acronyms_ambiguous and honorific_tails.
  • Merge rules: two packs setting the same scalar warn even when the values agree, a pack overrides the base silently, and segment_scripts unions like the other set fields.
  • ru and tr_az rows: the ru shape was described backwards (it is family/given/patronymic). The tr_az row now says four words ending in the marker.
  • Segmenters aren't offered every token. A wrong-type answer or a cut outside the token raises too.
  • ja without a segmenter now warns.
  • The stand-down example was a name the ru rule doesn't match at all.
  • Counts and links: "seven scripts" was a standing count that also left out the CJK vocabulary. The page cited Wrong parsing of vietnamese names #146, which is closed; it now cites Ship hi and bn locale packs for Latin-transliterated Indic honorifics #345.

locales.rst structure (b1ce51f5): subheadings throughout. The merge rules move to "Stacking packs" as a list. The stand-down note leaves the warning box. The Kapitan example now comes before the validation rule. Doctest variables get distinct names, and :doc: links become section :ref:s.

migrate.rst drift (f6c22760):

  • capitalization_exceptions advice: the guide's dataclasses.replace(..., capitalization_exceptions={...}) replaced every shipped mask, so phd repaired to PHD, and it failed mypy. It now links customize.rst's extend-the-defaults recipe.
  • 1.x docs link: it pointed at readthedocs stable, which now serves 2.3.0. It now targets /en/v1.4.0/.
  • Breaks after 2.0: "runs clean on 1.4 → runs on 2.0, four exceptions" is scoped to 2.0.0, and a new list covers later 2.x breaks 1.4 never warned about: TITLES.add raises, retired config names warn, non-mask exceptions raise, and unmatchable entries are dropped with a warning.
  • Behavior changes now point at every 2.x release-log section, with a new list of the shapes 2.2–2.4 moved on HumanName. Each 1.4.0 reading was measured with PYTHONSAFEPATH=1 from outside the worktree (asserting __file__), and each current reading is pinned by a doctest.
  • Small additions: the Lexicon fields with no v1 attribute are named, and the dean recipes are now doctests.

migrate.rst structure (71079efa): subheadings in "Before you upgrade", "Config map" and "Behavior changes". The particles flip warning moves up next to the table row that points to it; it used to sit about 200 lines below. Over-long cells are shortened, and section :ref:s are added.

Review rounds. Two reviewers checked the first five commits, and a third checked the fix commit. Every finding was verified against the parser, or against 1.4.0, 2.1.0, 2.2.0 and 2.3.0 installs, before fixing (fab01a13, f247082b):

  • migrate.rst: every 1.4.0 reading held, but my one-line explanations of why each changed didn't:
    • The Queen/Prince change was really two 2.3 mechanisms: a run of titles matched by its last title, and newly added titles that address by given name. Dr Harry is unchanged.
    • md simply left the exceptions map.
    • In an all-lowercase or all-caps name, only e reads as an initial; y still joins.
    • i links two surnames only in a mixed-case name.
    • The capitals and dotted-credential bullets now state their conditions, with unchanged counterexamples.
    • No listed shape changed in 2.2, so the subsection is "Changed in 2.3 and 2.4" and each bullet names its release.
    • The unmatchable-entry warning also covers entries made only of full stops, and capitalization_exceptions keys.
  • locales.rst:
    • Stand-down: under FAMILY_FIRST_GIVEN_LAST the declared order wins and the last word reads as the given name (for tr_az that's the separate oglu marker).
    • Segmenter: it's still asked when a word in another script sits beside the token (Dr. 高橋一郎).
    • Fragment validation: only particles_ambiguous, suffix_acronyms_ambiguous and honorific_tails are checked against the set they mark.
    • Script coverage: katakana ships no honorific.
    • Pack warnings: the scalar-conflict warning fires only between packs with different codes.

Out of scope, filed as #594: halfwidth katakana (タロウ) is unclassified, so 山田 タロウ stays positional where 山田 エミ reads family-first.

No behavior change. sphinx-build -b doctest docs (0 failures) and -b html both print nothing. A diff of inline literals per page shows nothing dropped except deliberate link retargets.

🤖 Generated with Claude Code

derek73 and others added 5 commits October 2, 2026 22:40
Five pages paraphrased the script_orders kana license as "kanji mixed
with kana" (concepts.rst's #592 wording, "mixing kanji with kana or the
two kanas", was narrower still). The predicate (_vocab's kana license,
plus HIRAGANA's own script_orders entry) admits any name within kanji
and kana that holds some kana and is not katakana alone: wholly
hiragana やまだ はなこ and unspaced さくらエミ read family-first along
with 山田 エミ and 高橋 みなみ, while ヤマダ タロウ stays positional
(measured). customize.rst (three sites), locales.rst, concepts.rst and
modules.rst now say "kanji and kana other than katakana alone".
usage.rst stated the condition right but omitted the katakana
exception from it, and gave the reason as "a transcription would have
been kana alone" -- it is katakana alone.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
Each measured on the tree:
- given_name_titles alone does not raise; the entry is inert (the
  check was dropped 2026-07-19 and AGENTS.md forbids its return).
- The validate-alone example used given_name_titles, which works fine
  in a fragment; the rule bites for particles_ambiguous,
  suffix_acronyms_ambiguous and honorific_tails, which raise
  ValueError when the word they mark is absent.
- Merge rules: two packs setting the same scalar warn even when the
  values agree, a pack overrides the base silently, and
  segment_scripts unions like the other set fields.
- The ru row described the shape backwards; it is family/given/
  patronymic, three words, the last patronymic-shaped and the middle
  not. The tr_az row now says four words ending in the marker, the
  first read as family (Aliyev Ilham Heydar oglu).
- A segmenter is not offered "every token": only an unspaced token the
  surname list could not divide, and not where a space, family comma
  or 间隔号 already divides the name; and a wrong-type answer or an
  out-of-token cut raises, not only the segmenter's own exceptions.
- ja without a segmenter now warns at construction.
- The stand-down example (Мицкевич Адам Юзеф) is a name the ru rule
  does not match at all; replaced with one it does, and the claim
  restated: on a matched name the declared order and the rotation
  agree, so they never compete.
- "seven scripts" was a standing count that also left out the CJK
  vocabulary; the list stays, the count goes.
- #146 (Vietnamese) closed 2026-08-07; locales.rst now cites the open
  pack issue #345, and customize.rst drops its "no vn pack yet
  (#146)" pointer.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- "What works without a pack": "Adding a word the defaults hold back"
  and "East Asian names need no pack", which now also links the
  opt-out (east-asian-defaults).
- "Using a pack": "Stacking packs" (the merge rules move here from the
  end of "Creating your own Locale", as a three-item list),
  "Finding packs by code" (the --locale equivalence stated once, not
  twice) and "Shipped packs".
- The family-first stand-down note leaves the shape-not-language
  warning box, which it is not about.
- "Creating your own Locale": the Kapitan example now follows the
  PolicyPatch introduction it illustrates, and the validate-alone rule
  gets "The lexicon fragment is validated on its own".
- Contributing item 3's three pack kinds are sub-bullets.
- :doc: links become section refs: concepts' containers section
  (newly labeled config-containers) and ambiguous-words.
- Doctest variables get distinct names (shaikh, sayyid_lex, sayyid,
  script_pack, kapitan_lex, corp_pack, kapitan) instead of rebinding
  name/lex/mine.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- The capitalization_exceptions advice, dataclasses.replace(lexicon,
  capitalization_exceptions={...}), replaced every shipped mask
  ("john smith phd" then repairs to PHD) and failed mypy; both sites
  now point at customize.rst's extend-the-defaults recipe (newly
  labeled case-exceptions).
- The 1.x docs link pointed at readthedocs "stable", which now serves
  2.3.0; it targets /en/v1.4.0/.
- "Runs clean on 1.4 under -W error ... it will run on 2.0, with four
  exceptions" is scoped to 2.0.0 and loses its count, and a second
  list names what later 2.x releases broke without a 1.4 warning:
  TITLES.add raising (2.2), retired config names warning (2.2), a
  non-mask capitalization_exceptions value raising (2.4), an
  unmatchable Constants entry dropped with a UserWarning (2.4).
- "Behavior changes" pointed only at the 2.0.0 release-log section and
  stopped at 2.1. It now points at every 2.x section and lists the
  shapes 2.2-2.4 moved on HumanName, each 1.4.0 reading measured with
  PYTHONSAFEPATH=1 from outside the worktree (nameparser.__file__
  asserted) and each current reading pinned by a doctest.
- Names the Lexicon fields with no CONSTANTS attribute
  (conjunctions_ambiguous, maiden_markers, surnames, honorific_tails).
- The dean recipes become doctests, with the default reading shown so
  they are non-vacuous.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
- "Before you upgrade" splits into what breaks without a warning, the
  silent == change, the re-pickle step, and the warned-removal table.
- "Config map" splits into "Vocabulary sets → Lexicon", "Renamed word
  lists (2.2)", "Default word lists are frozen" and "Behavior and
  render settings → Policy". The particles flip warning moves up
  beside the table row whose "see the warning below" it answers -- it
  sat about 200 lines further down -- and the rename caveat's pointer
  to it now says "above".
- "Behavior changes" splits into "Changed in 2.0", "East Asian names
  (2.1)", "Turning the East Asian readings off" and "Changed in
  2.2–2.4".
- The suffix_delimiter cell shrinks to one line plus a ref; its
  scalar-raises detail moves into prose below the table rather than
  being lost.
- Table rows and prose gain section refs: rendering-arguments,
  suffix-delimiters, brackets, strip-flags, east-asian-names,
  east-asian-defaults.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@derek73 derek73 added the docs Documentation fixes and updates label Oct 3, 2026
@derek73 derek73 self-assigned this Oct 3, 2026
@codecov

codecov Bot commented Oct 3, 2026 •

Copy link
Copy Markdown

Codecov Report

✅ All modified and coverable lines are covered by tests.
✅ Project coverage is 98.98%. Comparing base (7a9788a) to head (f247082).

Additional details and impacted files
@@           Coverage Diff           @@
##           master     #593   +/-   ##
=======================================
  Coverage   98.98%   98.98%           
=======================================
  Files          45       45           
  Lines        4220     4220           
=======================================
  Hits         4177     4177           
  Misses         43       43           

☔ View full report in Codecov by Harness.
📢 Have feedback on the report? Share it here.

🚀 New features to boost your workflow:
  • ❄️ Test Analytics: Detect flaky tests, report on failures, and find test suite problems.

derek73 and others added 2 commits October 2, 2026 22:50
migrate.rst -- every 1.4.0 reading held; the one-line causes did not:
- Queen/Prince: only a given-name title makes the lone name given
  ("Dr Harry" still reads last); "a form of address" was too broad.
- md -> MD: md left the exceptions map; it was not a mask change.
- "a one-letter connective ... is an initial" holds for e only (y still
  joins); "i links two surnames" holds in a mixed-case name only.
- "a capitalized credential" and "an unlisted dotted credential" were
  broader than the rules (Jack Ma, JACK MA and Jack X.Y.Z. unchanged);
  each now states its condition and an unchanged counterexample.
- No listed shape changed in 2.2 (2.2.0 reads all as 1.4, measured);
  the subsection is "Changed in 2.3 and 2.4" and each bullet names its
  release (2.3.0 measured too).
- The unmatchable-entry warning also covers entries that fold to empty
  (full stops) and capitalization_exceptions keys.
- The frozen-sets bullet links its own subsection, and the moved flip
  warning points forward to the rename it cites.

locales.rst:
- The family-first stand-down is not "never disagree" under
  FAMILY_FIRST_GIVEN_LAST, where the patronymic reads as the given
  name; both orders are now stated, and the stand-down names both
  rotations.
- The segmenter is still asked beside a Latin word ("Dr. 高橋一郎"); only
  a second East Asian word, the nakaguro, a family comma or a 间隔号
  divides the name first.
- Only particles_ambiguous, suffix_acronyms_ambiguous and
  honorific_tails are checked against their parent set;
  given_name_titles and conjunctions_ambiguous are not.
- Katakana ships no honorific, and the East Asian scripts ship no
  conjunction or particle.
- Scalar-conflict warnings fire between packs of different codes.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
…ixes

- migrate.rst: "(Queen, Prince, Princess)" read as the list of
  given-name titles; sir, dame, king and swami are too, and the 2.3
  change was two mechanisms -- a title run matched by its last title
  (Her Majesty Queen, Rev Sir: measured 2.2.0 last -> 2.3.0 first) and
  newly added given-name titles (Prince, Princess, Swami, Guru). Both
  named; no list claims completeness.
- locales.rst: under FAMILY_FIRST_GIVEN_LAST the last word reads as the
  given name, and for tr_az that is often the separate oglu/qizi
  marker, not the patronymic ("Aliyev Ilham Heydar oglu" gives given
  oglu).
- locales.rst: any non-CJK-classified neighbour leaves the segmenter
  asked -- Cyrillic and halfwidth katakana as well as Latin.
- locales.rst: "the East Asian scripts honorifics only" took katakana
  back in; the scripts are named.
- Rewraps a line the previous fix left long.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@derek73
derek73 merged commit d989c5e into master Oct 3, 2026
11 checks passed
@derek73
derek73 deleted the docs/locales-migrate-review branch October 3, 2026 20:05
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

docs Documentation fixes and updates

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant