diff --git a/docs/customize.rst b/docs/customize.rst index b3231724..aa900db1 100644 --- a/docs/customize.rst +++ b/docs/customize.rst @@ -321,7 +321,9 @@ listed below. - ``frozenset[PatronymicRule]`` - Reorders patronymic-shaped names via opt-in detectors — East Slavic formal order (``EAST_SLAVIC``) or Turkic reversed order - (``TURKIC``). Defaults to empty. + (``TURKIC``) — but stands down under a declared + ``FAMILY_FIRST`` or ``FAMILY_FIRST_GIVEN_LAST`` ``name_order``. + Defaults to empty. * - ``middle_as_family`` - ``bool`` - Folds ``middle`` into ``family`` instead of splitting them — diff --git a/docs/locales.rst b/docs/locales.rst index 6eaed404..fb1d9546 100644 --- a/docs/locales.rst +++ b/docs/locales.rst @@ -180,6 +180,12 @@ new naming rule belongs in. mixes traditions, parse the subsets separately with different parsers rather than enabling a pack over all of it. + A declared family-first ``name_order`` stands down the rotation + instead of competing with it: fold the pack onto a base parser + built with ``Policy(name_order=FAMILY_FIRST)`` and ``"Мицкевич + Адам Юзеф"`` reads family ``Мицкевич`` rather than the given-first + order the pack restores by default. + .. _segmenter-contract: Segmenters diff --git a/docs/release_log.rst b/docs/release_log.rst index 0477cdd3..c66a2e9f 100644 --- a/docs/release_log.rst +++ b/docs/release_log.rst @@ -10,31 +10,31 @@ Release Log - **Record a 2.0.0 change to HumanName.initials() that no release note had classified:** since 2.0.0 the facade initials each WORD of a name part, where 1.4.0 initialed a joined run as one group -- ``HumanName("Juan Velasquez y Garcia").initials()`` is ``J. V. G.`` and was ``J. V G.``; ``Abdul Salam Hassan`` is ``A. S. H.`` and was ``A S. H.``. Nothing changes in 2.3.0; the differential gate now compares ``initials()`` (#484) and this is what it found. See the ``differential-ledger, the initials view`` entry of ``docs/design/decisions.md`` - - **Fix a space-separated run of post-nominals rendering with a comma the writer never typed.** ``HumanName("John Smith MD PhD").suffix`` is ``MD PhD`` and was ``MD, PhD`` at every release since 1.4.0; ``Kenneth Clarke QC MP`` gives ``QC MP``, and the CJK honorific runs (``田中さん II``, ``김민준 박사 씨``) follow the same rule. This is a deliberate deviation from 1.4.0, which inserted a comma into a run the writer had spaced; the v1-parity suite is re-pinned to match. The comma forms are unchanged -- ``HumanName("Smith, MD, PhD").suffix`` is still ``MD, PhD`` -- because the separator is now the comma the writer typed rather than the shape of the comma segments, and a configured suffix delimiter still parts a run, as does a name word standing between two post-nominals. Round-tripping is fixed for these runs, which the 2.2.0 note below recorded as broken: ``str(HumanName("Smith, MD PhD"))`` is ``Smith MD PhD`` and re-parses to suffix ``MD PhD``, where the comma-written and space-written spellings of one run used to disagree. ``str()`` is still a rendering rather than a canonical form, and one name goes the other way: ``HumanName("Smith, John PhD I.")`` renders ``John Smith PhD I.``, which reads middle ``Smith PhD`` and family ``I.`` as it always has, so that name round-tripped only on the strength of the comma this fix removes. Thirteen names that predate this change move in the differential corpora (fourteen with the example this change adds), all in the same direction, and no other field view moves. See the ``C1`` entry of ``docs/design/decisions.md`` (closes #436, closes #437) + - **Fix a space-separated run of post-nominals rendering with a comma the writer never typed.** ``HumanName("John Smith MD PhD").suffix`` is ``MD PhD`` and was ``MD, PhD`` at every release since 1.4.0; ``Kenneth Clarke QC MP`` gives ``QC MP``, and the CJK honorific runs (``김민준 박사 씨``) follow the same rule. This is a deliberate deviation from 1.4.0, and the v1-parity suite is re-pinned to match. The comma forms are unchanged -- ``HumanName("Smith, MD, PhD").suffix`` is still ``MD, PhD`` -- because the separator is now the comma the writer typed; a configured suffix delimiter still parts a run, as does a name word standing between two post-nominals. Round-tripping is fixed for these runs, which the 2.2.0 note below recorded as broken: ``str(HumanName("Smith, MD PhD"))`` is ``Smith MD PhD`` and re-parses to suffix ``MD PhD``. ``str()`` is still a rendering rather than a canonical form. See the ``C1`` entry of ``docs/design/decisions.md`` (closes #436, closes #437) - - **Fix Parser.revise() splitting a space-separated suffix value into comma-separated entries.** ``Parser().revise(n, suffix="MD PhD").suffix`` is ``MD PhD`` and was ``MD, PhD``, and a name's own rendered suffix now revises back to itself for all but one of the 38 differential-corpus names that failed to on 2026-09-06 (368 of the 1117 carry a suffix) (recipe in the ``C1`` entry of ``docs/design/decisions.md``; the one left is a Korean honorific glued to an initial, which the value's own parse peels where the whole name kept it glued -- a word read on its own, not an entry boundary). A suffix value's entries are now derived from the value's own commas by the same rule a whole name uses: a comma parts two credentials and a space joins them, so ``revise(n, suffix="MD, PhD")`` is still two entries. ``revise(n, suffix="Ph. D.")`` renders ``Ph. D.`` where the 2.2.0 note below accepted ``Ph., D.``; the head-position merge that note describes still does not fire, the pair joining under the entry rule instead. One limit: a delimiter configured through ``extra_suffix_delimiters`` parts a value only where the value's own words read as a name with a tail segment, so in a run of post-nominals it stays a word; write a comma at the boundary instead. ``ParsedName.replace()`` is unchanged (closes #511) + - **Fix Parser.revise() splitting a space-separated suffix value into comma-separated entries.** ``Parser().revise(n, suffix="MD PhD").suffix`` is ``MD PhD`` and was ``MD, PhD``. A suffix value's entries are now derived from the value's own commas by the rule a whole name uses -- a comma parts two credentials and a space joins them -- so ``revise(n, suffix="MD, PhD")`` is still two entries, and ``revise(n, suffix="Ph. D.")`` renders ``Ph. D.`` where the 2.2.0 note below accepted ``Ph., D.``. One limit: a delimiter configured through ``extra_suffix_delimiters`` parts a value only where the value's own words read as a name with a tail segment, so in a run of post-nominals it stays a word; write a comma at the boundary instead. ``ParsedName.replace()`` is unchanged. See the ``C1`` entry of ``docs/design/decisions.md`` (closes #511) - - **Remove rai and cha from the default post-nominal acronyms, so a trailing Rai or CHA keeps the family name.** ``HumanName("Aishwarya Rai")`` gives last ``Rai``, where 2.0.0 through 2.2.0 gave suffix ``Rai`` and no last name at all -- 1.4.0's reading, restored -- and ``Lala Lajpat Rai`` gives middle ``Lajpat``, last ``Rai``. Both entries came in with a 2019 bulk import of Wikipedia post-nominals and were never reviewed against the surnames they collide with. The cost is that a genuine credential written after a full name is no longer recognized: ``John Smith RAI`` gives middle ``Smith``, last ``RAI``, and the comma forms swap ends -- ``John Smith, RAI`` gives first ``RAI``, last ``John Smith``, and ``Ahmad Jayadi, CHA`` first ``CHA``, last ``Ahmad Jayadi``. A caller who needs either back adds it: ``Lexicon.default().add(suffix_acronyms={"cha"})``. Five names move in the differential corpora, all of them radar-tier. See the ``suffix-acronym-collisions`` entry of ``docs/design/decisions.md`` (closes #342) + - **Remove rai and cha from the default post-nominal acronyms, so a trailing Rai or CHA keeps the family name.** ``HumanName("Aishwarya Rai")`` gives last ``Rai``, where 2.0.0 through 2.2.0 gave suffix ``Rai`` and no last name at all -- 1.4.0's reading, restored. The cost is that a genuine credential written after a full name is no longer recognized: ``John Smith RAI`` gives last ``RAI``, and ``John Smith, RAI`` gives first ``RAI``, last ``John Smith``. A caller who needs either back adds it: ``Lexicon.default().add(suffix_acronyms={"cha"})``. See the ``suffix-acronym-collisions`` entry of ``docs/design/decisions.md`` (closes #342) - - **Mark ba as an acronym that is also an ordinary name, so a bare trailing Ba keeps the family name.** ``HumanName("Anna Ba")`` gives last ``Ba`` and reports a suffix-or-name ambiguity, where 2.0.0 through 2.2.0 gave suffix ``Ba`` and no last name. The SPACED full-name form keeps the credential reading: ``John Smith BA`` still gives suffix ``BA``, now flagged, and the dotted ``John Smith B.A.`` is an unflagged suffix, the periods settling it. The COMMA forms move, and this is the marking's real cost: ``Smith, BA`` gives first ``BA``, and ``John Smith, BA`` gives first ``BA``, last ``John Smith``, where 2.0.0 through 2.2.0 gave suffix ``BA`` for both -- what ``Smith, Ed`` costs, which S2 already accepted for the other ambiguous acronyms. A bracketed or quoted ``John Smith (BA)`` falls through to nickname parsing, as the 2.0 note for ``ma``/``do`` below recorded for that pair. Write ``B.A.`` to keep the credential reading. BA is a common credential and Ba a real surname in Vietnamese and Senegalese Fula, which is the ``ma``/``Ma`` shape exactly. No corpus name moves (#342) + - **Mark ba as an acronym that is also an ordinary name, so a bare trailing Ba keeps the family name.** ``HumanName("Anna Ba")`` gives last ``Ba`` and reports a suffix-or-name ambiguity, where 2.0.0 through 2.2.0 gave suffix ``Ba`` and no last name. The spaced full-name form keeps the credential reading -- ``John Smith BA`` still gives suffix ``BA``, now flagged -- and the dotted ``John Smith B.A.`` is an unflagged suffix. The comma forms move, which is the marking's cost: ``John Smith, BA`` gives first ``BA``, last ``John Smith``, as ``Smith, Ed`` already did for the other ambiguous acronyms. Write ``B.A.`` to keep the credential reading. Ba is a real surname in Vietnamese and Senegalese Fula, the ``ma``/``Ma`` shape exactly (#342) - - **Fix a title run addressing by its first title rather than its last.** ``HumanName("Her Majesty Queen Elizabeth")`` gives first ``Elizabeth`` with an empty last name, where every release since 1.4.0 gave last ``Elizabeth``. Several titles written together are one form of address and the one that does the addressing is the last, so the run is now matched whole or by its last word: ``Reverend Mother Teresa``, ``Dr. Sir John`` and ``Mr Sir John`` move the same way, and ``Sir Sheikh abdul rahman`` gives first ``abdul rahman``. What does NOT move is a run whose last word addresses by surname: ``His Excellency Lord Duncan`` still gives last ``Duncan``, ``lord`` not being a given-name title. A caller's multi-word entry still matches as a phrase. Two names in the differential corpora read differently for this rule. See the ``H1`` entry of ``docs/design/decisions.md`` (closes #489) + - **Fix a title run addressing by its first title rather than its last.** ``HumanName("Her Majesty Queen Elizabeth")`` gives first ``Elizabeth`` with an empty last name, where every release since 1.4.0 gave last ``Elizabeth``. Several titles written together are one form of address and the one that does the addressing is the last, so the run is now matched whole or by its last word: ``Reverend Mother Teresa``, ``Dr. Sir John`` and ``Sir Sheikh abdul rahman`` move the same way. A run whose last word addresses by surname does not move: ``His Excellency Lord Duncan`` still gives last ``Duncan``, ``lord`` not being a given-name title. A caller's multi-word entry still matches as a phrase. See the ``H1`` entry of ``docs/design/decisions.md`` (closes #489) - - **Fix the leading title peel taking a name word and leaving a post-nominal to be the name.** ``HumanName("Dr King Jr")`` gives title ``Dr``, last ``King``, suffix ``Jr``, where every release since 1.4.0 gave title ``Dr King``, last ``Jr`` and no suffix at all; ``Dr. King MD`` moves the same way, and both now read as the comma spelling ``King, Dr Jr`` always has. A title addresses somebody, so the run leaves a name word standing and a post-nominal is not one. A name that is nothing but titles or nothing but post-nominals is untouched, the word given back having to be a name candidate: ``Marquess of Bath``, ``MD DDS`` and ``Jr. Ph. D.`` are unchanged, and so is a title written as one joined unit -- ``Prince of Wales Jr`` keeps title ``Prince of Wales`` rather than losing the title to make a name. Where the run's whole content is the word given back there is no title left, so ``Dr Jr`` gives first ``Dr``, suffix ``Jr`` and reports a title-or-name ambiguity. Three names in the differential corpora read differently for this rule. See the ``H3`` entry of ``docs/design/decisions.md`` + - **Fix the leading title peel taking a name word and leaving a post-nominal to be the name.** ``HumanName("Dr King Jr")`` gives title ``Dr``, last ``King``, suffix ``Jr``, where every release since 1.4.0 gave title ``Dr King``, last ``Jr`` and no suffix at all; ``Dr. King MD`` moves the same way, and both now read as the comma spelling ``King, Dr Jr`` always has. A name that is nothing but titles or nothing but post-nominals is untouched (``Marquess of Bath``, ``MD DDS``), and so is a title written as one joined unit: ``Prince of Wales Jr`` keeps title ``Prince of Wales``. Where the only word left is the title itself, ``Dr Jr`` gives first ``Dr``, suffix ``Jr`` and reports a title-or-name ambiguity. See the ``H3`` entry of ``docs/design/decisions.md`` - - **Fix a trailing abbreviated title reading as a name word.** ``HumanName("John Smith Prof.")`` gives title ``Prof.``, first ``John``, last ``Smith``, where every release since 1.4.0 gave last ``Prof.`` and lost the surname; ``John Smith Mr.``, ``John Smith Rev.``, ``John Smith Dr.`` and ``Andrew Perkins (Mgr.)`` move the same way. A run chains from the end (``John Smith Prof. Dr.`` gives title ``Prof. Dr.``), a leading title keeps its place (``Dr. John Smith Prof.`` gives title ``Dr. Prof.``), and the comma forms agree with the bare ones now -- ``Smith, John Prof.`` gives title ``Prof.``, first ``John``, last ``Smith`` where it gave middle ``Prof.`` at every release. The trailing title is transparent to the post-nominal reading, so ``John Smith Jr. Prof.`` gives suffix ``Jr.`` rather than promoting the generational suffix to the last name. What does NOT move: an unlisted abbreviation (``John Smith Xyz.`` keeps last ``Xyz.``), a bare title word (``John Smith Sir``, ``Mary Jane King``) and a post-nominal (``John Smith Esq.``). Only a listed title word wearing the abbreviation period is claimed -- the leading slot infers a title from the shape alone, the trailing slot never does. The reach is the whole title vocabulary, ordinary surnames in it included, so a period written behind one of them takes it out of the name: ``Mary Jane King.`` gives title ``King.``, first ``Mary``, last ``Jane``, where the bare ``Mary Jane King`` keeps last ``King``. That is accepted rather than prevented -- the period is a writing convention and not evidence about the word, and the bare spelling is what the trailing slot is protected from. A title run in FRONT of the name still does the addressing, so the trailing title is transparent to that reading too: ``Sir John Prof.`` gives title ``Sir Prof.``, first ``John`` -- ``Sir John`` plus a title -- and ``Dr. Smith Sir.`` gives title ``Dr. Sir.``, last ``Smith``. Where NO title stands in front, the trailing one decides the field: ``Smith Sir.`` gives first ``Smith`` and ``Smith Prof.`` gives last ``Smith``. Sixteen names in the differential corpora read differently for this rule, and a seventeenth for the same argument in a native script: ``毛 泽东 Dr.`` gives title ``Dr.``, first ``泽东``, last ``毛``, where 2.2.0 gave last ``Dr.`` and lost the family-first order. See the ``H5`` entry of ``docs/design/decisions.md`` (closes #316) + - **Fix a trailing abbreviated title reading as a name word.** ``HumanName("John Smith Prof.")`` gives title ``Prof.``, first ``John``, last ``Smith``, where every release since 1.4.0 gave last ``Prof.`` and lost the surname; ``John Smith Dr.``, ``John Smith Rev.`` and ``Andrew Perkins (Mgr.)`` move the same way, and the comma form ``Smith, John Prof.`` now agrees with the bare one where it gave middle ``Prof.``. A run chains from the end (``John Smith Prof. Dr.`` gives title ``Prof. Dr.``), a leading title keeps its place (``Dr. John Smith Prof.`` gives title ``Dr. Prof.``), and a post-nominal is read through it (``John Smith Jr. Prof.`` gives suffix ``Jr.``). Only a listed title word wearing the abbreviation period is claimed: an unlisted abbreviation (``John Smith Xyz.``), a bare title word (``John Smith Sir``) and a post-nominal (``John Smith Esq.``) do not move. The reach is the whole title vocabulary, ordinary surnames in it included, so ``Mary Jane King.`` gives title ``King.`` where the bare ``Mary Jane King`` keeps last ``King`` -- accepted rather than prevented, the period being evidence the bare spelling never gives. The same argument holds in a native script: ``毛 泽东 Dr.`` gives title ``Dr.``, first ``泽东``, last ``毛``, where 2.2.0 gave last ``Dr.``. See the ``H5`` entry of ``docs/design/decisions.md`` (closes #316) - - **Remove esq from the default post-nominal acronyms, and assert the two post-nominal sets disjoint.** ``HumanName("John Smith E.S.Q.")`` gives middle ``Smith``, last ``E.S.Q.``, where every release since 1.4.0 gave suffix ``E.S.Q.``. ``Esq``, ``Esq.``, ``ESQ`` and ``esq`` are unchanged, the post-nominal word list carrying every single-token spelling; the acronym entry's only unique coverage was the multi-dot spelling. Esquire is a contraction rather than an initialism, so the initialism set was never its home, and it was the one word in both post-nominal sets -- which is why the sets can now assert they do not overlap, a word in both being matched by two rules that normalize differently. A caller who needs it back adds it: ``Lexicon.default().add(suffix_acronyms={"esq"})``. One name moves in the differential corpora. See the ``suffix-acronym-collisions`` entry of ``docs/design/decisions.md`` + - **Remove esq from the default post-nominal acronyms, and assert the two post-nominal sets disjoint.** ``HumanName("John Smith E.S.Q.")`` gives middle ``Smith``, last ``E.S.Q.``, where every release since 1.4.0 gave suffix ``E.S.Q.``; ``Esq``, ``Esq.``, ``ESQ`` and ``esq`` are unchanged, the post-nominal word list carrying every single-token spelling. Esquire is a contraction rather than an initialism, and it was the one word in both post-nominal sets, which can now assert they do not overlap. A caller who needs the dotted spelling back adds it: ``Lexicon.default().add(suffix_acronyms={"esq"})``. See the ``suffix-acronym-collisions`` entry of ``docs/design/decisions.md`` - - **Fix the East Slavic and Turkic patronymic rotations overriding a declared family-first name order.** With ``patronymic_rules`` opted in and ``Policy(name_order=FAMILY_FIRST)``, ``Мицкевич Адам Юзеф`` gave last ``Адам`` through 2.2.0 and now gives last ``Мицкевич`` -- the reading the declaration asks for -- and ``oglu Ahmad Vali Ali`` with Turkic handling gave last ``Ahmad`` and now ``oglu``. The rotations exist to restore the given-first reading a family-first listing hides, so under a declared family-first order the declaration decides. No corpus name moves. See the ``O1`` entry of ``docs/design/decisions.md`` (closes #384) + - **Fix the East Slavic and Turkic patronymic rotations overriding a declared family-first name order.** With ``patronymic_rules`` opted in and ``Policy(name_order=FAMILY_FIRST)``, ``Мицкевич Адам Юзеф`` gave last ``Адам`` through 2.2.0 and now gives last ``Мицкевич`` -- the reading the declaration asks for -- and ``oglu Ahmad Vali Ali`` with Turkic handling gave last ``Ahmad`` and now ``oglu``. The rotations exist to restore the given-first reading a family-first listing hides, so under a declared family-first order the declaration decides. See the ``O1`` entry of ``docs/design/decisions.md`` (closes #384) - - **Fix CJK honorifics and names written with a full stop of any width.** ``HumanName("김민준 씨.").suffix`` is ``씨.`` (the fullwidth stop a Japanese or Chinese IME produces by default), with last ``김`` and first ``민준``, where every release since 2.1.0 gave first ``김``, middle ``민준``, last ``씨.`` and no suffix; the ideographic ``씨。`` and halfwidth ``씨。`` spellings move the same way, and so does a decomposed (NFD) ``김민준, 씨.`` from macOS-origin data -- the vocabulary lookup now composes NFC, so a decomposed ``Señor`` or ``née`` is recognized too. A period glued to a name word no longer breaks the name: ``양. 지훈`` gives last ``양.``, first ``지훈`` where 2.1.0 through 2.2.0 gave first ``양.``, middle ``지``, last ``훈``; ``김. 민준`` moves the same way; ``양 지훈.`` keeps the family-first order; ``김민준씨.`` and ``田中さん.`` peel their honorific as the stop-less spellings do; and a lone ``田中.`` is the family name where it read as a title. The period stays on the word it was written with -- nothing is rewritten. A Latin name written with ASCII periods is untouched: ``Smith. John`` still reads title ``Smith.``; a Latin or Cyrillic word wearing one of the three wider stops now reaches the vocabulary too, so ``Dr。 John Smith`` reads title ``Dr。`` where it read first ``Dr。``, and ``Smith, John V。`` reads suffix ``V。`` where it read middle ``V。`` -- the wide stop is not the shape the Latin initial veto reads, so the roman numeral wins where the ASCII ``V.`` stays a middle initial. This retires the 2.2.0 note below that read *period* strictly: the fullwidth ``김민준 씨.`` it named as unrecognized is recognized now. Seventeen names in the differential corpora read differently for this rule, every one of them a period-marked CJK form the corpora hold on the radar tier. One incompatibility, by decision: a ``Lexicon`` pickled by 2.1.x or 2.2.x that carries a caller-added entry the widened fold now changes -- a non-ASCII entry written with a fullwidth or ideographic stop, or in NFD -- no longer loads (``ValueError: incompatible Lexicon pickle: entries are not normalized``); the shipped vocabulary is unaffected, and the remedy is to rebuild the ``Lexicon`` from its source rather than unpickle it. See the ``cjk-full-stops`` entry of ``docs/design/decisions.md`` (closes #322, closes #323) + - **Fix CJK honorifics and names written with a full stop of any width.** ``HumanName("김민준 씨.").suffix`` is ``씨.`` (the fullwidth stop a Japanese or Chinese IME produces by default), with last ``김`` and first ``민준``, where every release since 2.1.0 gave first ``김``, middle ``민준``, last ``씨.`` and no suffix; the ideographic ``씨。`` and halfwidth ``씨。`` spellings move the same way. The vocabulary lookup now composes NFC, so a decomposed ``Señor`` or ``née`` from macOS-origin data is recognized too. A period glued to a name word no longer breaks the name: ``양. 지훈`` gives last ``양.``, first ``지훈`` where 2.1.0 through 2.2.0 gave first ``양.``, middle ``지``, last ``훈``, and ``김민준씨.`` peels its honorific as the stop-less spelling does. The period stays on the word it was written with; nothing is rewritten. A Latin name written with ASCII periods is untouched (``Smith. John`` still reads title ``Smith.``), but a Latin or Cyrillic word wearing one of the three wider stops now reaches the vocabulary: ``Dr。 John Smith`` reads title ``Dr。`` where it read first ``Dr。``. This retires the 2.2.0 note below that read *period* strictly. One incompatibility, by decision: a ``Lexicon`` pickled by 2.1.x or 2.2.x that carries a caller-added entry the widened fold now changes -- a non-ASCII entry written with a fullwidth or ideographic stop, or in NFD -- no longer loads (``ValueError: incompatible Lexicon pickle: entries are not normalized``); the shipped vocabulary is unaffected, and the remedy is to rebuild the ``Lexicon`` from its source rather than unpickle it. See the ``cjk-full-stops`` entry of ``docs/design/decisions.md`` (closes #322, closes #323) **Additions** - **Add the renunciate titles to the given-name title list, so a renunciate's one name is a given name.** ``HumanName("Swami Vivekananda")`` gives first ``Vivekananda`` with an empty last name, where every release since 1.4.0 gave last ``Vivekananda``; ``Guru Nanak``, ``Baba Ramdev`` and ``Lama Zopa`` move the same way, and so do the Devanagari and Bengali spellings added below. Two name words behind the title are unchanged -- ``Swami Vivekananda Saraswati`` keeps last ``Saraswati`` -- and a surname-retaining title is untouched: ``Rabbi Cohen`` still gives last ``Cohen``. ``venerable`` is deliberately not in the list, the traditions using it splitting on whether the family name survives. See the ``indic-honorifics`` entry of ``docs/design/decisions.md`` (closes #346) - - **Add prince and princess to the given-name title list, so a royal's one name is a given name.** ``HumanName("Prince Harry")`` gives first ``Harry`` with an empty last name, where every release since 1.4.0 gave last ``Harry``; ``Princess Anne``, ``Her Royal Highness Princess Anne`` and ``Prince Charles`` move the same way, and ``Prince abdul Rahman`` gives first ``abdul Rahman`` as ``Sir abdul Rahman`` does. Two name words behind the title are unchanged -- ``Prince Harry Windsor`` keeps last ``Windsor`` -- and ``Prince of Wales Jr`` is unchanged, the joined title addressing by its last word. ``lord`` and ``lady`` are deliberately NOT in the list: ``Lord Byron`` still gives last ``Byron``, and ``Lady Gaga`` last ``Gaga``, because each addresses by given name only as a courtesy style for children of the senior ranks (``Lord Peter``, ``Lady Diana``) and by title or surname for every peer and every wife (``Lord Byron``, ``Lady Thatcher``), which the list cannot express. The given-name collision is untouched: ``Prince Fielder`` now gives first ``Fielder`` (#348). See the ``H1`` entry of ``docs/design/decisions.md`` (closes #519) + - **Add prince and princess to the given-name title list, so a royal's one name is a given name.** ``HumanName("Prince Harry")`` gives first ``Harry`` with an empty last name, where every release since 1.4.0 gave last ``Harry``; ``Princess Anne`` and ``Her Royal Highness Princess Anne`` move the same way. Two name words behind the title are unchanged -- ``Prince Harry Windsor`` keeps last ``Windsor`` -- and so is ``Prince of Wales Jr``, the joined title addressing by its last word. ``lord`` and ``lady`` are deliberately NOT in the list: they address by given name only as a courtesy style (``Lord Peter``, ``Lady Diana``) and by surname for every peer and every wife (``Lord Byron``, ``Lady Thatcher``), which the list cannot express. The given-name collision is untouched: ``Prince Fielder`` now gives first ``Fielder`` (#348). See the ``H1`` entry of ``docs/design/decisions.md`` (closes #519) - **Add trailing honorifics as post-nominal vocabulary:** Latin ``rinpoche``, Devanagari ``जी``, ``साहब``, ``साहिब``, ``साहेब``, ``महाराज``, and Bengali ``সাহেব``, ``বাবু``, ``মহারাজ``. ``HumanName("Lama Zopa Rinpoche")`` reads title ``Lama``, first ``Zopa``, suffix ``Rinpoche``; ``नरेन्द्र मोदी जी`` reads last ``मोदी``, suffix ``जी``. They are recognized SPACED only and are deliberately absent from ``Lexicon.honorific_tails``: Banerjee, Mukherjee and Chatterjee end in the ``जी`` substring (``बनर्जी``, ``मुखर्जी``, ``चटर्जी``), so a glued peel would cut a real family name in two, and ``गांधीजी`` staying unpeeled is the accepted cost. Bengali ``বাবু`` is trailing where Devanagari ``बाबू`` is a leading title (#344, #343) @@ -42,9 +42,14 @@ Release Log - **Add Bengali honorifics -- the first Bengali vocabulary in the default lexicon (#343):** ``ড``, ``ডঃ``, ``ডক্টর``, ``ডাঃ``, ``ডা``, ``ডাক্তার``, ``শ্রী``, ``শ্রীমতী``, ``জনাব``, ``অধ্যাপক``, ``প্রফেসর``, ``বিচারপতি``, ``মাওলানা``, ``মুফতি``, ``আলহাজ্ব``, ``আলহাজ``, ``মিঃ``, ``মি``, ``মিসেস``, ``মোঃ``, ``মো``, ``মোসাঃ``, ``মোসা``, ``মোছাঃ`` and ``মোছা`` as titles, and ``স্বামী``, ``শ্রীল``, ``গুরু``, ``বাবা`` as given-name titles. ``ড. মুহাম্মদ ইউনূস`` reads title ``ড.``, first ``মুহাম্মদ``, last ``ইউনূস`` -- the vocabulary beats the initial reading -- while real initials are untouched: ``র. কে. নারায়ণ`` is unchanged. ``মোঃ আবদুল করিম`` reads title ``মোঃ``, first ``আবদুল``, last ``করিম``, the mirror of Latin ``Md``; the visarga spelling and the ``মো.`` period spelling both match, and the women's ``মোসাঃ``/``মোসা.`` rides the same pair of entries. ``ঠাকুর`` stays out, being Tagore. Latin transliterations (``Sri``, ``Pandit``, ``Mst``) are not added -- they collide with real given names where the native scripts cannot -- and belong to the opt-in packs of #345 (closes #343) - - **Add AmbiguityKind.GIVEN_OR_FAMILY, reported when a name of one name word had nothing to decide which field it is:** ``parse("Andrew")`` still gives given ``Andrew`` and now says that field was a convention rather than a reading -- one word gives the positional rule nothing to compare, so the library picks the given name under the default order and the family name under a declared family-first one, and ``detail`` names the field it picked. A trailing suffix does not decide it either: ``parse("Smith Jr.")`` reports it too, the suffix being peeled and the convention placing the one name word left. A name something DID decide stays silent -- ``"Dr. Smith"``, ``"Smith née Jones"``, ``"'Smitty' Jones"`` and ``"Smith, Andrew"`` -- and so do ``"abdul"`` and ``"de"``, where the bound given-name and particle vocabularies claimed the word, and ``"J."``, claimed by the initial's own shape. A name whose script settles the order is silent too: ``"毛泽东"`` reads family by convention of the writing system, not of this rule. Twenty-seven names in the differential corpora gain the report, and no field moves anywhere. See the ``O5`` entry of ``docs/design/decisions.md`` (closes #449) + - **Add AmbiguityKind.GIVEN_OR_FAMILY, reported when a name of one name word had nothing to decide which field it is:** ``parse("Andrew")`` still gives given ``Andrew`` and now says that field was a convention rather than a reading -- the library picks the given name under the default order and the family name under a declared family-first one, and ``detail`` names the field it picked. ``parse("Smith Jr.")`` reports it too, the suffix being peeled first. A name something DID decide stays silent -- ``"Dr. Smith"``, ``"Smith, Andrew"``, ``"abdul"`` (bound given-name vocabulary), ``"J."`` (an initial's shape) -- and so does ``"毛泽东"``, where the writing system settles the order. No field moves anywhere. See the ``O5`` entry of ``docs/design/decisions.md`` (closes #449) + + - **Add AmbiguityKind.TITLE_OR_NAME, reported when an input that is nothing but honorifics had its last word read as the name:** ``parse("Lord Chancellor")`` still gives title ``Lord``, family ``Chancellor``, and now says so -- this is a name parser, not a title parser, so handed a string with no name in it, it reads the last title word as one. ``"His Holiness"`` and ``"Dr. King"`` move the same way, ``king`` being title vocabulary. A title with an ordinary word behind it is silent (``"Dr. Smith"``, ``"King Charles"``), and so is a lone title word: ``parse("Dr.")`` is a title with no name beside it. The same convention on the post-nominal vocabulary reports the existing suffix-or-name: ``parse("Rinpoche")`` gives given ``Rinpoche`` and flags it, as does ``"QC MP"``. No field moves for this change; ``Dr King Jr``, ``Dr. King MD`` and ``Dr Jr`` gain the report from the title-peel fix above. See the ``H4`` entry of ``docs/design/decisions.md`` (closes #491) + + **Documentation** + + - **Document the family-first and East Asian input shapes beside the three Latin ones.** The input-shapes list in ``usage.rst`` grows from three forms to seven: forms 4 and 5 for a declared ``FAMILY_FIRST`` or ``FAMILY_FIRST_GIVEN_LAST`` order, and forms 6 and 7 for the native East Asian arrangements the script carries on its own. A comma or a Latin wrapper around a CJK name is named as tolerated input -- parsed best-effort, its handling changeable without notice -- and the ``customize.rst`` correspondence between forms 2 and 4 is written out. Behavior is unchanged (#469) - - **Add AmbiguityKind.TITLE_OR_NAME, reported when an input that is nothing but honorifics had its last word read as the name:** ``parse("Lord Chancellor")`` still gives title ``Lord``, family ``Chancellor``, and now says so -- this is a name parser, not a title parser, so handed a string with no name in it, it reads the last title word as one. ``"The Right Hon. the President of the Queen's Bench Division"`` and ``"His Holiness"`` move the same way, and so does ``"Dr. King"``, ``king`` being title vocabulary for the addressing forms. A title with an ordinary word behind it is silent (``"Dr. Smith"``, ``"King Charles"``), and so is a lone title word: ``parse("Dr.")`` is a title with no name beside it, the peel having taken the whole string, so no word was left standing to be read as a name and nothing was chosen. The same convention on the other vocabulary reports the existing suffix-or-name -- ``parse("Rinpoche")`` gives given ``Rinpoche`` and flags it, as does ``"QC MP"`` -- while ``"Jr."`` is unchanged and unflagged, reading as a title on its shape. The same doubt inside a joined unit reports too -- ``parse("Attorney General of Minnesota")`` reads title ``Attorney``, family ``General of Minnesota``, and whether ``General`` is a title is the fork. Eight names in the differential corpora gain a report from this change -- ten rows, two of those names sitting in two corpus files each -- and no field moves. Three more corpus names join the kind later in this cycle, from the title-peel fix above -- ``Dr King Jr``, ``Dr. King MD`` and ``Dr Jr``, where the run now leaves a title-vocabulary word standing; those move fields, for the reasons that bullet gives. ``Sir Jr`` reads the same way -- first ``Sir``, suffix ``Jr``, the same report -- but no differential corpus holds it, so it is ``Dr Jr``'s row that pins the reading for both. See the ``H4`` entry of ``docs/design/decisions.md`` (closes #491) * 2.2.0 - August 31, 2026 diff --git a/docs/usage.rst b/docs/usage.rst index 55262267..c7c11c18 100644 --- a/docs/usage.rst +++ b/docs/usage.rst @@ -331,6 +331,10 @@ syllable held as its separate jamo rather than as one codepoint. macOS filenames are the common source. Everything above works on decomposed input: script classification normalizes to NFC before deciding, so a decomposed name gets the same order rule as its composed twin. +Vocabulary lookup does the same before matching a word against +titles, honorifics and the rest, so a decomposed ``Señor`` or ``née`` +— macOS-origin data again — is recognized as readily as its composed +spelling. Splitting is the exception. An unspaced decomposed hangul name is ordered correctly but not split, because surname matching runs against @@ -393,6 +397,19 @@ without it. That is why ``김민준씨`` still divides into family 김 and given 민준, and why a configured Japanese segmenter is handed 山田太郎 rather than 山田太郎様. +The stop can be any width: the fullwidth ``.`` a Japanese or +Chinese input method produces by default, the ideographic ``。`` and +the halfwidth ``。`` all reach the honorific vocabulary as an ASCII +period does, so ``김민준 씨.`` gives suffix ``씨.``, family 김, given +민준. A period glued to an ordinary name word, not an honorific, is +likewise left where it was written rather than breaking the +segmentation that follows it: + +.. doctest:: + + >>> parse("양. 지훈").family, parse("양. 지훈").given + ('양.', '지훈') + Commas and Latin wrappers around a CJK name ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ diff --git a/nameparser/_pipeline/_classify.py b/nameparser/_pipeline/_classify.py index 4aa73964..b1eb885b 100644 --- a/nameparser/_pipeline/_classify.py +++ b/nameparser/_pipeline/_classify.py @@ -4,9 +4,11 @@ the structural-boundary test the marker pass applies -- see _tag_marker_runs). Produces: tokens with vocabulary tags added (text/span/role unchanged). -Reads: every Lexicon vocabulary field; no Policy FIELD is consulted -(is_initial does consult the _policy module's _NO_INITIALS constant, -which is not configuration -- nothing here varies by Policy value). +Reads: every Lexicon vocabulary field except surnames and +honorific_tails, which script_segment consumes upstream; no Policy +FIELD is consulted (is_initial does consult the _policy module's +_NO_INITIALS constant, which is not configuration -- nothing here +varies by Policy value). Tags emitted -- stable (API): "particle", "conjunction", "initial"; namespaced (unstable): "vocab:title", "vocab:given-title", diff --git a/nameparser/_policy.py b/nameparser/_policy.py index 8700d8df..6f4fbf3f 100644 --- a/nameparser/_policy.py +++ b/nameparser/_policy.py @@ -607,7 +607,10 @@ class Policy: #: disable one. segment_scripts: frozenset[Script] = frozenset({Script.HANGUL}) #: Opt-in detectors that reorder patronymic-shaped names - #: (EAST_SLAVIC, TURKIC); usually set via a locale pack. + #: (EAST_SLAVIC, TURKIC); usually set via a locale pack. A rotation + #: restores the given-first reading a family-first listing hides, so + #: under a declared FAMILY_FIRST or FAMILY_FIRST_GIVEN_LAST name_order + #: it stands down and the declaration decides. patronymic_rules: frozenset[PatronymicRule] = frozenset() #: Folds middle into family instead of splitting them (v1's #: middle_name_as_last) -- for data where unrecognized interior