Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
4 changes: 3 additions & 1 deletion docs/customize.rst
Original file line number Diff line number Diff line change
Expand Up @@ -321,7 +321,9 @@ listed below.
- ``frozenset[PatronymicRule]``
- Reorders patronymic-shaped names via opt-in detectors — East
Slavic formal order (``EAST_SLAVIC``) or Turkic reversed order
(``TURKIC``). Defaults to empty.
(``TURKIC``) — but stands down under a declared
``FAMILY_FIRST`` or ``FAMILY_FIRST_GIVEN_LAST`` ``name_order``.
Defaults to empty.
* - ``middle_as_family``
- ``bool``
- Folds ``middle`` into ``family`` instead of splitting them —
Expand Down
6 changes: 6 additions & 0 deletions docs/locales.rst
Original file line number Diff line number Diff line change
Expand Up @@ -180,6 +180,12 @@ new naming rule belongs in.
mixes traditions, parse the subsets separately with different
parsers rather than enabling a pack over all of it.

A declared family-first ``name_order`` stands down the rotation
instead of competing with it: fold the pack onto a base parser
built with ``Policy(name_order=FAMILY_FIRST)`` and ``"Мицкевич
Адам Юзеф"`` reads family ``Мицкевич`` rather than the given-first
order the pack restores by default.

.. _segmenter-contract:

Segmenters
Expand Down
31 changes: 18 additions & 13 deletions docs/release_log.rst

Large diffs are not rendered by default.

17 changes: 17 additions & 0 deletions docs/usage.rst
Original file line number Diff line number Diff line change
Expand Up @@ -331,6 +331,10 @@ syllable held as its separate jamo rather than as one codepoint. macOS
filenames are the common source. Everything above works on decomposed
input: script classification normalizes to NFC before deciding, so a
decomposed name gets the same order rule as its composed twin.
Vocabulary lookup does the same before matching a word against
titles, honorifics and the rest, so a decomposed ``Señor`` or ``née``
— macOS-origin data again — is recognized as readily as its composed
spelling.

Splitting is the exception. An unspaced decomposed hangul name is
ordered correctly but not split, because surname matching runs against
Expand Down Expand Up @@ -393,6 +397,19 @@ without it. That is why ``김민준씨`` still divides into family 김 and
given 민준, and why a configured Japanese segmenter is handed 山田太郎
rather than 山田太郎様.

The stop can be any width: the fullwidth ``.`` a Japanese or
Chinese input method produces by default, the ideographic ``。`` and
the halfwidth ``。`` all reach the honorific vocabulary as an ASCII
period does, so ``김민준 씨.`` gives suffix ``씨.``, family 김, given
민준. A period glued to an ordinary name word, not an honorific, is
likewise left where it was written rather than breaking the
segmentation that follows it:

.. doctest::

>>> parse("양. 지훈").family, parse("양. 지훈").given
('양.', '지훈')

Commas and Latin wrappers around a CJK name
^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^

Expand Down
8 changes: 5 additions & 3 deletions nameparser/_pipeline/_classify.py
Original file line number Diff line number Diff line change
Expand Up @@ -4,9 +4,11 @@
the structural-boundary test the marker pass applies -- see
_tag_marker_runs).
Produces: tokens with vocabulary tags added (text/span/role unchanged).
Reads: every Lexicon vocabulary field; no Policy FIELD is consulted
(is_initial does consult the _policy module's _NO_INITIALS constant,
which is not configuration -- nothing here varies by Policy value).
Reads: every Lexicon vocabulary field except surnames and
honorific_tails, which script_segment consumes upstream; no Policy
FIELD is consulted (is_initial does consult the _policy module's
_NO_INITIALS constant, which is not configuration -- nothing here
varies by Policy value).

Tags emitted -- stable (API): "particle", "conjunction", "initial";
namespaced (unstable): "vocab:title", "vocab:given-title",
Expand Down
5 changes: 4 additions & 1 deletion nameparser/_policy.py
Original file line number Diff line number Diff line change
Expand Up @@ -607,7 +607,10 @@ class Policy:
#: disable one.
segment_scripts: frozenset[Script] = frozenset({Script.HANGUL})
#: Opt-in detectors that reorder patronymic-shaped names
#: (EAST_SLAVIC, TURKIC); usually set via a locale pack.
#: (EAST_SLAVIC, TURKIC); usually set via a locale pack. A rotation
#: restores the given-first reading a family-first listing hides, so
#: under a declared FAMILY_FIRST or FAMILY_FIRST_GIVEN_LAST name_order
#: it stands down and the declaration decides.
patronymic_rules: frozenset[PatronymicRule] = frozenset()
#: Folds middle into family instead of splitting them (v1's
#: middle_name_as_last) -- for data where unrecognized interior
Expand Down