feat(policy-ocr): suggest the customer from the printed insured name
The office books customers surname-first ("WAGONER, PAMELA") and carriers
print them given-name-first ("PAMELA DENISE WAGONER"), so the review screen
made staff retype a name the parser had already read. Comparing normalized
token sets makes the two orderings the same thing.
Only on the zero-hit path, where the policy number found nothing and a human
has to pick a customer anyway. The suggestions are written to a new
`customerSuggestions` column rather than `matchCandidates`, which the review
screen reads as policy-number hits, and they never set `matchedCustomerId` or
`confident` — matching on `Policy.policyNumber` is unchanged.
Two tiers, drawn where the real book has cliffs: EXACT (identical token sets)
and PARTIAL (containment, >=2 shared tokens, surname present). Of 1536
customers, 1487 have a distinct token set, so EXACT cross-person collisions
are ~0; loosen to surname + first given name and 131 (8.5%) collide, and 185
surnames are shared by 524 customers, which is why one token is never enough
and the surname must be printed explicitly. Replaying every book row as a
carrier would print it: 97.9% top-ranked correct, 1.2% a different row, all
but two of those the same human on a duplicate or variant row.
Normalization folds accents (OCR's MUNOZ reaches the book's MUÑOZ), drops
initials, Spanish particles, JR/S.A. DE C.V., and any token with a digit —
ANA prints the phone hard against the name as `Ph.3102001538`. Names over 8
tokens or 80 characters are refused outright, because GMX's especificación
has no field labels and the parser has handed its whole first page over as
`insuredName`.
Not used for utility statements: there the registrant genuinely is not the
customer, so the same trick would be wrong rather than noisy.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
@@ -134,6 +134,55 @@ customers do occur (one group policy bound by two related parties), and
|
||||
picking arbitrarily would silently book the wrong coverage against the wrong
|
||||
person.
|
||||
|
||||
### Name suggestions on the zero-hit path
|
||||
|
||||
When the policy number finds nothing — the new-policy case, where a human has
|
||||
to pick a customer anyway — `name-matcher.ts` ranks the customer book against
|
||||
the printed insured name and the review screen offers the top three as
|
||||
one-click buttons above the picker. They are written to
|
||||
`policy_ocr_documents.customerSuggestions`, deliberately **not** to
|
||||
`matchCandidates`, so a name hint can never be read as a policy-number hit.
|
||||
Nothing sets `matchedCustomerId`; the rule above is unchanged.
|
||||
|
||||
The problem is only ordering: the office books customers surname-first
|
||||
(`WAGONER, PAMELA`) and carriers print them given-name-first
|
||||
(`PAMELA DENISE WAGONER`), so a string compare never hits while a **token-set**
|
||||
compare does. Names are normalized (accents folded, so OCR's `MUNOZ` reaches
|
||||
the book's `MUÑOZ`; initials, `DE`/`LA`/`Y`, `JR`, `S.A. DE C.V.` and any token
|
||||
containing a digit dropped — ANA prints the phone hard against the name as
|
||||
`Ph.3102001538`). Two tiers:
|
||||
|
||||
| Tier | Rule |
|
||||
|---|---|
|
||||
| `EXACT` | identical token sets, any order |
|
||||
| `PARTIAL` | one set contains the other, ≥2 shared tokens, **and** the surname is present |
|
||||
|
||||
Both thresholds come from measuring the real book (1536 customers):
|
||||
|
||||
- 1487 distinct token sets, so `EXACT` cross-person collisions are ~0
|
||||
- loosen to surname + first given name and 131 customers (8.5%) collide —
|
||||
the book holds `MCWILLIAMS, BRIAN MICHAEL` *and* `MCWILLIAMS, BRIAN`
|
||||
- 185 surnames are shared by 524 customers, so one token is never evidence;
|
||||
hence the ≥2 floor and the explicit surname requirement, which is what stops
|
||||
`JERRY MARILYN` reaching `ESTRADA, JERRY & MARILYN` on given names alone
|
||||
|
||||
Replaying every book row as a carrier would print it (given-name-first, joint
|
||||
spouse dropped): 97.9% top-ranked correct, 0.9% no suggestion, 1.2% a
|
||||
different row — and all but two of those are the same human on a duplicate or
|
||||
variant row (`MOLNAR, JANOS` vs `MOLNAR, JANOS`, `IBARRA, ISMAEL &`). The
|
||||
two genuine wrong-person cases are `CUADROS, JORGE JR` against three
|
||||
`CUADROS, JORGE H.`, and they appear as a tie in the list rather than as a
|
||||
single answer.
|
||||
|
||||
A blob is refused outright (>8 tokens or >80 characters): GMX's especificación
|
||||
has no field labels and the parser has been seen handing its whole first page
|
||||
over as `insuredName`, which would find a surname somewhere in the prose.
|
||||
`(SIN NOMBRE)` — 14 rows the migration left — is skipped on both sides.
|
||||
|
||||
**Not used for utility statements.** There the registrant genuinely is not the
|
||||
customer (the `CATT, RANDY` finding above), so the same trick would be wrong,
|
||||
not merely noisy.
|
||||
|
||||
## GMX ships two unrelated documents for the same policy
|
||||
|
||||
The office downloads both from the same portal, and either can land in a
|
||||
@@ -457,6 +506,17 @@ term, the excluded sections, the column-position split on the driver's policy,
|
||||
its different section order, and — for both faces — that a doubled or tripled
|
||||
input yields one set of coverages and one driver rather than one per copy.
|
||||
|
||||
`apps/api/src/policy-ocr/name-matcher.spec.ts` — 21 cases on the customer name
|
||||
suggestions, every fixture name lifted from the real book: the reversed name,
|
||||
the printed middle name, the exact row outranking the row that merely contains
|
||||
it, the joint account reached from one spouse (and refused when only given
|
||||
names are printed), the Spanish double surname with the comma in either place,
|
||||
the 54 rows with no comma at all, the `(SIN NOMBRE)` placeholder, a bare shared
|
||||
surname, and the page-sized blob. Four more in
|
||||
`policy-matcher.service.spec.ts` pin the wiring: suggestions on the zero-hit
|
||||
and unreadable-number paths, no book read at all when the policy number hits,
|
||||
and one book read across a batch.
|
||||
|
||||
Four of the GMX cases are regression tests for ways the parser can silently attach
|
||||
the *wrong* value rather than none — a neighbouring coverage's prose read as
|
||||
a deductible, the page-level `DEDUCIBLES:` paragraph read as one, a coverage
|
||||
|
||||
Reference in New Issue
Block a user