feat(policy-ocr): suggest the customer from the printed insured name
The office books customers surname-first ("WAGONER, PAMELA") and carriers
print them given-name-first ("PAMELA DENISE WAGONER"), so the review screen
made staff retype a name the parser had already read. Comparing normalized
token sets makes the two orderings the same thing.
Only on the zero-hit path, where the policy number found nothing and a human
has to pick a customer anyway. The suggestions are written to a new
`customerSuggestions` column rather than `matchCandidates`, which the review
screen reads as policy-number hits, and they never set `matchedCustomerId` or
`confident` — matching on `Policy.policyNumber` is unchanged.
Two tiers, drawn where the real book has cliffs: EXACT (identical token sets)
and PARTIAL (containment, >=2 shared tokens, surname present). Of 1536
customers, 1487 have a distinct token set, so EXACT cross-person collisions
are ~0; loosen to surname + first given name and 131 (8.5%) collide, and 185
surnames are shared by 524 customers, which is why one token is never enough
and the surname must be printed explicitly. Replaying every book row as a
carrier would print it: 97.9% top-ranked correct, 1.2% a different row, all
but two of those the same human on a duplicate or variant row.
Normalization folds accents (OCR's MUNOZ reaches the book's MUÑOZ), drops
initials, Spanish particles, JR/S.A. DE C.V., and any token with a digit —
ANA prints the phone hard against the name as `Ph.3102001538`. Names over 8
tokens or 80 characters are refused outright, because GMX's especificación
has no field labels and the parser has handed its whole first page over as
`insuredName`.
Not used for utility statements: there the registrant genuinely is not the
customer, so the same trick would be wrong rather than noisy.
Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
This commit is contained in:
+7
@@ -0,0 +1,7 @@
|
||||
-- Ranked customers whose name matches the printed insured name, for the
|
||||
-- documents whose policy number found nothing and therefore need a customer
|
||||
-- picked by hand. Kept in its own column rather than folded into
|
||||
-- `matchCandidates`, which the review screen reads as policy-number hits —
|
||||
-- a name is a suggestion and must never be able to masquerade as a match.
|
||||
ALTER TABLE `policy_ocr_documents`
|
||||
ADD COLUMN `customerSuggestions` JSON NULL;
|
||||
@@ -476,6 +476,13 @@ model PolicyOcrDocument {
|
||||
/// normal; >1 means the policy number is shared across customers and a
|
||||
/// human must pick.
|
||||
matchCandidates Json?
|
||||
/// `CustomerNameSuggestion[]` — customers whose name matches the printed
|
||||
/// insured name, ranked. A SUGGESTION, never a match: it is deliberately
|
||||
/// kept out of `matchCandidates` so the review screen cannot mistake a
|
||||
/// name hint for a policy-number hit, and it never sets
|
||||
/// `matchedCustomerId`. Only populated when the policy number found
|
||||
/// nothing, which is exactly when staff have to pick a customer by hand.
|
||||
customerSuggestions Json?
|
||||
/// Text, not VARCHAR(191): this carries the parser's whole note trail, and
|
||||
/// a multi-section ANA policy runs past 191 characters routinely. Silently
|
||||
/// truncating it drops the tail notes, which are the ones that say what
|
||||
|
||||
Reference in New Issue
Block a user