Aglet

Triage Inconsistent Identifiers Across Records

Identifier inconsistency may be a formatting difference, a source-specific key, a reused value, or a true entity conflict. Triage should list the fields involved, show representative variants, and map the joins or reports that depend on them. Bound the issue before selecting a canonicalization rule.

Establish what is happening

  1. Inventory key variants

    Collect observed identifier values with their source, field name, type, length, casing, whitespace, separators, and first-seen time. Group obvious representations together, but leave ambiguous values separate. Preserve examples where two records look similar yet may identify different entities.

  2. Map dependent joins

    Trace each identifier into joins, uniqueness constraints, URLs, exports, and aggregate queries. Check whether variants create missed matches, false matches, one-to-many expansions, or silent overwrites. Record the first downstream result that changes when a variant is normalized or left untouched.

  3. Bound the population

    Count affected records by source, identifier family, date range, and processing revision. Compare current and historical distributions to find an introduction boundary. Mark records whose identity cannot be inferred safely, since they need a separate review path from formatting-only variants.

What to carry forward

Deliver a scope brief listing identifier families, observed variants, dependent joins, affected cohorts, and ambiguous cases. Triage is complete when a reviewer can reproduce the grouping and see which records require investigation. Do not declare a canonical value until identity evidence supports it.

Keep the decision with the work.

Use a Work Item in Aglet to record the problem, the evidence you have, and the next decision. Add an owner and priority, then keep updates in the discussion so the next person can follow the reasoning.

Create an account See the product workflow