Skip to content

Candidate: CPE Name Matching (NIST IR 7696)

NIST Interagency Report 7696, Common Platform Enumeration: Name Matching Specification 2.3, defines a procedure that compares a source well-formed name (WFN) with a target WFN and returns one of:

  • EQUAL
  • SUBSET
  • SUPERSET
  • DISJOINT

The name relation is aggregated from per-attribute set relations by a documented rule:

  • any attribute DISJOINT makes the name relation DISJOINT;
  • all attributes EQUAL makes it EQUAL;
  • all SUBSET-or-EQUAL makes it SUBSET; and
  • all SUPERSET-or-EQUAL makes it SUPERSET.

The matching relation is directional: wildcards are permitted in the source and undefined in the target.

NIST also provides a reference implementation.

The Rule

The raw CPE name-matching rule is not an equivalence relation.

It compares the sets of products denoted by two CPE names and reports one of four set-theoretic relationships:

EQUAL, SUBSET, SUPERSET, or DISJOINT.

SUBSET and SUPERSET capture directional containment; EQUAL and DISJOINT are symmetric.

Therefore the raw CPE name-matching result must not itself be treated as an equivalence partition over names.

Modeling CPE names directly as equivalence classes would force a directional containment relation into a symmetric identity relation.

That translation is INVALID.

This resembles the modeling error exposed by the Grype feasibility case, but it is not the same failure mode.

The Grype lifting produced a single binary classification (covered / not-covered), whose two-block partition offered little structure beyond labeled set comparison.

CPE instead supplies a directional, multi-valued matching relation aggregated over multiple attributes.

The candidate partition therefore arises only at a derived level:

Assets are grouped when they have identical match signatures over an externally fixed family of CPE criteria.

Overview

CVE
v
NVD applicability statement
v
fixed CPE criteria K
v
assets
v
normative match signatures
v
scanner match signatures


raw CPE relation
    = directional
    = not partitioned

Study A
    = NIST IR 7696 criterion-match signatures
    = derived equivalence
    = equality / mutual-refinement conformance

Study B
    = NIST IR 7696 + NIST IR 7698 applicability composition
    = separately admitted translation

structural surplus
    = independent empirical question

Two Possible Studies

A. CPE criterion-match profile equivalence sig(a) = vector of individual criterion matches Normative basis: NIST IR 7696.

B. CVE applicability-profile equivalence sig(a) = vector/results of complete applicability expressions Normative basis: NIST IR 7696 + NIST IR 7698.

A is cleaner mathematically and directly grounded in NIST IR 7696. B is closer to the security consequence but requires CPE Applicability Language/NVD configuration semantics also.

Study B is groundable, not merely harder. NIST IR 7698, CPE Applicability Language 2.3, defines the composition of CPE name-matching results into higher-level applicability statements. NVD configurations instantiate related applicability logic over CPE match criteria.

A and B are therefore two composable translations: Study A audits the base criterion-matching layer; Study B audits the applicability logic built on top of it. Each requires its own translation-admissibility check.

Admissible Construction

Do not partition the CPE names. Partition the assets by their criterion-match signature under the matching rule.

Fix a family of CPE criteria,
e.g., all leaf criteria from a preregistered set of NVD CVE configurations:
    K = { K1, K2, ... , Km }

Asset universe:
    A = { a1, a2, ... , an }   (each asset carries a CPE)

Match signature of asset a:
    sig(a) = ( m1(a), m2(a), ... , mm(a) )
    where mi(a) = "does Ki match a under NIST IR 7696"

Normative criterion-match signature:
    sig_NIST(a) is computed according to the normative CPE 2.3
    matching semantics of NIST IR 7696.
    The NIST-maintained reference implementation is used as an executable
    oracle for that specification and independently checked against selected
    normative examples.

Declared partition pi_declared:
    group assets by equality of sig_NIST(a)

Implemented criterion-match signature:
    sig_S(a) is computed from scanner S's decisions over the same fixed K

Implemented partition pi_implemented:
    group assets by equality of sig_S(a)

Audit:
    compare pi_declared and pi_implemented in both refinement directions.
    Equality is the conformance target:
    pi_declared refines pi_implemented and
    pi_implemented refines pi_declared.

Witness = an asset pair that the normative and implemented partitions
classify differently.

Report witnesses by direction:

- split witness:
  the normative partition co-classifies the pair,
  but the implemented partition separates it;

- merge witness:
  the normative partition separates the pair,
  but the implemented partition co-classifies it.

Either witness establishes divergence in criterion-match profile.
Whether that divergence establishes a vulnerability-applicability defect
depends on the higher-level applicability semantics being audited.

The directional containment relation is recorded honestly in the translation manifest as relation_kind = directional, and the equivalence lives only at the derived asset-signature level, where it genuinely is symmetric and transitive.

The derived signature equality is a genuine equivalence relation and therefore induces a partition.

Relationships among CPE criteria may constrain which signatures are possible, but whether those constraints yield useful refinement structure between declared and implemented signature partitions is an empirical question for the feasibility test.

Illustration

Three criteria that cross-cut, over six assets:

pi_declared (per NIST IR 7696):
    { a1, a2 }     sig = {K1}
    { a3, a4 }     sig = {K1, K2}
    { a5, a6 }     sig = {K2, K3}

A scanner that spuriously matches K2 to a2 yields:
pi_implemented:
    { a1 }
    { a2, a3, a4 }
    { a5, a6 }

These two partitions are incomparable:
neither refines the other.

Split witness:
    (a1, a2)

The normative partition groups a1 and a2 together,
but the implementation separates them.

This is the direction detected by one-directional
declared-to-implemented faithfulness.

Merge witness:
    (a2, a3)

The normative partition separates a2 and a3,
but the implementation groups them together.
This is not detected by the one-directional faithfulness test;
it appears under the reverse refinement direction.

For applicability conformance, both directions matter.
The audit target is therefore equality of
pi_declared and pi_implemented,
equivalently mutual refinement in both directions.
Report split witnesses and merge witnesses separately.

Admissibility Checklist

  • SOURCE: PASS. NIST IR 7696 supplies the normative matching semantics. The NIST-maintained reference implementation provides an executable oracle for those semantics, subject to checks against normative examples.
  • DERIVED EQUIVALENCE: PLAUSIBLE. For a fixed externally grounded criteria family K, the normative matching rule determines a match signature for each asset. Equality of complete signatures is symmetric, transitive, and reflexive, and therefore induces an equivalence relation over assets. This is criterion-match profile equivalence. It must not yet be described as full vulnerability-applicability identity, because complete CVE applicability may depend on higher-level logical composition beyond individual criterion matches.
  • SYMMETRY: PASS ONLY AT THE DERIVED LEVEL. The raw CPE name-matching relation is directional and must remain relation_kind = directional; forcing it into a symmetric relation is INVALID. Equality of complete asset match signatures is symmetric and may legitimately induce the derived partition. The names themselves are never partitioned by the directional match relation.
  • TRANSITIVITY: PASS AT THE DERIVED LEVEL. Equality of complete match signatures is transitive. No transitivity claim about the raw four-valued CPE matching result is required to construct the derived equivalence.
  • CONFORMANCE ORIENTATION: EQUALITY / MUTUAL REFINEMENT. For this applicability-oriented case, both splitting and merging can represent implementation defects. A split may correspond to a missed match or false negative. A merge may correspond to a spurious match or false positive. Therefore the conformance target is: pi_declared == pi_implemented equivalently: pi_declared refines pi_implemented and pi_implemented refines pi_declared. Split witnesses and merge witnesses must be reported separately.
  • MULTI-CLASS: PLAUSIBLE; MUST BE DEMONSTRATED. With m Boolean criterion-match decisions, at most 2^m distinct signatures are possible. The diagnostic passes this gate only if the externally fixed criteria family K and asset universe actually induce at least three nonempty signature classes. The criteria must not be selected to manufacture that result.
  • SECURITY RELEVANCE: PASS, core. CPE matching is part of the standardized machinery used to connect platform descriptions to applicability statements. NVD CVE configurations use CPE match criteria to describe the products/platforms to which a vulnerability applies.
  • IMPLEMENTATION: PASS. NIST IR 7696 supplies the normative matching semantics. The NIST-maintained reference implementation provides an executable oracle for those semantics, subject to checks against normative examples. A CPE-consuming scanner supplies the implementation behavior under audit. Both executable sides are observable.
  • BASELINE: LIVE RISK. Diffing a scanner's per-criterion decisions against the reference implementation is a strong conventional baseline. The surplus must be the cross-criterion partition structure that a per-criterion diff does not assemble automatically.

Candidate added value includes:

  • classifying the overall relation as equal, refinement, reverse refinement, or incomparable;
  • distinguishing split witnesses from merge witnesses;
  • showing that one criterion error simultaneously splits one normative class and crosses another normative boundary; and
  • deriving a repair obligation at the class level rather than merely identifying an incorrect individual match decision. Restating the spec relation does not count.

Additional guard, not on the original checklist but load-bearing here:

  • CRITERIA-FAMILY PRINCIPLE: the family K must be source-grounded or representatively sampled, not chosen to manufacture cross-cutting structure. Choosing K to produce a nice partition is the same circularity that using observed scanner output to pick the relation would be. Fix K from an external source (preferably one CVE configuration or a preregistered NVD slice) before observing any scanner behavior.
  • LOGICAL-COMPOSITION BOUNDARY: Study A audits individual CPE criterion-match profiles. It must not silently flatten a complete CVE applicability expression and then claim to have audited vulnerability applicability. If AND, OR, negation, or other configuration semantics determine the security decision, those semantics belong to Study B and require their own source-grounded translation.

Verdict

CPE is currently the strongest candidate for the second feasibility test.

SOURCE, STUDY A: PASS, NIST IR 7696
SOURCE, STUDY B: PLAUSIBLE, NIST IR 7696 + NIST IR 7698
SECURITY RELEVANCE: PASS
IMPLEMENTATION ORACLE: PASS
DIRECTIONAL SEMANTICS: PASS, explicitly preserved
DERIVED EQUIVALENCE: PLAUSIBLE
CONFORMANCE ORIENTATION: EQUALITY / MUTUAL REFINEMENT
MULTI-CLASS: PLAUSIBLE / must be demonstrated on fixed K
STRUCTURAL ANALYTICAL SURPLUS: ORGANIZATIONAL ONLY; judged against pre-registred RUBRIC
BASELINE ADVANTAGE: STRONG LIVE RISK

No informational-surplus claim is made: all partition outputs are computable from the same complete evidence available to the strongest conventional baseline. The open question is whether the structural representation materially improves diagnostic compression, localization, unification, or repair guidance.

SYMMETRY passes only through the asset-signature construction, and that constraint is a feature: getting it wrong is itself a Layer-1 INVALID, so the case exercises the translation-admissibility layer honestly.

It is not the gentle, natively partition-shaped case: that is PURL canonicalization, which is a clean single-level equivalence but is the candidate most exposed to the BASELINE objection, because canonicalization-equivalence is exactly what a PURL-normalizing semantic diff already computes, and it audits a shallow lexical/representational dimension.

CPE reaches closer to the actual security decision by auditing criterion-match profiles used in applicability processing. It also has the potential to induce richer multi-class structure, at the cost of two design obligations: the criteria family must be fixed independently, and criterion-match equivalence must not be overstated as full vulnerability-applicability identity. The strongest live threat remains the reference-implementation baseline.

Positioned against the others:

  • PURL canonicalization = the floor test. Does the partition machinery run natively on a clean equivalence at all.
  • CPE match-signature = the criterion-matching test. Does the machinery add surplus on the standardized matching relation used as an input to vulnerability-applicability processing, where the baseline is a real matcher and the normative grounding is strong.

They are complementary, not competing. If a single diagnostic case has to carry the go/no-go, CPE is currently the strongest candidate for forcing the added-value question into the open. Its normative grounding removes much of the source-underdetermination risk, and a fixed criteria family gives the method a genuine opportunity to produce more than two nonempty classes. A negative result would therefore bear more directly on structural surplus than the Grype binary-coverage case did.

Next Task

Begin with Study A. Fix a small criteria family K from one external source (e.g., the CPE match criteria associated with one preregistered CVE configuration) before looking at any scanner.

Use NIST IR 7696 only for the first diagnostic. Do not interpret the resulting criterion-match partition as full CVE applicability.

If Study A is ADMISSIBLE and the derived construction is sound, proceed to Study B using NIST IR 7698 and the corresponding higher-level applicability semantics. Study B does not require Study A to demonstrate structural surplus.

Take a set of assets with known CPEs. Compute pi_declared with the NIST reference implementation and pi_implemented with one real scanner. Check three things:

  1. does pi_declared have more than two classes (multi-class holds);
  2. do the two partitions actually diverge (there is something to audit);
  3. does the partition comparison produce a materially better analytical product from the same evidence than the per-criterion reference-implementation diff, such as mutual-refinement classification, split-versus-merge distinction, incomparability, or a class-level repair obligation.

Admissibility must be decided independently from structural surplus. If the source semantics, criteria-family construction, and derived signature equivalence pass the admissibility checks, the translation is ADMISSIBLE whether or not the scanner partitions diverge.

If (1) holds, the case is genuinely multi-class.

If (2) also holds, there is a conformance divergence to analyze.

If (3) also holds, the case provides evidence of partition-native analytical surplus relative to the conventional baseline.

If (3) fails, the case may still be ADMISSIBLE, but the partition analysis has not demonstrated added value over the baseline: a clean, early boundary result, which is exactly what the diagnostic is for.

Structural-Surplus Rubric

Input parity:

SE and the conventional baseline receive the same complete evidence.

No credit for:

  • detecting a mismatch already visible in the baseline;
  • recomputing a value already present in the baseline;
  • renaming a conventional mismatch;
  • aggregating results without reducing diagnostic burden.

Candidate surplus requires at least one preregistered analytical product that materially reduces diagnostic burden or improves fault localization or repair guidance relative to the strongest conventional baseline:

  1. Compression A finite witness set or structural summary replaces multiple independent mismatch reports while preserving the relevant fault.

  2. Population-level localization The analysis identifies a class-level pattern such as systematic splitting, merging, refinement, reverse refinement, or incomparability.

  3. Repair obligation The structural result yields a general correction condition applying to a class of cases rather than one observed mismatch.

  4. Cross-case unification Multiple individual discrepancies are shown to arise from one common structural failure.

  5. Boundary diagnosis The method distinguishes cases where partition analysis contributes no additional analytical value and returns NO_STRUCTURAL_SURPLUS.

Meaningful Compression

Compression is credited only when the structural output represents multiple independently reported baseline discrepancies through one shared diagnosis or repair condition.

Conceptually:

baseline: N independent mismatch records

SE: k witnesses + one shared structural diagnosis

Meaningful compression requires k < N and, more importantly, that the reduced witness set preserves the fault pattern or repair obligation represented by the baseline mismatches. The quantitative threshold for crediting compression must be fixed before prospective execution.

Open preregistration item: The decision rule for crediting meaningful compression must be fixed before prospective execution. This is the highest-priority remaining rubric parameter because it is the principal subjective threshold that could otherwise be influenced by observed results.

Tentative Design

Study A

Purpose: admissibility and soundness control.

Expected:

  • ADMISSIBLE
  • SOUND
  • possibly NO_STRUCTURAL_SURPLUS

Interpretation: demonstrates the method can decline to claim added value

Question: Can we translate and audit the base relation correctly? Expected: yes, probably no surplus.

Study B

Purpose: Test whether population-level structural diagnosis improves the analytical product over the strongest conventional applicability-conformance baseline.

BASELINE B receives the same complete evidence as SE and reports:

  • per-criterion normative and implemented match results;
  • criterion-level conformance;
  • normative and implemented applicability results;
  • applicability-level conformance; and
  • error-layer classification, including:
  • criterion mismatch that changes applicability;
  • criterion mismatch with applicability preserved;
  • criteria conformant but applicability divergent; and
  • fully conformant.

SE receives the same evidence and additionally computes:

  • normative and implemented partitions;
  • mutual-refinement relationships;
  • split and merge witnesses;
  • population-level structural classification; and
  • candidate class-level repair obligations.

Outcome:

  • STRUCTURAL_SURPLUS, or
  • NO_STRUCTURAL_SURPLUS

according to the preregistered Structural-Surplus Rubric.

Question: Given identical complete evidence, does structural organization materially improve diagnosis? Expected: Unknown; actual surplus test.

Both: ADMISSIBILITY is independent of outcome. NO_STRUCTURAL_SURPLUS is a valid result.