Skip to content

Project Plan: Structural Explainability Audit of Vulnerability-Matching Identity

Project Decision

Target the MSR 2027 Registered Reports track with a Stage 1 protocol submission by November 20, 2026; the abstract is due November 13, 2026.

The project instantiates the existing Structural Explainability operational-identity audit in one external cybersecurity domain: identity decisions made when SBOM and VEX information is matched to software components by vulnerability-management tools.

The project began as one self-contained repository: se-verification-vulnerability-matching

The immediate objective is to test whether the existing SE operational-identity construct exposes consequential, inspectable failures or underdetermination in production security systems.

Central Research Contribution

The proposed contribution is an exploratory feasibility study of finite conformance auditing for security-relevant identity semantics in software vulnerability matching, demonstrated on software vulnerability matching.

The unit of contribution is identity conformance, not scanner disagreement.

A conventional differential or metamorphic test may show that changing an input representation changes a scanner result. The SE audit must add structural explanation by identifying:

  • the two compared inputs;
  • the applicable identity commitment and its provenance;
  • the operational relationship induced by the implementation;
  • the identity axis on which the relationships diverge;
  • the affected security behavior; and
  • a finite, reproducible witness of the divergence.

Example form:

Inputs A and B identify the same content-addressed artifact under the applicable object-identity commitment, but the scanner separates them after a locator transformation because repository_url participates operationally in its PURL comparison. The witness therefore localizes the divergence to locator/object identity rather than merely reporting changed output.

Scientific Boundary

This work belongs to Structural Explainability / Operational Identity. It asks:

  • What entity is the security control treating as the same entity?
  • What rule or convention supports that treatment?
  • Does the implementation honor that commitment?
  • If no sufficient commitment exists, where is identity underdetermined?
  • Can the divergence be expressed as an inspectable finite witness?

It does not ask whether a security claim, evidence chain, or composed system remains assured under composition or change. That is the separate Structural Assurability work.

The abstract and introduction should state this boundary explicitly.

Research Questions

RQ1. Identity commitments What identity commitments governing vulnerability matching are formally specified, documented by tools, established by community convention, or left underspecified?

RQ2. Operational identity What operational identity relations do real SCA/VEX implementations induce under controlled identity transformations?

RQ3. Structural divergence Where do operational identity relations diverge from the applicable identity commitments, and on which structural identity axes do those divergences occur?

RQ4. Security consequences What cybersecurity consequences result from those divergences, including missed findings, inappropriate suppression, duplicated findings, incorrect scope propagation, and indeterminate behavior?

RQ5. Explanatory value, if supported by the data Does structural positioning localize or explain identity-sensitive failures more precisely than conventional input/output differential testing?

RQ5 is not required to justify the study. It should be answered only if the study design and resulting evidence support a defensible comparison.

Two Principal Failure Classes

The study will distinguish at least two classes of findings.

Identity nonconformance

An applicable identity commitment exists, and the implementation's behavior conflicts with it.

Identity underdetermination

A security decision depends on a rule of sameness that the governing specification, documentation, or convention does not adequately determine.

Underdetermination is a result, not a failed experiment. It may prove as important as direct nonconformance.

Study Objects

The initial implementations are:

  • Grype;
  • Trivy; and
  • their associated VEX-processing and vulnerability-database behavior.

The study will use controlled pairs of SBOM, VEX, package, image, and component representations. Candidate transformation families include:

  • PURL qualifiers present, absent, reordered, or changed;
  • OCI mutable tag versus immutable digest;
  • alternate locator for the same artifact;
  • changed artifact hash with stable name and version;
  • CPE versus PURL representation;
  • version syntax and normalization;
  • architecture;
  • namespace;
  • package versus subcomponent relationship;
  • product scope versus component scope; and
  • other transformations justified by the standards, tool documentation, public defects, or the frozen sampling protocol.

The final registered transformation set should contain approximately 20-40 controlled cases, with one deliberately isolated identity-relevant change per pair whenever possible.

Identity-commitment model

The declared side and operational side must remain separate.

Declared or expected side

Each case will record the source and status of the relevant identity commitment. The core provenance classes are:

  • formal_standard;
  • tool_documentation;
  • community_convention; and
  • underspecified.

An engineering-validation fixture may also carry a clearly separated known expectation derived from an already-public defect report or documented example. Such a fixture validates the harness; it is not prospective research evidence.

implementation_induced is not a declaration-provenance class. Implementation-induced identity is an observation produced by the experiment.

Operational side

The scanner adapters will record what each implementation actually does, including where available:

  • vulnerability association;
  • VEX applicability;
  • suppression decision;
  • finding duplication or disappearance;
  • scope propagation;
  • matcher or provenance information; and
  • errors, warnings, or indeterminate outcomes.

Audit relationship

The SE audit compares the applicable commitment with the operational relationship and produces a positioned witness when they diverge or when the decision is underdetermined.

Each adjudicated witness should record:

  • case identifier;
  • inputs A and B;
  • transformation family;
  • expected identity relationship;
  • commitment provenance and source;
  • observed identity relationship;
  • scanner and scanner version;
  • vulnerability-database snapshot;
  • SE identity axis;
  • security consequence;
  • reproduction procedure;
  • reviewer disposition; and
  • adjudication notes.

Repository design

Create one repository inside the structural-explainability organization:

se-verification-vulnerability-matching

The repository should be self-contained for reviewers and conform to se-constitution. The scientific basis is SE-210 and its executable companion, se-verification-operational-identity.

Use the existing SE-210 implementation as the reference audit core rather than reimplementing it. Pin an exact release or commit instead of depending on main.

Bring in se-regimes, se-mapspec, or se-contract-kit only if a concrete implementation requirement justifies the dependency. They are not prerequisites merely because they exist.

Proposed repository layout:

se-verification-vulnerability-matching/
├── contracts/
├── cases/
├── data/
│   └── validation/
├── protocol/
├── results/
├── src/
├── tests/
├── CITATION.cff
├── README.md
└── SE_MANIFEST.toml

Responsibilities:

  • contracts/: experimental identity commitments, provenance, and source references;
  • cases/: controlled transformation-case definitions;
  • data/validation/: only already-known public cases used to validate the machinery;
  • protocol/: research questions, inclusion/exclusion criteria, freeze policy, adjudication procedure, and analysis plan;
  • results/: prospective results, contains no prospective observations until the registered study may begin;
  • src/: scanner adapters, transformation execution, observation collection, SE-210 adapter, and witness production; and
  • tests/: unit, integration, schema, and reproducibility tests.

One documented command should run the validation audit end to end.

Do not create a separate vulnerability-matching-spec repository at this stage. The identity commitments are experimental research data. Promoting them to a general domain specification before the study would incorrectly imply that the project is authoring the standard it intends to examine.

Engineering validation before Stage 1

Known public cases may be reproduced before the registered report is submitted, strictly as engineering validation.

Initial fixtures:

  1. The Grype OCI PURL repository_url matching defect documented in Issue #3657 and fixed through PR #3659.
  2. Trivy's documented PURL qualifier behavior, including the asymmetric qualified/unqualified comparison and architecture-sensitive comparison.

These cases answer only:

Can the machinery detect and correctly position a known identity-sensitive behavior?

They must not be counted as prospective findings, included in the discovery corpus, or used to answer the research questions.

Registered-report boundary

Before Stage 1, the project may:

  • implement and test the harness;
  • define transformation families;
  • write and validate schemas;
  • freeze scanner and database-snapshot procedures;
  • reproduce already-public validation fixtures;
  • test synthetic inputs that verify machinery rather than answer the RQs;
  • write the protocol, adjudication guide, and analysis code; and
  • confirm that the complete workflow is reproducible.

Before Stage 1, the project must not:

  • execute the prospective discovery corpus;
  • inspect prospective scanner outcomes;
  • tune transformations in response to prospective results;
  • move newly observed cases into validation data; or
  • populate results/ with evidence intended to answer the RQs.

The prospective corpus, inclusion rules, exclusions, transformations, expected analyses, and stopping rules must be frozen before execution.

Reproducibility strategy

The study must freeze or record:

  • exact scanner versions and immutable container or binary digests;
  • scanner configuration;
  • vulnerability-database snapshot identifiers and acquisition times;
  • SBOM and VEX formats and versions;
  • all input artifacts and content hashes;
  • transformation definitions;
  • operating environment;
  • commands and exit status;
  • raw scanner output;
  • normalized observations;
  • generated witnesses; and
  • human adjudication records.

The harness should preserve raw evidence and derive normalized results without overwriting the source observations.

The protocol must specify what happens if an implementation cannot be fully frozen, a database snapshot cannot be redistributed, or a scanner exposes insufficient matcher provenance. These limitations should be recorded rather than repaired through undocumented inference.

Human adjudication

Machine-produced candidates are not automatically findings. Each candidate witness requires human review to determine whether it represents:

  • confirmed nonconformance;
  • confirmed underdetermination;
  • permitted difference;
  • tool defect unrelated to identity;
  • experimental artifact;
  • insufficient evidence; or
  • another predeclared disposition.

For reliability, a subset should be independently reproduced and adjudicated without exposing the expected-answer field until after the independent judgment is recorded.

The protocol should freeze:

  • reviewer instructions;
  • evidence available to the reviewer;
  • independent-review sample size or sampling rule;
  • disagreement-resolution procedure;
  • conditions for consulting tool maintainers; and
  • rules for responsible disclosure of previously undocumented defects.

Division of work

Principal investigator

The principal investigator owns:

  • the SE identity-conformance model;
  • research questions and study design;
  • transformation families and sampling protocol;
  • identity-commitment provenance model;
  • expected outcomes where an applicable commitment exists;
  • audit harness and scanner adapters;
  • SE-210 integration;
  • scanner and database freeze strategy;
  • adjudication protocol;
  • analysis plan; and
  • registered-report manuscript.

Graduate researcher

The graduate researcher will:

  • execute the frozen study protocol;
  • independently check a defined subset of results;
  • reproduce surprising cases manually;
  • inspect relevant standards and tool documentation;
  • maintain the evidence corpus;
  • adjudicate candidate witnesses under the frozen guide;
  • classify divergences and consequences;
  • visualize the results;
  • document anomalies and experimental limitations; and
  • contribute to interpretation and reporting.

This is substantive empirical research while keeping theory construction and experimental design under the principal investigator's control.

Analysis outputs

The primary empirical representation will be a matrix with fields equivalent to:

  • Commitment
  • Provenance
  • Operational identity
  • Divergence type
  • SE axis
  • Security consequence
  • Reproduction status

Planned summaries may include:

  • cases by commitment provenance;
  • outcomes by scanner and transformation family;
  • nonconformance versus underdetermination;
  • security consequences by SE identity axis;
  • adjudication disposition and reviewer agreement; and
  • detailed finite witnesses for representative failures.

Counts alone are insufficient. The paper's distinctive evidence is the positioned witness and the explanatory relationship between commitment, operational behavior, identity axis, and security consequence.

MSR fit and manuscript framing

The empirical objects should be presented as public software-security repositories and artifacts:

  • scanner source repositories and releases;
  • matching implementations and their evolution;
  • SBOM and VEX documents;
  • vulnerability advisories and databases;
  • issue reports and fixes; and
  • reproducible transformation and witness datasets.

Structural Explainability supplies the analytical method. Cybersecurity supplies the consequence. The manuscript must remain visibly grounded in empirical software-repository and software-evolution research to satisfy MSR scope.

The central claim should be:

We introduce a finite conformance audit for security-relevant identity semantics and demonstrate it on software vulnerability matching.

The manuscript does not claim:

  • a new universal definition of software identity;
  • a general vulnerability-matching standard;
  • a solution to scanner inconsistency; or
  • a new cybersecurity framework.

Publication plan

Springer Journal of Empirical Software Engineering (ESME)

Primary venue:

  • MSR (Mining Software Repositories) 2027 Registered Reports
  • Abstract: November 13, 2026
  • Stage 1 report: November 20, 2026
  • Stage 1 decision target: February 4, 2027
  • MSR presentation: April 26-27, 2027, Dublin
  • Completed study to EMSE: September 30, 2027, following the registered-report process

Official track page: https://2027.msrconf.org/track/msr-2027-registered-reports

The Stage 1 submission is a protocol, not a results paper. The project should use this feature deliberately: establish the method, instrument, corpus-construction rules, and analysis before observing the prospective outcomes.

Milestones

September 2026: establish the project

  • Create se-verification-vulnerability-matching.
  • Apply se-constitution requirements.
  • Inspect and pin the SE-210 executable core.
  • Define the minimal scanner-adapter and witness interfaces.
  • Record the registered-report boundary in the repository.
  • Add the Grype and Trivy public validation fixtures.

October 2026: freeze the protocol

  • Complete the transformation taxonomy and case schema.
  • Complete the identity-commitment provenance model.
  • Freeze scanner-version and database-snapshot procedures.
  • Freeze inclusion, exclusion, sampling, and stopping rules.
  • Complete the adjudication guide and reliability procedure.
  • Complete the planned analyses and empty result templates.
  • Confirm that one command reproduces the validation cases.
  • Draft the six-page Stage 1 report.

November 2026: submit Stage 1

  • Submit the abstract by November 13.
  • Submit the registered report by November 20.
  • Tag and archive the submitted protocol and implementation state.
  • Preserve the unexecuted prospective corpus boundary.

December 2026-February 2027: revision and preparation

  • Respond to reviews.
  • Revise the protocol without examining prospective results.
  • Finalize training material for independent execution and adjudication.
  • Tag the accepted protocol and freeze the executable study release.

After Stage 1 authorization: execute the study

  • Run the frozen prospective corpus.
  • Preserve raw outputs and normalized observations.
  • Independently reproduce the registered subset.
  • Adjudicate candidate witnesses.
  • Follow the disclosure protocol for new defects.
  • Produce the registered analyses and visualizations.

By September 30, 2027: complete the article

  • Complete the EMSE manuscript.
  • Publish the reproducibility package permitted by licenses and disclosure constraints.
  • Clearly distinguish preregistered analyses from any labeled exploratory follow-up.

Risks and controls

Risk: the work appears to be ordinary differential testing

Control: Make the declared commitment, structural identity axis, and positioned finite witness load-bearing in the protocol, data model, and paper.

Risk: no formal identity rule exists

Control: Treat underdetermination as a registered finding class. Do not manufacture a normative rule merely to permit conformance scoring.

Risk: the project becomes another SE infrastructure effort

Control: Use SE-210 and the existing executable audit. Add other SE packages only when required by a concrete implementation need.

Risk: scanner or database evolution harms reproducibility

Control: Pin immutable versions, preserve identifiers and hashes, archive permitted inputs and outputs, and document any components that cannot be redistributed.

Risk: known cases contaminate prospective findings

Control: Isolate them in data/validation/, label them as public engineering fixtures, and exclude them from every RQ analysis.

Risk: tool opacity prevents causal claims

Control: Separate directly observed behavior from inferred matching logic. Record insufficient provenance as a limitation or indeterminate case rather than overstating the mechanism.

Risk: MSR scope is unclear

Control: Center public repositories, implementation evolution, security metadata, reproducible artifacts, and empirical analysis. Use SE as the analytical method rather than presenting the paper primarily as a new formalism.

Definition of ready for Stage 1 submission

The project is ready when:

  • the repository exists and conforms to se-constitution;
  • one command runs the complete engineering-validation audit;
  • the SE-210 adapter produces positioned witnesses;
  • two or three already-public validation cases pass end to end;
  • the transformation families and prospective case-generation rules are frozen;
  • the commitment-provenance model is frozen;
  • scanner versions and database-snapshot procedures are frozen;
  • the inclusion, exclusion, and stopping rules are written;
  • the human adjudication and independent-reproduction procedures are written;
  • the analysis plan is complete;
  • the prospective corpus has not been executed;
  • results/ contains no prospective RQ evidence; and
  • the six-page registered report clearly establishes MSR relevance.

First actions

  1. Create the se-verification-vulnerability-matching repository in structural-explainability.
  2. Add the minimal constitution-conforming project skeleton.
  3. Inspect public API of se-verification-operational-identity and pin exact version or commit to use.
  4. Define the case, commitment, observation, adjudication, and witness schemas.
  5. Implement the Grype repository_url case as the first engineering-validation fixture.
  6. Implement the Trivy qualifier behavior as the second engineering-validation fixture.
  7. Write the prospective-study boundary into protocol/freeze-policy.md before adding discovery cases.
  8. Open the Stage 1 manuscript with the contribution statement, RQs, MSR relevance, and the explicit Structural Explainability versus Structural Assurability boundary.