NIST TEVV-Athlon: Research Mapping
Source and Purpose
- NIST framework overview and request for input
- NIST AI 200-2 initial public draft
- Public-comment deadline for the initial draft: October 6, 2026.
NIST's TEVV-Athlon Framework provides a flexible, four-stage methodology for designing and conducting AI assessments based on organizational evaluation objectives.
The framework connects assessment goals to measurement Blocks, uses Events and Tools to collect relevant evidence, and provides a process for synthesizing and interpreting results.
This pilot draws on that methodology to investigate the applicability and limitations of claim-relative Structural Assurability.
The objective is to understand how the theory can help express relationships among evaluative claims, observations, evidence, evaluator conditions, and assumptions.
The investigation may also identify where the theory itself requires clarification or further development.
Initial Correspondence
| TEVV-Athlon stage | NIST methodology | Structural Assurability research |
|---|---|---|
| Articulate & Organize | Establish assessment goals and identify the system attributes or characteristics of interest. | Examine how an evaluation question can be represented as an explicit, claim-relative research question. |
| Define & Construct | Define measurement Blocks and the evidence needed to assess the selected characteristics. | Investigate the distinctions relevant to a specified claim and the evidence needed to distinguish them. |
| Apply & Measure | Select Events and Tools and collect evidence for the Blocks. | Model available observations and the conditions under which an evaluator obtains them. |
| Synthesize & Interrogate | Analyze the collected evidence, interpret results, and inform organizational decisions. | Examine what the specified observations establish under explicit modeling assumptions and which conclusions require additional justification. |
Initial Investigation
The pilot begins with the illustrative query-violation assessment in Section 3 of the NIST draft.
That example provides concrete assessment components through which to examine claim-relative Structural Assurability.
The first experiment introduces two researcher-defined claims and constructs finite models of their observation conditions.
The resulting witness pairs illustrate conditional claim-resolution properties of those models.
Further investigation will examine sampled interactions, human annotations, evaluator conditions, and the relationship between observed measurements and broader evaluative claims.
NIST's discussion of measurement validation and Goodhart's Law in Sections 4.5 and 4.6 provides context for that work.
Research Status
The initial mapping and finite-model experiment are documented. The investigation remains open.