INL “Reviewing Smarter” RFI: What “Every Output Traceable to Source Documents” Could Mean
Date: August 25, 2026 · Author: Dmitrii Zatona
On June 23, 2026, Idaho National Laboratory posted a request for information on SAM.gov asking vendors about AI tools to accelerate nuclear licensing review for an international regulator. The document is explicit about the constraint that comes with acceleration: outputs must stay grounded, and there must be “every output traceable to source documents.” That phrase, with its siblings fully cited, auditable, and defensible, appears at three separate levels of the document. Read as engineering, the wording is compatible with at least four different properties of sharply increasing strength — and the document does not say which one it means. Only the weakest of them ships in the tooling this article surveys. This article maps the four readings, what tooling exists for each, and the engineering problems living in the gaps between them.
1. What was published
The notice is INL-26-047, “*Request for Information* Reviewing Smarter: AI Tools for Building International Nuclear Regulatory Capacity,” posted on SAM.gov on June 23, 2026 by Battelle Energy Alliance, LLC — the management and operating contractor that runs Idaho National Laboratory for the Department of Energy under contract DE-AC07-05ID14517 . On SAM.gov’s own classification the posting is a Special Notice, not a Sources Sought; “Request for Information” lives in the title string. The notice was published twice the same day: a first version posted at 14:48 UTC with a July 10 response deadline was replaced 64 minutes later by the surviving version with a July 8 deadline, and the superseded PDF was deleted from SAM.gov at the moment of replacement, leaving a 215-byte size difference as the only documented delta. Responses were due July 8, 2026, by email, and the notice auto-archived the same day. Questions and submissions were directed to two INL addresses, Katya Le Blanc and Chase Egbert.
The document states in bold that it is not a request for proposals and commits the issuing authority to no procurement; it describes itself as input to a potential pilot and deployment pathway, with an anticipated sequence of formal solicitation, a scoped pilot, and production deployment. As of August 25, 2026, no follow-on solicitation or award traceable to the notice appears in SAM.gov or USAspending searches. That statement carries a stronger-than-usual caveat: Battelle Energy Alliance is an M&O contractor, so an engagement contracted through it would be a BEA subcontract — a category that does not appear in federal award databases by design — and BEA runs its solicitations through a vendor portal that is not publicly searchable. If the anticipated solicitation ran through that channel, it would leave no trace in the sources this article can search.
Method note: this is the second procurement teardown in a series; the first examined the FBI Threat Screening Center’s “traceable lineage” notice .
2. What the requirement actually says
The RFI’s purpose section sets the frame: design certification and combined license reviews have historically taken years; the initiative aims to compress review, research, and drafting timelines by years while keeping the applicant–regulator interface safe, traceable, and defensible. Acceleration, it says, must not come at the expense of rigor, explainability, or independent regulatory judgment: outputs must remain fully cited, auditable, and grounded in source documents so that both sides of the interface can rely on them. The scope of interest, in full:
The issuing authority is interested in AI tools and platforms that support international regulator’s licensing review function, including but not limited to:
- Intelligent search and discovery across large repositories of technical, safety, and licensing documentation
- Understanding and relating requirements, guidance, licensing precedent, and prior regulatory reviews
- Identification of resolved topics, open items, and novel or safety-significant issues for prioritized review
- Mapping of national/U.S. licensing bases to IAEA and Western European Nuclear Regulators Association requirements
- Automated drafting and generation of review-supporting documentation, with full citation and source traceability
- Explainable reasoning, with every output traceable to source documents
- Other capabilities as relevant for international regulator and/or applicant assistance
Read as an engineering list: intelligent search and discovery is retrieval. Understanding and relating requirements, guidance, licensing precedent is cross-document linking. Identification of resolved topics, open items, and novel or safety-significant issues is classification and prioritization — technically the heaviest item on the list, and the one where the provenance question is sharpest, because the output that matters is a judgment (this issue is novel, this one is resolved) whose derivation someone will later need to see. Mapping of national/U.S. licensing bases to IAEA and WENRA requirements is a crosswalk across four distinct normative corpora — national law, U.S. precedent, IAEA safety requirements, WENRA reference levels — and the crosswalk is itself a derived artifact with provenance of its own. Automated drafting with full citation and source traceability is retrieval-augmented generation with a citation layer. And every output traceable to source documents is the phrase whose meaning the rest of this article is about.
One distinction has to be made before the tooling can be evaluated, because the tooling ships one property and the wording can be read as either. Citation correctness is a relation between an output and a source: the source supports the statement. The attribution literature defines it exactly that way — output “verified against an independent, provided source” (Rashkin et al., Computational Linguistics 49(4)) — and it is checkable by a person with no access to the generating system. Assembly provenance is a relation between an output and a process: these sources, in this form, were what the system actually used to produce this output. The two come apart in both directions. RARR (ACL 2023) “automatically finds attribution for the output of any text generation model” — after generation, for text the sources played no role in producing; post-hoc citation is a standard, benchmarked architecture, not an edge case (ALCE , EMNLP 2023). In the other direction, retrieved context that shaped an answer can go uncited. Citation precision and citation recall are separately defined, separately measured properties (Liu et al. , Findings of EMNLP 2023); neither is guaranteed by the existence of a citation mechanism, and neither one measures assembly.
The RFI asks for traceability at three levels of the document: the scope bullet above; the live-demonstration criteria, which include citation and source traceability for every answer, checked in real time on a non-trivial licensing scenario; and the assessment criteria, whose second item is traceability and explainability: every output cited and auditable. A property demanded three times, at demonstration and at assessment, reads as load-bearing. Which property it is, the text does not fix.
A second thing the text leaves open is who is asking. The document consistently says “the issuing authority” and never equates that phrase with INL, BEA, DOE, or any named regulator. The authority is interested in tools supporting an international regulator’s licensing review function; a demonstration criterion is applicability to international partners deploying nuclear energy for the first time; one section describes down-selected respondents helping to approach potential partners and assess their interest, while a later section anticipates a pilot on the regulator’s own documents and production in the regulator’s sovereign cloud tenant, supported through existing bilateral nuclear cooperation frameworks — none named. The potential partners of the earlier section and the definite article of the later one sit side by side in the text. The identity of the regulator, and whether “the issuing authority” denotes the laboratory itself, a program office, or a partner authority, is not established by the document or by any public source consulted for this article.
3. The assurance ladder
“Traceable to source documents” admits at least four operational readings. Each one is a different property, with a different verifier and a different thing that must be trusted.
Rung 1 — checkable citation. The property: a reviewer with access to the corpus can open the cited source now and confirm that it supports the statement. The verifier is a human, at review time, inside the system’s trust domain. What is trusted: the tool’s display and the corpus’s current state. The tool class is the citation layer of grounded-generation products. This rung is on the shelf.
Rung 2 — recorded execution. The property: the system retains a record of assembly — which chunks of which documents, which index, which model and parameters produced the output. The verifier is the operator, plus whoever the operator grants access. What is trusted: the operator’s store, because the record is mutable by its keeper. The tool class is observability and tracing. This rung exists as raw material, examined in Section 4.2.
Rung 3 — reproducibility. The property: the recorded tuple can be re-executed to the same output. The verifier is anyone holding the full execution environment. What is trusted: environment pinning and deterministic decoding — neither of which is offered contractually today (Section 4.4). No surveyed tool class ships this as a contract.
Rung 4 — verifiable attestation. The property: a counterparty who does not trust the operator can verify a bounded, signed statement about assembly — that a specific key attested a specific tuple, and, where a transparency log participates, that the attestation entered an append-only structure by a point in time. The verifier is any party, adversarially. What is trusted: cryptographic assumptions plus explicitly stated residual ones. The machinery exists, standardized, in other domains — certificates, software artifacts, media files — with no standardized counterpart for generative pipelines identified in the survey behind this article (Section 4.2).
The RFI’s words are compatible with all four readings. Fully cited reads naturally at rung 1. Auditable is used at every rung; the word leaves open which store is audited and who keeps it. Defensible — over review cycles that run years, corpora that live in revisions for decades, and a two-sided interface between an applicant and a regulator in different jurisdictions — is the word that reads furthest up the ladder, and independent regulatory judgment names the verifier question without answering it. Nothing in the text selects a rung, and this article does not impute one. What can be stated as engineering: each rung up is qualitatively more expensive than the last, and the distance between what rung 1 checks and what rung 4 proves is where the following four problems live.
4. Four engineering problems
4.1 Citation correctness is not assembly provenance
What the shipping citation layers deliver is documented precisely, and it differs by vendor. Anthropic’s Citations returns cited text with character or page offsets into the supplied documents; the API extracts the cited text itself, so citations are “guaranteed to contain valid pointers to the provided documents” — while the span selection is model output, parsed into that format. AWS Bedrock’s citation objects carry deterministic offsets locating the cited span in the generated answer, with retrieved-chunk references on the source side. Google Vertex’s grounding metadata returns byte-offset segments plus per-reference scores expressing “confidence that the reference supports the claim” — the schema itself types the source-to-claim attachment as a scored assertion rather than a record. Azure OpenAI’s On Your Data returns no offsets at all: inline [doc1] markers anchored by prompt guidance, with retrieved-but-uncited documents surfacing under an UNCITED_REFERENCE label. That last layer also demonstrates the lifecycle of this contract: Azure has scheduled On Your Data for retirement on October 14, 2026. A citation format at this layer is a product surface, not a standard.
None of the four records which context actually shaped which output tokens — Azure disables even logprobs in grounded mode — and none binds a citation to a revision of the source corpus. The citation edge is produced by the same decoding process that produces the prose, with the cited text then extracted deterministically in three of the four products.
The consequence is direct. Rung 1’s check — a reviewer confirming that the cited source supports the sentence — can pass in full on an answer whose actual assembly is unrecorded, including an answer whose citations were attached after the fact. Correctness and provenance are independent properties. Verifying the first, at any scale and with any diligence, establishes nothing about the second.
4.2 A record is not evidence
Rung 2 exists as raw material. The OpenTelemetry GenAI semantic conventions — the closest thing to a standard trace schema — define span types for retrieval and tool execution with named attributes; the conventions carry a Development stability badge, and nothing in them addresses integrity of the captured trace. Products record on these shapes and keep the records as workflow assets: LangSmith ’s hosted service “retains trace data for 180 days from ingestion” and then deletes it; Bedrock’s invocation logging is disabled by default, writes to the operator’s own bucket, and has a delete-configuration API. The U.S. federal control catalog names the missing properties exactly: NIST SP 800-53 AU-9 requires protecting audit information from modification and deletion, and AU-10 requires “irrefutable evidence” of who produced what — with the catalog’s own verifier being authorized individuals, not an outside party.
Whether an operator-held record is enough depends entirely on who must rely on it. The RFI describes a two-sided interface — applicant tooling and regulator tooling — with production in a sovereign cloud tenant. It does not state that either side must verify the other’s records; the verifier model is among the things the text leaves open. If the answer is ever “a party outside the operator’s trust domain,” the machinery for bounded verifiable assertions already exists, and its members differ in ways that matter. Certificate Transparency (RFC 9162 ) and Sigstore’s Rekor are transparency logs: append-only Merkle trees with inclusion proofs and consistency proofs. in-toto is signed attestations checked against a layout policy, with no log inherent to the model; SLSA v1.0 layers build-isolation levels onto such attestations. C2PA is signed manifests cryptographically bound to asset bytes, with no append-only log in the model. Every member proves a statement of the same bounded shape: a key attested a specific tuple, and — where a log participates — the attestation was included by a point in time in a history that has not been rewritten. No member proves that capture was complete, that the attested process executed as described, or that a cited source semantically supports anything. A signed retrieval trace, if one existed, would inherit exactly those limits.
The diagram’s right side is one possible verifier model, not a requirement the RFI states. What the RFI does state is defensible. Read as engineering, a defensibility claim implies an audience for the defense; whom that audience includes is a decision the wording leaves open, and each candidate answer prices out at a different rung.
4.3 No surveyed contract propagates a source change
Licensing corpora live in revisions, on regulated cadences, for decades (Section 5). The sync semantics of the retrieval layer, per the vendors’ own documentation, are forward-only. Bedrock re-ingests changed documents — re-parsed, re-chunked, re-embedded — and removes deleted ones from the vector store, with only job history recording that a sync happened. Azure AI Search detects changes by timestamp, and “An indexer doesn’t track object deletion”; its native deletion-detection path requires blob versioning to be disabled, by design. Vertex AI Search refresh “replaces existing documents with updated documents with the same ID.” Anthropic’s Files API is the outlier: “Files cannot be modified or renamed after upload,” so a file id denotes immutable content — but no version chain links a replacement file to its predecessor, and the citation object identifies documents by request index and title, not by file id.
Checked across all four schemas: no citation object carries a revision or timestamp field that a consumer holding a past answer could test against the current corpus, and no vendor documents any mechanism that notifies consumers of past answers, or marks them stale, when a source changes. The standards layer can express the event without obliging anyone to emit it: W3C PROV defines invalidation — “Invalidation is the start of the destruction, cessation, or expiry of an existing entity” — as vocabulary, with no emission requirement; OpenLineage’s lifecycle-change and version facets are observations attached to runs; C2PA accretes manifests on the edited asset without reaching derivatives already distributed. The closest genuine infrastructure is Databricks’ Delta Sync index , which keeps a vector index synced to a versioned Delta table — a freshness contract from source to index, not from source to the outputs derived from it. Across the twelve vendor documentation sets and specifications surveyed for this article, no standardized, interoperable contract was identified that states what a source revision implies for artifacts derived from the prior revision (as of August 25, 2026; the survey list is in the sources). A literature search on the same date found no published work on citation staleness in derived artifacts of this kind; the finding rests on the specifications surveyed, not on the state of the literature.
The bite is concrete. A draft review section assembled in year two of a licensing proceeding cites a source paragraph superseded in year four; under 10 CFR 50.71(e) , the underlying safety analysis report is re-revised on a cadence not exceeding 24 months for the life of the license. Which derived artifacts a given revision touches, and who is obliged to find out, has no specified answer anywhere in the surveyed stack.
4.4 An archive is not a replay
At rung 3, auditable means the run can be reconstructed. Two obstacles stand, one of policy and one of mechanism. Policy: OpenAI’s seed feature is best-effort — “Determinism is not guaranteed” — and is conditioned on a backend fingerprint the caller can observe but not pin. Anthropic’s API reference states that results “will not be fully deterministic” even at temperature zero, and deprecates the temperature parameter for its newer models, so greedy decoding cannot even be requested. Mechanism: floating-point addition is not associative, and parallel GPU execution does not fix a reduction order; the resulting run-to-run variability in deep-learning inference is documented in the HPC literature (SC’24 Workshops ) — and the reduction order is not part of any evidence bundle any surveyed system records. Archiving the tuple — prompt, chunks, model id, parameters — therefore yields an archive, not a replay contract.
One middle case deserves its place on the ladder. WORM storage — S3 Object Lock in compliance mode, where a locked object version “can’t be overwritten or deleted by any user, including the root user” — lets an operator bind its own hands and gives an auditor a configuration to inspect. It exports no proof: the verifier still trusts the storage provider and the capture pipeline in front of it. That hardens rung 2 without reaching rung 4. The strongest current machinery in the vicinity, confidential-computing inference — TEE-attested serving stacks with a transparency ledger of stack artifacts (Azure confidential inferencing , preview, with NVIDIA GPU attestation ) — attests which software served the answer, not which corpus revision and chunks assembled it. Execution integrity for the stack exists in preview form; for the assembly of a specific answer, nothing equivalent appears in the documentation surveyed here.
5. The scale
The corpora these tools would work over are documented in agency numbers. For the NuScale design certification — the first small modular reactor design the NRC certified — the Department of Energy’s announcement puts it plainly: “The 12,000-page application took less than 42 months to review,” with more than two million pages of supporting documents available for audits. NuScale’s own docketed lessons-learned report counts “over a quarter million review hours, about two million pages of documentation” plus roughly 100 gigabytes of test data and some 40 meetings of the Advisory Committee on Reactor Safeguards. The application itself moved through six full docketed revision states, Revision 0 (December 2016) to the certified Revision 5 (July 2020); the AP1000 certified design moved from Revision 15 (2006) to Revision 19 (2011) through one amendment. On the regulator’s side, the NRC’s public ADAMS library holds “more than 3 million full-text documents” released since 1999, growing by several hundred a day — and ADAMS accession numbers function as document identifiers inside the CFR’s own incorporation-by-reference tables. Meanwhile the time budget is being compressed by statute and order: the ADVANCE Act’s §207 sets 18 months to a safety evaluation and environmental statement and 25 months to a final combined-license decision for qualifying applications, “to the maximum extent practicable,” and EO 14300 sets an 18-month deadline for new-reactor decisions.
These figures carry one operational point and no more: any assurance procedure that relies on humans re-verifying citations one at a time operates against corpora measured in millions of pages, revision histories measured in years, and obligations measured in decades — on a clock that legislation is shortening. They establish no error rate and no cost; no such arithmetic appears in this article.
6. What has been tried
Two related but distinct artifacts define the current state of AI-for-licensing work at the national laboratories, and the primary record keeps them apart. The first is a demonstration: on March 26, 2026, DOE announced that Everstar’s Gordian platform, built on Microsoft Azure and working with INL and Argonne, converted the preliminary documented safety analysis of NRIC’s generic high-temperature gas reactor into sections equivalent to an NRC license application: “The final 208-page document took one day to generate,” against a stated typical four to six weeks of manual drafting; expert review characterized the output’s quality, with no quantitative rubric published. The antecedent INL–Microsoft collaboration (July 2025, NRIC-funded) states the scope of this class of tool exactly: it “automates the process of constructing licensing documents for subsequent human verification” and performs no analyses.
The second artifact is an evaluation methodology. The NRIC white paper INL/RPT-25-87369 (March 2026) is the source of the widely repeated 50-percent figure: its abstract states the approach “has the potential to reduce both document development time and regulatory review cycles” by as much as half. The document proposes the measurement rather than reporting one — SME checklist scoring against a DOE-approved benchmark safety analysis, followed by a DOE mock review — and the tool under evaluation there is Microsoft’s generative-AI permitting accelerator, not Gordian. As published, the potential is a stated projection with a defined test plan behind it.
On the regulator’s side, the NRC’s AI use-case inventory lists a deployed text-retrieval-and-generation tool supporting licensing and oversight work; the agency’s FY26 AI strategic plan states that it “leverages quality, traceable datasets (e.g., ADAMS, EDW) but seeks greater centralization and traceability”; and the NRC-commissioned regulatory gap assessment (October 2024) reads the existing framework as one that “calls for control of information as well as auditability, but not explicit explainability.” The properties in the RFI’s wording — auditability, traceability, explainability — appear in the regulator’s own current documents as open work items, not settled expectations.
7. Who would build it
The RFI’s mandatory eligibility criteria describe its intended respondent pool precisely, and they are worth reading in full:
- Nuclear AI experience. Demonstrated, deployed experience applying AI specifically to nuclear energy, licensing, compliance, or safety documentation (commercial, national laboratory, or regulatory deployments). Academic or purely conceptual claims do not satisfy this criterion.
- U.S. regulatory pedigree. Tool has been developed, validated, deployed, or otherwise substantively engaged with the U.S. Department of Energy (DOE) and/or U.S. Nuclear Regulatory Commission (NRC), or a U.S. national laboratory operating under DOE/NRC frameworks. Respondents must describe the nature and current status of that engagement.
- Reactor-class coverage. Demonstrated applicability across reactor types from micro to large (see Section 2).
- Security and sovereignty. Ability to deploy within various international sovereign cloud tenant or otherwise meet the issuing authority’s data residency, security, and confidentiality requirements.
- Availability for live demonstration within the RFI period (see Section 6).
The filter selects for vendors already deployed in nuclear documentation work under DOE- or NRC-adjacent frameworks. The public record names the participants in the work described in Section 6 — Everstar and Microsoft, with INL and Argonne — in DOE’s own announcements. Participation in that cited work is all the record establishes; it does not establish that any named vendor satisfies the five criteria, several of which — reactor-class coverage, sovereign-cloud deployment, demonstration availability — no agency publication reviewed for this article addresses. Whether the anticipated down-select occurred, and with whom, is not publicly visible; per Section 1, a BEA subcontracting process would leave no public trace.
8. What does not exist as a standard
The teardown as a table:
| Reading of “traceable” | Closest existing tooling | What it lacks |
|---|---|---|
| Rung 1: checkable citation | Citation layers of grounded generation (extracted spans, offsets) | Binding to a corpus revision; any record of assembly |
| Rung 2: recorded execution | OTel GenAI conventions (Development status); tracing products; vendor invocation logs | Integrity: stores are operator-mutable, opt-in, retention-bounded |
| Rung 3: reproducibility | Seed parameters; archived tuples; WORM storage | A replay contract: vendors disclaim determinism; execution environment not pinnable |
| Rung 4: verifiable attestation | CT/Rekor logs; in-toto/SLSA attestations; C2PA manifests; TEE serving-stack attestation (preview) | Any standardized application to the assembly tuple of a generated answer |
Every row exists somewhere; in the tooling surveyed for this article, no row is complete for this use. The record-keeping obligations that such machinery would serve are not hypothetical — they are already codified in the corpora the RFI’s scope itself names: U.S. licensing bases, IAEA requirements, WENRA reference levels. U.S. quality-assurance records must be “identifiable and retrievable” (10 CFR 50 Appendix B , Criterion XVII); licensees must keep “adequate safeguards against tampering with, and loss of records” (10 CFR 50.71(d) ); departures from a certified design must stay “available for audit until the date of termination of the license” (10 CFR 52.63(b)(2) ). The IAEA’s management-system requirements make records “readable, complete, identifiable and easily retrievable” (GSR Part 2 ), WENRA’s reference levels track the same wording, and GSR Part 1 obliges the regulator itself to formally record the basis of its decisions. Which of these frameworks would govern the unidentified regulator is likewise not established; they are cited here as the corpora the RFI’s scope names, not as that regulator’s governing law. In every instrument reviewed, the verification model named in the text is institutional — officer certification, approval levels, regulator audit; no provision specifying how a party outside that institutional relationship verifies record integrity was identified in the instruments reviewed for this article. That seam is where the ladder’s upper rungs would attach. And the RFI’s own demonstration criteria name applicability to partners deploying nuclear energy for the first time — regulators that the IAEA’s milestones guidance describes as building review capacity while their programs advance. In that setting, the question of what a regulator can independently verify reads less like a corner case of the wording than like its center.
A disclosure: I am working on a specification in this area — verifiable provenance for generated outputs — and that work is separate from, and not linked in, this analysis.
The RFI’s wording is on the page, three times, at three levels of the document. Which of the four readings it means is not — and above the first rung, nothing surveyed here puts the tooling on the shelf.
If you are building a response to this document and the provenance or attestation piece has to be designed and built, that is contract work I take on.