NSF Solicitation 26-512: Building Data Infrastructure for AI-Driven Scientific Discovery
According to the National Science Foundation, solicitation 26-512, Unlocking Dataset Value for AI-Enabled Scientific Discovery, is now open for proposals.
Shane Barrett·updated July 27, 2026

The program is not a model-development call in the narrow sense: it targets the data infrastructure and research workflows required to make existing scientific datasets operational for AI systems. For ML researchers, the relevant unit of progress is therefore the pipeline—metadata, integration, provenance and automated analysis—not a new architecture claim.
The dataset, not the model, is the intervention
NSF frames the solicitation around extracting additional scientific value from datasets collected for earlier purposes. The proposed work may apply AI to feature extraction, metadata generation and the integration of multiple datasets. It may also augment or harmonize existing collections so they can support AI data pipelines and automated analysis.
That scope matters because scientific datasets commonly fail at the interface between storage and inference. A model may be parameter-efficient in isolation yet remain unusable when its inputs lack consistent schema, adequate metadata or cross-dataset alignment. The solicitation directs attention to these upstream constraints. It treats dataset preparation as research infrastructure rather than a preprocessing afterthought.
The agency also calls for robust pipelines that support automated analysis of existing datasets and comparable AI use cases. This shifts the evaluation problem. A credible proposal will need to establish whether a workflow remains valid beyond a single curated benchmark split, rather than merely reporting a favorable model score.
Security, integrity and governance are explicit requirements
NSF states that proposals should address dataset security and integrity, along with governance and processes through which scientific communities can contribute to datasets. These are not administrative appendices. They are conditions on reproducibility.
For code-oriented teams, this implies that a released training script alone would be incomplete evidence. The operational artifact is the full data path: ingestion, feature extraction, metadata generation, dataset integration and the controls that preserve the integrity of each step. Ablation studies can isolate a model component; they cannot compensate for untracked changes in the input corpus or ambiguous provenance.
The solicitation encourages use of existing resources, including NSF data platforms, the NSF Integrated Data Systems and Services program, the NSF-led National AI Research Resource, the Genesis Mission platform and other national infrastructure. It also leaves room for partnerships with philanthropy, private industry and non-profit organizations.
What practitioners should inspect
The practical question is whether a dataset can support repeated machine analysis without requiring a bespoke reconstruction for every project. Teams considering a submission should test their assumptions at the pipeline level: which fields are machine-actionable, where metadata must be generated, what integration step changes the latent representation of the data, and how integrity is retained after automated processing.
This is less visible work than training a new foundation model, but its computational overhead and failure modes are easier to audit. The same principle appears in data-dependent performance systems outside research; standardized feeds are a prerequisite for tools such as GoCharting’s GoChamp competition platform, even though scientific governance imposes a different set of constraints.
NSF’s stated objective is AI-enabled scientific discovery and innovation, including interdisciplinary investigation and work beyond the original motivation for collection. The decisive evidence will be whether funded workflows make that reuse measurable, governed and reproducible—not whether they attach AI terminology to an existing archive.